Chatterbox

Premium

The voice cloning engine behind TextToSpeechAI cloned voices

Moderate Speed
Very Good Quality
Yes Cloning
15 Languages

About Chatterbox

Chatterbox is an open-source, zero-shot voice cloning model with MIT licensed code and weights. On TextToSpeechAI it powers every cloned voice: upload 6 to 60 seconds of clear audio and Chatterbox reproduces that voice in any of the languages listed on the voice cloning page, with no training step.

Key Features

Zero-Shot Voice Cloning

Clone a voice from 6 to 60 seconds of audio. No training run is needed.

Cross-Lingual

Clone from an English sample and generate speech in any of the 15 languages on the cloning page.

Measured Speaker Similarity

In our own tests Chatterbox matched the reference speaker more closely than the other licensed cloning engines we evaluated.

Commercially Licensed

MIT license for both code and model weights.

Use Cases

Voice cloning for content creation Multilingual narration in one voice Character voices for games Personal and brand voices

How to Use Chatterbox

  1. 1

    Create a free account

    Sign up for TextToSpeechAI to receive free starter credits.

  2. 2

    Upload a reference clip

    Open the voice cloning page and upload 6 to 60 seconds of clear speech from the voice you have the right to clone.

  3. 3

    Enter your text and language

    Type the text to speak and pick one of the 15 supported output languages.

  4. 4

    Generate the speech

    TextToSpeechAI renders the text in the cloned voice with Chatterbox on our GPU servers.

  5. 5

    Download or use the API

    Download the finished audio, or automate generation through the TextToSpeechAI REST API.

Chatterbox API

Generate speech programmatically using the TextToSpeechAI REST API.

curl -X POST "https://api.texttospeechai.com/v1/generate/" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "This is my cloned voice, created from a short recording.",
    "voice": "en_US-kristin-medium"
  }'

Frequently Asked Questions

Chatterbox is an open-source text-to-speech model that clones a voice from a short reference recording without any per-voice training. TextToSpeechAI uses it for every cloned voice.

Yes. Both the Chatterbox code and its model weights are released under the MIT license. You still need the right to use the voice you clone.

Upload between 6 and 60 seconds of clear speech from a single speaker. A quiet recording without music or background voices gives the closest match.

A cloned voice can speak the 15 languages listed on the voice cloning page: English, Spanish, French, German, Italian, Portuguese, Polish, Russian, Dutch, Turkish, Japanese, Korean, Chinese, Arabic and Hindi. The reference clip can be in a different language from the text you generate.

Cloned speech is rendered on our GPU servers and usually takes around a minute for a short passage, because the reference voice is processed again for each generation. Longer texts take longer.

Creating a cloned voice is free. Generating speech with it is charged per character, like any other voice, and new accounts receive free starter credits to try it.

Yes. Create a cloned voice through the REST API at api.texttospeechai.com with your account token, then reference it in your generation requests.

Technical Specs

  • Generation Speed Moderate
  • Output Quality Very Good
  • Voice Cloning Supported
  • Languages 15
  • GPU VRAM 4-8GB
  • Credits/1000 chars 5

Try Chatterbox Now

Generate your first audio free. No credit card required.

Start Free