Chatterbox
PremiumThe voice cloning engine behind TextToSpeechAI cloned voices
About Chatterbox
Chatterbox is an open-source, zero-shot voice cloning model with MIT licensed code and weights. On TextToSpeechAI it powers every cloned voice: upload 6 to 60 seconds of clear audio and Chatterbox reproduces that voice in any of the languages listed on the voice cloning page, with no training step.
Key Features
Zero-Shot Voice Cloning
Clone a voice from 6 to 60 seconds of audio. No training run is needed.
Cross-Lingual
Clone from an English sample and generate speech in any of the 15 languages on the cloning page.
Measured Speaker Similarity
In our own tests Chatterbox matched the reference speaker more closely than the other licensed cloning engines we evaluated.
Commercially Licensed
MIT license for both code and model weights.
Use Cases
How to Use Chatterbox
-
1
Create a free account
Sign up for TextToSpeechAI to receive free starter credits.
-
2
Upload a reference clip
Open the voice cloning page and upload 6 to 60 seconds of clear speech from the voice you have the right to clone.
-
3
Enter your text and language
Type the text to speak and pick one of the 15 supported output languages.
-
4
Generate the speech
TextToSpeechAI renders the text in the cloned voice with Chatterbox on our GPU servers.
-
5
Download or use the API
Download the finished audio, or automate generation through the TextToSpeechAI REST API.
Chatterbox API
Generate speech programmatically using the TextToSpeechAI REST API.
curl -X POST "https://api.texttospeechai.com/v1/generate/" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"text": "This is my cloned voice, created from a short recording.",
"voice": "en_US-kristin-medium"
}'
Frequently Asked Questions
Technical Specs
- Generation Speed Moderate
- Output Quality Very Good
- Voice Cloning Supported
- Languages 15
- GPU VRAM 4-8GB
- Credits/1000 chars 5