Fish Audio: ten seconds of audio to clone a voice that speaks 83 languages
Updated on September 6, 2026
Fish Audio is an AI text-to-speech and voice cloning platform that grew out of the open-source Fish Speech project. A 10-second sample is enough to rebuild a voice, and the resulting clone can speak 83 languages with automatic language detection. The free tier covers roughly seven minutes of audio per month, while commercial publishing sits behind the paid plans. YouTubers, podcasters and game studios treat it as a budget-friendly alternative to ElevenLabs.
- Voice cloning from a 10-second sample
- 83 languages handled by one model
- Inline tags for whispering, laughing, shouting
- S2 model released with open weights
- Pricing far below direct rivals
- Free tier limited to personal projects
- Interface available in English only
- Community voice library varies in quality
Fish Audio clones a voice, then makes it whisper or burst out laughing
Fish Audio runs on the in-house S2.1 Pro model, which generates natural speech across 83 languages with a time-to-first-audio around 90 milliseconds, quick enough for a live voice agent to hold a real conversation. You paste a script, pick from a library of over 2 million voices or build your own from a short recording, and the file lands a few seconds later.
The fun part is the bracket syntax. Drop an instruction like whispers nervously or laughs softly straight into the text and the model follows it word by word (yes, the laughing tag genuinely works, and it is unsettling the first time). With streaming latency under 300 ms, the API ranks among the fastest voice generation tools currently available.
From a bedroom gaming GPU to a $52 million seed round
Fish Speech, the open-source ancestor of the platform, has passed 31,000 stars on GitHub. Its author, Shijia Liao, a former NVIDIA researcher and longtime VTuber fan, trained his first models on a single gaming graphics card at home. The company now reports 8 million accounts and has announced a $52 million seed round led by Coreline Ventures and Capital Today.
The open roots are still visible. The S2 model, trained on 10 million hours of audio, shipped with open weights and fine-tuning code, and it won blind listening tests against Seed-TTS and MiniMax-Speech at release. Few companies in the speech field play that transparency card.
Fish Audio in practice, from free minutes to the pay-as-you-go API
Fish Audio can be tested without a credit card. The free plan grants 8,000 credits per month, about seven minutes of audio, but stays restricted to personal projects. Monetizing a YouTube video or podcast requires a paid subscription with commercial rights. Developers also get a free variant of S2.1 Pro through the API under fair-use conditions.
| Plan | Price | What you get |
|---|---|---|
| Free | $0 | About 7 min per month, personal use |
| Plus | $15 per month | About 200 min, commercial rights, API |
| Pro | $100 per month | 2 million credits, 3 team seats |
| API | About $15 per million characters | Pay-as-you-go, real-time streaming |
Frequently asked questions
Is Fish Audio free?
Yes, Fish Audio has a permanent free plan with 8,000 monthly credits, roughly seven minutes of audio, capped at 500 characters per generation. It is restricted to personal projects and excludes API access. Publishing monetized content on YouTube, podcasts or client work requires one of the paid plans, which include full commercial rights.
Can I use Fish Audio voices commercially?
Commercial rights come with the paid plans, starting at the Plus tier. The free plan is explicitly limited to personal, non-monetized projects under the terms of service. Cloned voices also carry a consent requirement, so only use recordings of your own voice or audio you are licensed to reproduce before publishing anything.
Fish Audio or ElevenLabs, which one should you pick?
Fish Audio bets on price and fine-grained emotion control, with an API billed around $15 per million characters, several times below its rival. ElevenLabs keeps a fuller studio suite, including video dubbing and sound effects. For high-volume multilingual voiceover on a budget, Fish Audio holds a clear cost advantage.
Can you self-host Fish Audio?
Yes, the S2 model was released with open weights, fine-tuning code and a production-ready inference engine. You can run it on your own hardware, including for commercial work under the repository license, something closed voice platforms do not allow. The hosted S2.1 Pro model, however, remains available only through the official API.
Verdict: Voicing an audiobook, a game character or a YouTube channel without booking a studio is where Fish Audio shines. Video makers, indie developers and podcasters on tight budgets get a convincing multilingual voice for a fraction of what the big names charge.
