small 3,
pixtral 12b
Unlimited FREE
HAPPY TIME: Mistral: small 3, pixtral 12b
Unlimited FREE
Multilingual TTS // Cabina.AI

CosyVoice - Multilingual TTS and Voice Cloning on Cabina.AI

The CosyVoice multilingual TTS model, ready to run without a single GPU to manage.
CosyVoice is the multilingual text-to-speech and voice cloning system built by the FunAudioLLM team at Alibaba Tongyi SpeechTeam, released under an open Apache-2.0 license. As a cosyvoice ai model, it clones a voice from a short reference clip and speaks it back in nine languages and more than 18 Chinese dialects, streaming the first audio chunk in about 150ms.
On Cabina.AI, you get that same engine without cloning a repository, provisioning a GPU, or configuring CUDA. CosyVoice sits inside one subscription alongside every other model on the platform, so the same account that runs your voice cloning also runs your language models, image generation, and everything else you need.
Start generating with CosyVoice
СosyVoice voice cloning
Drop your voice clip here
A short, clear recording is all CosyVoice needs to clone it
Choose File
Voice Cloning

Clone a voice, zero training required

Upload a short reference clip and CosyVoice reproduces that voice speaking your text. The same cloned voice can then speak in nine languages and more than 18 Chinese dialects, streaming the first audio chunk in about 150ms.
Start cloning on Cabina.AI

Capabilities

What the CosyVoice TTS model can do

At its core, CosyVoice is a multilingual TTS model that reproduces a speakers voice from a short reference sample and then speaks new text in that voice, across languages the original speaker may never have used. The cosyvoice voice cloning pipeline works zero-shot: give it one clip, and it can render that same voice in 9 languages and more than 18 Chinese dialects, without any per-speaker fine-tuning.

Real-time streaming

For real-time voice agents and conversational AI, CosyVoice streams the first audio chunk in about 150ms, fast enough that a spoken response starts before the sentence has finished generating on the backend.

Phoneme-level pronunciation

Pronunciation is controllable at the phoneme level, using Pinyin or CMU notation, which matters for technical and domain-specific vocabulary where a generic TTS engine tends to guess wrong.

Instruction-based control

Instruction-based controls let you adjust emotion, speed, volume, and dialect directly through the prompt, useful for audiobook and long-form narration that needs to shift tone across a chapter rather than read everything in one flat register.

Typical uses

Taking one recorded voice and localizing it into multiple languages for the same piece of content, powering voice agents that need to respond in under a second, narrating long-form audio with emotional variation, and handling text with domain-specific pronunciation that general TTS models mishandle.

Watch It In Action

See CosyVoice clone a voice in real time

Two short demos of CosyVoice running on Cabina.AI, from a first-person voice clone to a multilingual experiment on iconic voices.
Demo 01
Clone your voice, then speak any language
One short recording is enough for CosyVoice on Cabina.AI to reproduce your voice in French, German, Chinese, and Japanese, with adjustable emotion and intonation along the way.
Demo 02
Same voice, three new languages
We ran the experiment on three iconic voices, Scrooge McDuck in Spanish, Winston Churchill in French, and a classic quote in German, and CosyVoice kept every voices character and charisma intact.
Benchmarks
Reported results for the Fun-CosyVoice3-0.5B checkpoint give a concrete picture of how this cosyvoice speech model performs across languages and difficulty levels:
1.21%
Chinese · character error rate (CER)
78% speaker similarity
2.24%
English · word error rate (WER)
71.8% speaker similarity
6.71%
Hard test cases · error rate
75.8% speaker similarity
CER and WER measure how closely the generated speech matches the target text, while speaker similarity measures how closely the cloned voice matches the reference speaker. Lower error rates and higher similarity scores indicate stronger performance.

Try other LLMs similar to CosyVoice

Cosyvoice Pricing

Included in your Cabina.AI subscription

CosyVoice pricing on Cabina.AI is token-based, and it works the same way across every model on the platform: fewer tokens spent per generation means more usage from the same plan.
Each CosyVoice (Zero Shot) generation costs 0.09 tokens. On the Growth plan, which includes 2,100 tokens for $18.99/month, that works out to an average of about 22,626 generations, the only tier where Cabina.AI displays this figure directly.
Cabina.AI full plan lineup:

Starter

Easy to start

$4.99

per month

450 tokens/mo

All LLM functionalities · Basic email support · Unlimited access

Basic

For regular personal use

$9.99

per month

1,000 tokens/mo

All LLM functionalities · Basic email support · Unlimited access

Premium

Ultimate power for professionals

$99.99

per month

14,000 tokens/mo

All LLM functionalities · Priority email support · Unlimited access

If a monthly subscription is not what you want, Pay-As-You-Go gives $20 for 1,800 tokens ($1 ≈ 90 tokens), with a $3 minimum purchase. Tokens stay active for 100 days from your last payment, and there is no subscription commitment.
Whichever plan you choose, the same token pool that runs CosyVoice also runs every other model bundled into Cabina.AI, so a single subscription covers your voice cloning alongside language models, image generation, and the rest of the catalog.
Cosyvoice Online

Getting started with CosyVoice online

Running CosyVoice online through Cabina.AI takes a few minutes, with no repository to clone and no inference server to configure.
1
Create a Cabina.AI account.
2
Choose a plan: start on the Free tier (50 tokens forever) to try things out, pick a subscription tier that fits your expected usage, or use Pay-As-You-Go if you would rather not commit to a monthly plan.
3
Select CosyVoice from the model list inside your dashboard.
4
Provide a reference voice clip and the text you want spoken, then generate.
Each generation costs 0.09 tokens, drawn from the same token pool used across every other model in your subscription, so there is nothing extra to set up between CosyVoice and the rest of the catalog.

Frequently asked questions

Can I use CosyVoice for free?

Yes, up to a point. The Free plan gives 50 tokens forever, enough to test cloning and multilingual generation before deciding on a paid tier. Because each CosyVoice (Zero Shot) generation costs 0.09 tokens, that free allocation supports a handful of generations without any payment.

Is CosyVoice open source?

Yes, CosyVoice is released under the Apache-2.0 license by the FunAudioLLM team at Alibaba Tongyi SpeechTeam. The code and model weights are public on GitHub, which is also why it can run on platforms like Cabina.AI without a separate licensing negotiation.

Does CosyVoice support commercial use?

The Apache-2.0 license permits commercial use, including hosting it as part of a paid product. That is how Cabina.AI offers it: commercial use is already covered by the license, so there is no extra commercial tier or contract required.

How does CosyVoice pricing compare to self-hosting it from GitHub?

Self-hosting means provisioning a GPU, installing CUDA, and managing concurrency scaling yourself, on top of whatever compute costs you incur. Running CosyVoice through Cabina.AI trades that setup for a per-generation token cost, with the underlying infrastructure managed for you.

Are there rate limits or latency issues with CosyVoice in streaming mode?

Self-hosted CosyVoice deployments can see latency and throughput vary under heavy concurrent load, since a single GPU has a limit to how many streaming requests it can serve smoothly at once. On Cabina.AI, that scaling is handled on the infrastructure side rather than left to you to size and monitor.

Do I pay for unanswered requests?

No, you are not charged for requests that do not receive a response.

How do I Cancel a Subscription?

You can cancel your subscription at any time. Your current subscription will remain active until the end of the paid period without any restrictions.

CosyVoice brings multilingual, zero-shot voice cloning and low-latency streaming without the GPU management that comes with self-hosting the open-source weights. It runs on the same Cabina.AI subscription as every other model on the platform, so switching between voice cloning, transcription, and everything else does not mean juggling separate accounts.
Start using CosyVoice on Cabina.AI now, or explore the Free plan first if you want to test it before committing.