Skip to content
← Tags
Speech & AudioMediaTopA

Kokoro-82M: an Apache-2.0 TTS trained for a thousand dollars

An open weight TTS model from hexgrad at 82M parameters under Apache-2.0, repository hexgrad/kokoro (9,061 stars, verified through the GitHub API on 2026-09-29). We list it as the cost floor and the licence floor of the audio domain: what deserves to be remembered is not a leaderboard position but the fact that it published the full arithmetic of what one TTS deployment costs. Total training cost was 1000 dollars, being 500 A100 80GB GPU hours for each of v0.19 and v1.0; v0.19 (2024-12-25) trained on under 100 hours of audio with one language and ten voices, while v1.0 (2025-01-27) used a few hundred hours across eight languages and 54 voices. On price, two independent third party sources point at the same order of magnitude: ArtificialAnalysis records 65 cents per million characters on Replicate and DeepInfra lists 80 cents per million characters, which works out to roughly three to five cents per finished hour of audio. That price line is the reference floor for every comparison we make in the audio domain: whatever costs more than it is buying voice cloning, dialect coverage, emotion control, streaming latency or compliance backing, not the ability to be listened to at all. The technical form is a decoder-only StyleTTS 2 plus ISTFTNet stack with no diffusion and an unpublished encoder, G2P through the author maintained misaki library, espeak-ng on the system side and 24kHz output; input text accepts inline phoneme overrides such as [Kokoro](/kˈOkəɹO/), the same idea as the pinyin and CMU phoneme level pronunciation correction in CosyVoice 3, turning misreading from a probability problem into declarative configuration, except the scope here is a word rather than a whole prosody. The decoder-only shape also decides what it cannot do: there is no zero shot voice cloning, the 54 voices are trained in and a reference clip will not grow a new one, which is a division of labour rather than a defect. The compliance disclosure is close to unique in TTS: the model card states that only permissively licensed or copyright free audio was used, lists public domain audio, Apache and MIT licensed audio and synthetic audio generated by closed source TTS models as separate classes, explicitly excludes synthetic audio from open TTS models and custom voice cloning, publishes a CC BY attribution table, and publishes the model SHA256 so you can verify the weights you downloaded are the same ones. Boundaries: no voice cloning, so reference driven use cases are out; 24kHz is enough for broadcast and reading but not for music grade masters; the eight languages and 54 voices have not grown since v1.0 and Mandarin usability needs your own audition; the project is close to dormant with the last push on 2025-08-06, so bugs are yours to fix; and its fame has grown a phishing surface, with the model card naming lookalike domains containing kokoro as unrelated to the project, the only real entries being hexgrad/Kokoro-82M on Hugging Face and hexgrad/kokoro on GitHub, so never pay a search result that calls itself the Kokoro official site. Graded B (confirmed): what is confirmed is the cost and market price chain with independent third party sources, not audio quality.

<$1API market price (per million characters, 2025-04 model card)Confirmed · 2025-04
ProducthexgradSiteRepo
Kokoro-82M: an Apache-2.0 TTS trained for a thousand dollars
Speech & AudioMediaTopA

VibeVoice: a 7.5Hz token rate buys 90 minutes of speech in one pass

The open frontier speech model family from Microsoft, repository microsoft/VibeVoice (MIT, 54,531 stars, verified through the GitHub API). We catalogue it as one asset rather than three because TTS, Realtime and ASR share a single technical core: both the acoustic and the semantic tokenizer are continuous rather than quantised into discrete codebooks, and the frame rate is pushed down to 7.5Hz. That 7.5Hz is where every capability comes from. Ninety minutes of audio at 25Hz is 135,000 tokens and does not fit a 64K context, while at 7.5Hz it is 40,500, so single pass synthesis of 90 minutes and single pass transcription of 60 minutes are two directions of the same fact rather than two separate engineering feats. Five product lines each carry their own ceiling: TTS-1.5B at 90 minutes per run with up to four speakers, Realtime-0.5B at roughly 300ms to first packet, ASR-7B transcribing 60 minutes in one pass while emitting who said what and when with custom hotwords across more than 50 languages, ASR-Streaming emitting text as speech arrives, and ASR-BitNet using heterogeneous quantisation to compress 4.62GB into 1.58GB with RTF under 1 on three or more CPU threads and no GPU. The ASR line is the one worth remembering: conventional long form transcription chains slicing, recognition, diarisation and timestamp alignment, and the global context is lost at the slice boundary, while VibeVoice-ASR fuses recognition, diarisation and timestamps into a single generation, which is the difference between usable and unusable for meeting minutes and support QA. The architecture is next token diffusion, with the LLM owning conversational direction and the diffusion head owning acoustic detail, because turn consistency over a long dialogue is a language problem and only the LLM can hold four speakers distinct across 90 minutes. One fact must be recorded plainly: on 2025-09-05 the team removed the VibeVoice-TTS code from the repository, stating that usage inconsistent with its declared intent had been observed; the weights remain on Hugging Face but the Quick Try section reads Disabled, the official inference scripts are gone and today the TTS path runs through community reproductions, with the model page advising against commercial or real world use without further testing. Every public move since has been on the ASR side, so the commercially usable half is ASR rather than TTS. Its position in tables published by others must also be reported as written: the CosyVoice 3 README lists it at test-zh CER 1.16 with similarity 74.4 and test-en WER 3.04 with similarity 68.9, against CosyVoice3 base at 1.21 and 78.0. Short clip zero shot cloning is not its lane; the three comparisons that matter are how long a single run can be, whether four speakers stay distinct, and whether one hour of meeting audio can be processed in one pass, and on those three the open source field has almost no rival. Boundaries: no official TTS inference path; MIT covers code but not the usage limits on the model page; long form failure is global, since a 90 minute single pass has no retry one segment at a time fallback; the official Risks and Limitations section names deepfakes and requires disclosure of AI involvement. Graded B (confirmed): what is confirmed is the frame rate arithmetic, the published ceilings of each line, the code removal and the licence boundary, not audio quality, which we only relay without recomputing or blind listening.

90 / 60 minSingle long-audio cap (TTS 4 speakers / ASR single pass)Confirmed · 2026-09
ResearchMicrosoftSiteRepo
VibeVoice: a 7.5Hz token rate buys 90 minutes of speech in one pass
Speech & AudioMediaTopC

Fun-CosyVoice3-0.5B: open TTS that turns controllability into an interface

The third generation LLM based speech synthesis system from the FunAudioLLM group at Alibaba, 0.5B parameters, released as Fun-CosyVoice3-0.5B-2512 with base and RL weight sets plus training and inference scripts; the code repository has moved to QwenAudio/CosyVoice (Apache-2.0, 23,794 stars, verified through the GitHub API on 2026-09-29). We list it as the open source top row for speech in the audio domain, and the reason is not that it sounds most human but that it gives each of the four failure modes of production TTS an explicit interface: misread polyphonic characters go through pronunciation correction with Chinese pinyin or English CMU phonemes written straight into the input, wrong readings of numbers and symbols go through the built in text normalisation, collapsed long sentences go through RAS repetition aware sampling, and low latency goes through bidirectional streaming with an official first packet at 150ms, which is the right order of magnitude to sit inside a realtime conversational agent. Three readings matter on the metric sheet. The test-zh speaker similarity of 78.0 is the highest among 0.5B open models and above the human baseline of 75.5, yet still below the closed source Seed-TTS at 79.6, so open source first holds while world first does not. The test-hard column is the real strength, base 6.71 and RL 5.44 being the lowest in the table and better than the closed source 7.59, and hard is exactly long sentences and tongue twisters, so that lead maps to usable real scripts. English needs a discount: base test-en similarity 71.8 sits below the human baseline 73.4 and only the RL tier WER of 1.68 recovers it. Base and RL are two post training weights of one architecture, with all three error rates falling (zh CER down 33 percent, hard CER down 19 percent) and all three similarities slipping slightly, which is a good trade for broadcast and support workloads where one misread character is an incident, while voice fidelity work for a specific IP should audition the RL tier first. The repository also ships GRPO training scripts and a triton plus TensorRT-LLM runtime claimed at four times the speed of HF transformers, so post training is a path others can continue rather than a one off delivery, and the language surface covers nine languages with more than eighteen Chinese dialect accents, a tier the closed APIs still barely offer. Boundaries: the licence is two layered, Apache-2.0 for code while the weight terms live on the model pages and must be checked separately before commercial use; every reading comes from the evaluation of the authors themselves with no third party blind listening table to cross check and we did not recompute; zero shot cloning similarity comes from standard sets while real deployments depend on the noise and style of the reference recording. Graded C (vendor claim).

78.0test-zh speaker similarity (0.5B open-weight tier, human baseline 75.5)Vendor Claim · 2025-12
ProductAlibaba Tongyi FunAudioLLM (QwenAudio)SiteRepo
Fun-CosyVoice3-0.5B: open TTS that turns controllability into an interface
Speech & AudioMediaTopC

Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data

The new generation speech synthesis model from ElevenLabs, shipping in two variants: Eleven v4 for highest quality and Eleven v4 Turbo for real time use at a median inference latency of about 100 ms, with an official footnote stating that application and network latency are excluded. The vendor positions it above v3 on output quality, voice accuracy, consistency, emotion, delivery, audio tags and language coverage, and recommends migrating. We list it as the top closed source voice row in the audio domain because this generation changed the optimisation target from sounding better to sounding more like the source: timbre, cadence, delivery and other mannerisms are reproduced together, and the vendor concedes that accuracy and personal preference are not always the same thing and that you may still prefer how a voice sounded in v3. The engineering consequence is hard, since the burden of data cleaning moves back to the user. The showcase samples are explicitly labelled raw, with no EQ, compression, normalisation, de-essing or plosive removal, and the vendor says that if you hear clipping or plosives they are very likely in the training data and the model simply captured them accurately. What is deliverable comes down to three numbers: about 100 ms on Turbo; 10,000 characters per request, roughly ten minutes of audio, against 5,000 on v3 so long form text carries twice as many segment seams; and a coverage claim of 90 plus languages against a FAQ that enumerates 87, a gap we keep visible rather than smoothing away. Four boundaries matter. Cross language accent handling is a default behaviour change and not a toggle, so generating in a language that differs from the reference produces fluent native sounding speech in the target language, and the vendor calls a switchable version a research project with no timeline; any persona that depends on a native accent carried into a second language must be tested first. The control surface narrowed to Stability and Similarity, the Style and Speed sliders are gone, SSML is unsupported, and fine grained delivery moves to audio tags which the vendor admits are not perfect yet. Continuous model evolution is stated in writing, with training continuing after launch and behaviour possibly shifting over time, so teams treating a voice as a brand asset need periodic re-testing. Voice Design voices may also be less performative on v4. It is closed and not self-hostable, data must pass through ElevenLabs, and cloning compliance sits with the user; the self-hostable counterparts are CosyVoice 3, VibeVoice and Kokoro for speech and YuE2 for music under a non-commercial weight licence. This entry and the Eleven v3 already catalogued here are two generations of the same commercial stack rather than a replacement. Graded C (vendor claim): no verifiable third party benchmark, and no blind listening test by us.

~100msv4 Turbo median inference latency (model side, excl. app and network)Vendor Claim · 2026-09
ProductionElevenLabsSite
Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data
Speech & AudioMediaTopC

Eleven v3: the commercial audio stack that covers speaking, listening and singing

ElevenLabs' flagship speech synthesis model, officially its most emotive and expressive: 70+ languages, a 5,000-character request cap, dramatic delivery and natural multi-speaker dialogue, with a separate Text to Dialogue API making multi-character dialogue an endpoint, not a prompt trick. We list it as a top row in audio for the stack behind it: the only commercial one covering speaking, listening and singing, with four TTS tiers, three ASR tiers including a Medical build and two music models. The TTS tiers are mutually exclusive: v3 is most expressive but non-realtime at 5,000 characters, Flash v2.5 is lowest latency (about 75ms, 50% cheaper per character) and Multilingual v2 is most stable for long text at 29 languages. On ASR, Scribe v2 covers 90+ languages with diarisation for 32 speakers. Latency figures are model-side only, excluding network and upstream LLM time, not SLAs. Limits: you split long text and own continuity across the seams; closed and not self-hostable, a hard barrier where data sovereignty matters. No speech board among the ones we sync, so listed alongside Gemini 3.8 Flash TTS without ranking. Graded C (vendor-stated): no WER recomputation or blind listening test by us.

70+Language coverageVendor Claim · 2026-09
ProductElevenLabsSite
Eleven v3: the commercial audio stack that covers speaking, listening and singing
Speech & AudioMediaTopC

Gemini 3.8 Flash TTS: turning director-level voice direction into API parameters

Google's speech generation model (gemini-3.8-flash-tts), positioned for studio-grade voice fidelity, expressive acting and long-form stability, the most overlooked being the last: audiobook failure is not an ugly sentence but timbre drift by hour three. The design separates scopes: text is strictly a script while performance direction travels on another channel and is never read aloud, fixing "(sighs)" that used to be spoken; round style goes into speech_metadata.style, word-level emotion into pipe tags. Four voice sources: prebuilt, an extended library, voice design IDs (describe it, then reuse) and replication IDs; design and replication arrived only in 3.8. Specs: 130 languages with automatic language detection, WAV by default at 24 kHz, at most 2 speakers per request and only with prebuilt voices. Billing is per audio token at 25 tokens per second, $0.50 input and $9.00 output per million through 2026-12-31 then $1.00/$18.00. Limits: 24 kHz is not mastering grade; replication compliance rests with the user. No speech board among the 12 we sync, so listed alongside ElevenLabs v3 without ranking. Graded C (vendor-stated): no intelligibility or blind listening test by us.

All supportedVoice design + cloning + multi-speakerVendor Claim · 2026-09
ProductGoogleSite
Gemini 3.8 Flash TTS: turning director-level voice direction into API parameters