Skip to content
← Tags

#ElevenLabs (3)

Speech & AudioMediaTopC

Eleven Music v2.5: the closed-source row that turns music generation into a callable pipeline

The music generation model from ElevenLabs, currently at v2.5, released 2026-09-11 with the release post last updated 2026-09-20. We list it as the engineering-side top row for music in the audio domain, which contrasts with rather than duplicates the Suno entry already catalogued here: the value of Suno concentrates inside the product, a Studio multi-track timeline, Custom Models, up to twelve stems and MIDI export, while Eleven Music splits comparable capability into endpoints, namely compose, stream, a structured composition plan, scoring an uploaded video, uploading existing audio, stem separation, finetunes and section level inpainting, plus a marketplace where creators license tracks to each other. Choosing between them therefore needs no audio quality comparison, only one question: is this track something a person sits and adjusts, or something a pipeline requests in volume. The API surface is nine endpoints rather than one generate call, and two details matter for procurement: both compose and stem-separation take a sign_with_c2pa flag that applies to mp3 output so outgoing files can carry content credentials, and output format is tied to subscription tier, with mp3_44100_192 requiring Creator or above and pcm_44100 requiring Pro or above on stem-separation and video-to-music while compose goes up to mp3_48000_320, so the threshold for lossless differs per endpoint. The trap worth memorising is that v2.5 is already the interface default yet the model_id enum of music_v1, music_v2 and music_v2_5 still defaults to music_v1, and output_format=auto resolves per model to mp3_44100_128 on v1 and mp3_48000_192 on v2, so a minimal call that omits model_id silently gets the oldest generation at a lower bitrate; pin both in production code and name mp3_48000_320 when 320kbps is required. The two plan schemas are not interchangeable and the wrong pairing is a hard error: music_v1 takes MusicPrompt while music_v2 and music_v2_5 take CompositionPlan, whose chunks carry a text field with square bracket section names, lyric lines and curly brace inline directions. On the control surface prompt and composition_plan are mutually exclusive, and the request takes music_length_ms from 3000 to 600000, that is three seconds to ten minutes, plus force_instrumental, finetune_id and seed. Rights and commercial use are the most structured part of this line: a multi-year agreement with Universal Music Group was announced alongside v2.5 and the vendor states it is separate from Music 2.5; every track is yours on every plan including Free, Free allows commercial use provided ElevenMusic is credited, lossless downloads are capped at five per day on Free and 400 per month on Pro, tracks built on another artist song through Audio Reference cannot be downloaded, the Marketplace sells licences by usage type with creator earnings starting at 25 percent, and no licence permits distribution to streaming platforms such as Spotify. The official evidence for v2.5 over v2 is a self-run blind test in which v2.5 won the majority of 47,885 paired takes with the widest gap in vocal-led and acoustic-heavy genres. Boundaries: the API is paid-subscription only; vocals are documented for English, Spanish, German and Japanese with no Mandarin, so a Chinese language song is more practical on YuE2 or Suno; seed does not guarantee reproducibility; and it is closed source with no weights, so the self-hostable alternatives here are YuE2 for music and VoiceStudio or Kokoro-82M for speech. Graded C (vendor claim): quality was not recomputed and we ran no blind listening test, but the API contract is verifiable documentation and every clause of it is listed in the body.

10 minAPI 单曲时长上限(music_length_ms 3000-600000)Vendor Claim · 2026-09
ProductionElevenLabsSite
Eleven Music v2.5: the closed-source row that turns music generation into a callable pipeline
Speech & AudioMediaTopC

Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data

The new generation speech synthesis model from ElevenLabs, shipping in two variants: Eleven v4 for highest quality and Eleven v4 Turbo for real time use at a median inference latency of about 100 ms, with an official footnote stating that application and network latency are excluded. The vendor positions it above v3 on output quality, voice accuracy, consistency, emotion, delivery, audio tags and language coverage, and recommends migrating. We list it as the top closed source voice row in the audio domain because this generation changed the optimisation target from sounding better to sounding more like the source: timbre, cadence, delivery and other mannerisms are reproduced together, and the vendor concedes that accuracy and personal preference are not always the same thing and that you may still prefer how a voice sounded in v3. The engineering consequence is hard, since the burden of data cleaning moves back to the user. The showcase samples are explicitly labelled raw, with no EQ, compression, normalisation, de-essing or plosive removal, and the vendor says that if you hear clipping or plosives they are very likely in the training data and the model simply captured them accurately. What is deliverable comes down to three numbers: about 100 ms on Turbo; 10,000 characters per request, roughly ten minutes of audio, against 5,000 on v3 so long form text carries twice as many segment seams; and a coverage claim of 90 plus languages against a FAQ that enumerates 87, a gap we keep visible rather than smoothing away. Four boundaries matter. Cross language accent handling is a default behaviour change and not a toggle, so generating in a language that differs from the reference produces fluent native sounding speech in the target language, and the vendor calls a switchable version a research project with no timeline; any persona that depends on a native accent carried into a second language must be tested first. The control surface narrowed to Stability and Similarity, the Style and Speed sliders are gone, SSML is unsupported, and fine grained delivery moves to audio tags which the vendor admits are not perfect yet. Continuous model evolution is stated in writing, with training continuing after launch and behaviour possibly shifting over time, so teams treating a voice as a brand asset need periodic re-testing. Voice Design voices may also be less performative on v4. It is closed and not self-hostable, data must pass through ElevenLabs, and cloning compliance sits with the user; the self-hostable counterparts are CosyVoice 3, VibeVoice and Kokoro for speech and YuE2 for music under a non-commercial weight licence. This entry and the Eleven v3 already catalogued here are two generations of the same commercial stack rather than a replacement. Graded C (vendor claim): no verifiable third party benchmark, and no blind listening test by us.

~100msv4 Turbo 中位推理延迟(模型侧,不含应用与网络)Vendor Claim · 2026-09
ProductionElevenLabsSite
Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data
Speech & AudioMediaTopC

Eleven v3: the commercial audio stack that covers speaking, listening and singing

ElevenLabs' flagship speech synthesis model, officially its most emotive and expressive: 70+ languages, a 5,000-character request cap, dramatic delivery and natural multi-speaker dialogue, with a separate Text to Dialogue API making multi-character dialogue an endpoint, not a prompt trick. We list it as a top row in audio for the stack behind it: the only commercial one covering speaking, listening and singing, with four TTS tiers, three ASR tiers including a Medical build and two music models. The TTS tiers are mutually exclusive: v3 is most expressive but non-realtime at 5,000 characters, Flash v2.5 is lowest latency (about 75ms, 50% cheaper per character) and Multilingual v2 is most stable for long text at 29 languages. On ASR, Scribe v2 covers 90+ languages with diarisation for 32 speakers. Latency figures are model-side only, excluding network and upstream LLM time, not SLAs. Limits: you split long text and own continuity across the seams; closed and not self-hostable, a hard barrier where data sovereignty matters. No speech board among the ones we sync, so listed alongside Gemini 3.8 Flash TTS without ranking. Graded C (vendor-stated): no WER recomputation or blind listening test by us.

70+语言覆盖Vendor Claim · 2026-09
ProductElevenLabsSite
Eleven v3: the commercial audio stack that covers speaking, listening and singing