Skip to content
← Tags

#Music Generation (4)

Speech & AudioMediaTopC

Eleven Music v2.5: the closed-source row that turns music generation into a callable pipeline

The music generation model from ElevenLabs, currently at v2.5, released 2026-09-11 with the release post last updated 2026-09-20. We list it as the engineering-side top row for music in the audio domain, which contrasts with rather than duplicates the Suno entry already catalogued here: the value of Suno concentrates inside the product, a Studio multi-track timeline, Custom Models, up to twelve stems and MIDI export, while Eleven Music splits comparable capability into endpoints, namely compose, stream, a structured composition plan, scoring an uploaded video, uploading existing audio, stem separation, finetunes and section level inpainting, plus a marketplace where creators license tracks to each other. Choosing between them therefore needs no audio quality comparison, only one question: is this track something a person sits and adjusts, or something a pipeline requests in volume. The API surface is nine endpoints rather than one generate call, and two details matter for procurement: both compose and stem-separation take a sign_with_c2pa flag that applies to mp3 output so outgoing files can carry content credentials, and output format is tied to subscription tier, with mp3_44100_192 requiring Creator or above and pcm_44100 requiring Pro or above on stem-separation and video-to-music while compose goes up to mp3_48000_320, so the threshold for lossless differs per endpoint. The trap worth memorising is that v2.5 is already the interface default yet the model_id enum of music_v1, music_v2 and music_v2_5 still defaults to music_v1, and output_format=auto resolves per model to mp3_44100_128 on v1 and mp3_48000_192 on v2, so a minimal call that omits model_id silently gets the oldest generation at a lower bitrate; pin both in production code and name mp3_48000_320 when 320kbps is required. The two plan schemas are not interchangeable and the wrong pairing is a hard error: music_v1 takes MusicPrompt while music_v2 and music_v2_5 take CompositionPlan, whose chunks carry a text field with square bracket section names, lyric lines and curly brace inline directions. On the control surface prompt and composition_plan are mutually exclusive, and the request takes music_length_ms from 3000 to 600000, that is three seconds to ten minutes, plus force_instrumental, finetune_id and seed. Rights and commercial use are the most structured part of this line: a multi-year agreement with Universal Music Group was announced alongside v2.5 and the vendor states it is separate from Music 2.5; every track is yours on every plan including Free, Free allows commercial use provided ElevenMusic is credited, lossless downloads are capped at five per day on Free and 400 per month on Pro, tracks built on another artist song through Audio Reference cannot be downloaded, the Marketplace sells licences by usage type with creator earnings starting at 25 percent, and no licence permits distribution to streaming platforms such as Spotify. The official evidence for v2.5 over v2 is a self-run blind test in which v2.5 won the majority of 47,885 paired takes with the widest gap in vocal-led and acoustic-heavy genres. Boundaries: the API is paid-subscription only; vocals are documented for English, Spanish, German and Japanese with no Mandarin, so a Chinese language song is more practical on YuE2 or Suno; seed does not guarantee reproducibility; and it is closed source with no weights, so the self-hostable alternatives here are YuE2 for music and VoiceStudio or Kokoro-82M for speech. Graded C (vendor claim): quality was not recomputed and we ran no blind listening test, but the API contract is verifiable documentation and every clause of it is listed in the body.

10 minAPI 单曲时长上限(music_length_ms 3000-600000)Vendor Claim · 2026-09
ProductionElevenLabsSite
Eleven Music v2.5: the closed-source row that turns music generation into a callable pipeline
Speech & AudioMediaTopC

Suno: the platform that pulls the second half of music production inside one product

The most complete commercial platform on the text-to-song route: a style description, your own lyrics, a hummed melody or a recorded riff all work as a starting point, a full song with vocals and instrumentation arrives in seconds, and the same product then extends it, edits sections, restyles it, extracts stems and remasters it. We list it as the top closed entry for music in the audio domain on workflow completeness rather than peak audio quality. Most generative music products cover drafting and picking a take and stop there; Suno connects the rest through stem extraction with up to 12 stems, MIDI export and Suno Studio, and Studio speaks conventional DAW semantics rather than adding another prompt box, with a multitrack timeline, take lanes, comping, manual BPM to settle tempo drift and per clip transpose and speed. Custom Models trains up to three private style variants from six or more tracks you own, and Voices, formerly Personas, generates in your own singing timbre with a verification step. The tier split matters: the free plan covers creation (generation, lyrics, Cover, crop and fade, audio upload) while stems, Add Vocals, Voices and Custom Models require Pro or Premier and Studio is Premier only on desktop web, so real cost modelling should assume Premier. Version numbers do not track quality on third party benchmarks either: on WildSongBench v5 scores 6.8721, above v6 at 6.5562 and v6 Wild at 6.4195, with v4.5 at 6.6995 and v5.5 at 6.7150, and the vendor itself asks for capability based rather than version based description. Limits: closed with no self-hosting, so unreleased melodies and lyrics must be uploaded; no stable version semantics, meaning a regeneration can sound different after a model update; commercial rights follow the tier; control granularity sits at section and style level with no editable chord track, so theory level edits require Studio multitrack re-arrangement or an open model such as YuE2 that exposes an ABC score. Graded C (vendor claim); the benchmark numbers are submitted by m-a-p and have not been recomputed by us.

6.8721WildSongBench SongBench 均分(v5,第三方测)Vendor Claim · 2026-09
ProductionSunoSite
Suno: the platform that pulls the second half of music production inside one product
Speech & AudioMediaTopC

YuE2-3B: open song generation that exposes the score as an interface

The open song generation model from m-a-p, 3B parameters, weights under CC-BY-NC-4.0, turning lyrics and a style prompt into a complete song with vocals and accompaniment at 48 kHz stereo. We list it as the open-weight top row for music in the audio domain on the strength of a combination that is close to unique among its peers: open weights plus an editable intermediate representation. One AR-NAR Mixture-of-Transformers backbone writes an ABC score (melody and chords) and semantic tokens, flow matching then produces acoustic latents and a VAE decodes them, so the score is an artefact that a person or an agent can read and edit instead of a black box whose only control is another sample. Three cot modes (full, melody, off) map onto composing, covering and direct generation, and the official agentic editing demo runs nine turns across fourteen versions from Mandarin pop to English jazz. Read the benchmark protocol carefully: on 192 WildSongBench prompts the best-of-8 SongBench average of 6.9632 sits above Suno v5 at 6.8721 and Suno v6 at 6.5562, but that is eight candidates with selection against a delivered single candidate; MuLan and AllMusicCaps, the two style-text alignment measures, still favour Suno v5, and PER at 8.44 percent trails Suno v6 Wild at 7.45 and MiniMax Music 3 at 6.27. Cost is the most concrete advantage: a 3.6 minute song in 71 seconds on an RTX 4090 24GB with a peak of 11.18 GiB, and 373 songs per hour on an H800 with vLLM at AR concurrency 32. Limits: non-commercial licence; identity preservation in covers comes almost entirely from a supplied score (CLEWS mAP collapses to 0.006 without one); the benchmark runs on the legacy VAE while the default release is the newer one; no technical report yet and no arena result. Graded C (vendor claim), not recomputed by us.

6.9632WildSongBench SongBench 均分(best-of-8)Vendor Claim · 2026-09
ResearchMultimodal Art Projects (m-a-p)SiteRepo
YuE2-3B: open song generation that exposes the score as an interface
Speech & AudioMediaTopC

Eleven v3: the commercial audio stack that covers speaking, listening and singing

ElevenLabs' flagship speech synthesis model, officially its most emotive and expressive: 70+ languages, a 5,000-character request cap, dramatic delivery and natural multi-speaker dialogue, with a separate Text to Dialogue API making multi-character dialogue an endpoint, not a prompt trick. We list it as a top row in audio for the stack behind it: the only commercial one covering speaking, listening and singing, with four TTS tiers, three ASR tiers including a Medical build and two music models. The TTS tiers are mutually exclusive: v3 is most expressive but non-realtime at 5,000 characters, Flash v2.5 is lowest latency (about 75ms, 50% cheaper per character) and Multilingual v2 is most stable for long text at 29 languages. On ASR, Scribe v2 covers 90+ languages with diarisation for 32 speakers. Latency figures are model-side only, excluding network and upstream LLM time, not SLAs. Limits: you split long text and own continuity across the seams; closed and not self-hostable, a hard barrier where data sovereignty matters. No speech board among the ones we sync, so listed alongside Gemini 3.8 Flash TTS without ranking. Graded C (vendor-stated): no WER recomputation or blind listening test by us.

70+语言覆盖Vendor Claim · 2026-09
ProductElevenLabsSite
Eleven v3: the commercial audio stack that covers speaking, listening and singing