About this model
ElevenLabs Music v2 is the text-to-music generation model from ElevenLabs, which also offers text-to-speech and speech-to-text models. It turns a text prompt describing genre, mood, instrumentation and lyrics into a finished track, with sung vocals or as an instrumental, and is available both in the company's creative apps and through its music API.
Compared with its predecessor ElevenLabs Music (Music v1), ElevenLabs describes v2 as offering improved prompt adherence, composition, prompt understanding, multilingual output and vocal delivery. It also adds capabilities v1 lacked: Audio Reference, where a short uploaded clip of up to about 30 seconds steers sound, production style, instrumentation, tempo and mood without copying the source; improved inpainting, which regenerates a selected section such as a bridge while leaving the chorus untouched; and long-form composition that builds intro, verse and chorus in sequence rather than short clips. The company also notes more natural vocal performances and denser delivery patterns, including fast rap.
For developers, composition plans expose an ordered list of up to 30 chunks, each with its own styles, lyrics and duration, and durations are configurable from 3 seconds to 10 minutes. Vocals are supported across 59 languages, with native-like quality claimed in 11.
The newer ElevenLabs Music v2.5 builds on v2 with improved audio quality and prompt adherence while keeping the same feature set, and v2 remains available. Related audio models on this catalog include ElevenLabs Sound Effects and ElevenLabs TTS v3.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 4d ago