ElevenLabsElevenLabs·🎵 Music Generation·↑ Newer: ElevenLabs Music v2.5·New

ElevenLabs Music v2

anonymized
Try on Venice.ai ↗
Quick reference
ElevenLabs Music v2 — TLDR
  • 🧠 Text-to-music model generating full tracks with vocals or instrumental-only
  • 🆕 Successor to Music v1, with better prompt adherence and composition
  • 🔧 Inpainting regenerates a chosen section without touching the rest
  • 📚 Composition plans build songs section by section, up to 30 chunks
  • 🌐 Vocals supported in 59 languages, native-like quality in 11
  • 📏 Configurable track length from 3 seconds to 10 minutes
  • 🎯 Audio Reference guides style from a clip of roughly 30 seconds
  • 🔒 Built with artists, labels and publishers; cleared for commercial use
💰 Pricing
$0.690 – $6.90
per track
📅 On Venice since
Sep 15, 2026
5 days ago
Provider

ElevenLabs is a software company specializing in natural-sounding speech synthesis and audio generation powered by deep learning. The company has established itself as a leading force in AI-driven voice technology, building tools that span text-to-speech,…

Read full profile →
8 models on Venice
6 music · 1 tts · 1 asr
Since Feb 22, 2026

About this model

ElevenLabs Music v2 is the text-to-music generation model from ElevenLabs, which also offers text-to-speech and speech-to-text models. It turns a text prompt describing genre, mood, instrumentation and lyrics into a finished track, with sung vocals or as an instrumental, and is available both in the company's creative apps and through its music API.

Compared with its predecessor ElevenLabs Music (Music v1), ElevenLabs describes v2 as offering improved prompt adherence, composition, prompt understanding, multilingual output and vocal delivery. It also adds capabilities v1 lacked: Audio Reference, where a short uploaded clip of up to about 30 seconds steers sound, production style, instrumentation, tempo and mood without copying the source; improved inpainting, which regenerates a selected section such as a bridge while leaving the chorus untouched; and long-form composition that builds intro, verse and chorus in sequence rather than short clips. The company also notes more natural vocal performances and denser delivery patterns, including fast rap.

For developers, composition plans expose an ordered list of up to 30 chunks, each with its own styles, lyrics and duration, and durations are configurable from 3 seconds to 10 minutes. Vocals are supported across 59 languages, with native-like quality claimed in 11.

The newer ElevenLabs Music v2.5 builds on v2 with improved audio quality and prompt adherence while keeping the same feature set, and v2 remains available. Related audio models on this catalog include ElevenLabs Sound Effects and ElevenLabs TTS v3.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 4d ago