Black Forest LabsBlack Forest Labs·🎬 Video Generation·New

Flux 3

anonymized
Try on Venice.ai ↗
Quick reference
Flux 3 — TLDR
  • 🆕 Black Forest Labs' first multimodal Flux, here in image-to-video form
  • 🎬 Animates a still input image into video with synchronized native audio
  • 🧠 One multimodal model generates images, video and audio together
  • 🔧 Sibling modes cover text-to-video, video-to-video, and keyframe transitions
  • 💬 Multilingual dialogue plus sounds tied to on-screen physical events
  • 📚 Clips can be agentically chained into multi-shot, minutes-long sequences
  • 🎯 BFL reports better complex-prompt handling and text than earlier Flux
💰 Pricing
$0.940 – $6.38
per generation
📅 On Venice since
Aug 4, 2026
0 days ago
Provider

Black Forest Labs is a generative AI company based in Freiburg im Breisgau, Germany, founded by former members of Stability AI. The lab is best known for developing the Flux family of text-to-image models, which generate images from natural language prompts…

Read full profile →
6 models on Venice
3 video · 2 image · 1 inpaint
Since Nov 25, 2025

About this model

Flux 3 (image-to-video) is the animation entry point into Black Forest Labs' Flux 3 line: you supply a still image and a prompt, and the model produces a video clip with natively generated, synchronized audio. It sits alongside Flux 3 for pure text prompting and Flux 3 First Last Frame for controlled transitions between defined keyframes — all served from the same underlying model rather than separate specialised networks.

Unlike the still-image Flux 2 generation — Flux 2 Pro, Flux 2 Max and the editing-focused Flux 2 Max Edit — Flux 3 is described by BFL as a single multimodal model that mixes modalities and can generate images and video-with-audio jointly, from text alone or from image and video references.

BFL states that in preliminary evaluations conducted during midtraining, Flux 3 already showed significant improvement over earlier Flux versions in handling complex prompts and generating text. For video specifically, the family supports video-to-video reference transfer, generative video-audio continuation, keyframe control, high style diversity from camcorder-style footage to animation, and agentic chaining of clips into sequences lasting several minutes, with visual references used to keep characters consistent across scenes.

Audio is generated as part of the same pass rather than dubbed afterwards, covering multilingual dialogue and sound effects tied to physical events visible in the frame — a capability with no counterpart anywhere in the image-only Flux 2 family.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 3h ago