About this model
Flux 3 (image-to-video) is the animation entry point into Black Forest Labs' Flux 3 line: you supply a still image and a prompt, and the model produces a video clip with natively generated, synchronized audio. It sits alongside Flux 3 for pure text prompting and Flux 3 First Last Frame for controlled transitions between defined keyframes — all served from the same underlying model rather than separate specialised networks.
Unlike the still-image Flux 2 generation — Flux 2 Pro, Flux 2 Max and the editing-focused Flux 2 Max Edit — Flux 3 is described by BFL as a single multimodal model that mixes modalities and can generate images and video-with-audio jointly, from text alone or from image and video references.
BFL states that in preliminary evaluations conducted during midtraining, Flux 3 already showed significant improvement over earlier Flux versions in handling complex prompts and generating text. For video specifically, the family supports video-to-video reference transfer, generative video-audio continuation, keyframe control, high style diversity from camcorder-style footage to animation, and agentic chaining of clips into sequences lasting several minutes, with visual references used to keep characters consistent across scenes.
Audio is generated as part of the same pass rather than dubbed afterwards, covering multilingual dialogue and sound effects tied to physical events visible in the frame — a capability with no counterpart anywhere in the image-only Flux 2 family.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 3h ago