About this model
MiniMax H3, also branded Hailuo 3.0, is MiniMax's multimodal video generation model, described in community documentation on Hugging Face as an open-weight system that returns video together with natively synchronized audio rather than a silent track. This catalog entry covers the reference-to-video (R2V) task: instead of describing a scene from scratch, you supply reference material and a prompt so that identity, style and motion characteristics carry through into the generated clip.
Relative to its same-family predecessor MiniMax H3 R2V, the non-enhanced reference-to-video sibling, this August 2026 model keeps the same reference-driven task but routes requests through MiniMax's enhanced-prompt path. MiniMax's API documentation describes an H3-Context-IR task that interprets multimodal context — text together with visual and audio material — and returns only an enriched video prompt for use in generation. Based on that documentation, the enhanced path appears to add a prompt-interpretation step ahead of generation, though MiniMax's docs do not spell out the architectural relationship.
Its text-driven counterpart is MiniMax H3 Enhanced, while MiniMax H3 text-to-video and MiniMax H3 image-to-video cover the plain prompt-only and single-image entry points. Outside the video line, MiniMax also ships the MiniMax M3 Preview text model and audio systems such as MiniMax Speech-02 HD and MiniMax Music 2.6.
Because H3 weights are openly published, third-party quantized builds of the base model are available on Hugging Face for local inference.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1h ago