About this model
MiniMax H3 (text-to-video) is the prompt-driven entry point to MiniMax's H3 video generation family, released in July 2026. Where earlier MiniMax video systems were positioned as standalone clip generators, MiniMax's own developer documentation frames H3 as a "new-generation open general-purpose multimodal video model," signalling a shift toward one model handling several conditioning modes rather than a set of separate task-specific checkpoints.
This variant takes a written description and synthesises a video clip with no visual input required, making it the most open-ended of the three H3 surfaces. Its siblings share the same underlying generation stack but change the conditioning: MiniMax H3 animates a supplied still frame, while MiniMax H3 R2V uses reference material to carry subject or style identity into the output. Choosing between them is mostly a question of how much control you want over the starting appearance.
Because detailed technical specifications for H3 are not published in the primary documentation available here, this page avoids quoting resolution, duration or benchmark figures from secondary sources. What is verifiable is the family framing and release timing.
H3 sits within a broader MiniMax lineup that includes the multimodal text model MiniMax M3 Preview, the MiniMax Speech-02 HD voice system, and MiniMax Music 2.6 — components that can be combined into an end-to-end script, narration and video pipeline.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 23h ago