About this model
MiniMax H3 Max is the top tier of MiniMax's H3 text-to-video family, released in August 2026 alongside its counterpart MiniMax H3 Max (image-to-video). It takes a written prompt and returns a generated clip with synchronized audio, sitting at the head of a line that began with MiniMax H3 in July 2026 and continued with MiniMax H3 Enhanced in early August.
The underlying H3 system is described by MiniMax as a general-purpose, omni-modal generative model: it accepts a unified context mixing text, images, video and audio, and produces video with native stereo audio at resolutions up to 2K and durations up to 15 seconds. Rather than layering a soundtrack on afterwards, dialogue, sound effects and music are modelled jointly with the picture. Reference-driven and first/last-frame modes exist as separate family members, including MiniMax H3 R2V Enhanced.
On MiniMax's own platform, the hosted pipeline adds two server-side stages beyond the raw weights: a context-refinement step that rewrites multimodal instructions into the representation the model consumes best, and a regeneration step that takes a 768p result plus its original context and re-renders it at native 2K. Base H3 weights are openly released under the MiniMax H3 Community License, with regional use restrictions in the license text.
Because MiniMax has not published separate documentation for the Max tier, the quality, speed or duration differences relative to the standard and Enhanced H3 text-to-video variants are best verified directly against the provider's current API reference.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 20h ago