About this model
MiniMax H3 Max (image-to-video) is the highest tier in MiniMax's H3 video generation family, dated August 2026 in this catalog. It takes a source image plus a natural-language motion description and animates it into a coherent clip. Its text-driven counterpart is MiniMax H3 Max.
The underlying H3 system, announced by MiniMax in July 2026, is described by the lab as a general-purpose omni-modal generation model that jointly understands text, images, video and audio in one context and generates video with native stereo audio at up to 2K resolution and 15 seconds in length. Audio is generated jointly with the video rather than added in a separate post-processing step. MiniMax cites components including Contextual Omni Representation, H3-VAE, an H3-Omni Transformer and in-context regeneration.
Within the family, this model sits above the original MiniMax H3 image-to-video endpoint from July 2026 and the August MiniMax H3 Enhanced and MiniMax H3 R2V Enhanced tiers, which extend the same base model to prompt-enhanced and reference-driven workflows.
MiniMax has not published tier-by-tier benchmark figures for the Max variant, so no performance numbers are quoted here; capabilities such as 2K output and multimodal referencing derive from the H3 model card and MiniMax's own documentation.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 20h ago