About this model
MiniMax H3 Enhanced is the higher-fidelity text-to-video endpoint in MiniMax's H3 line, the company's general-purpose omni-modal generation family. The base sibling MiniMax H3 already turns a prompt alone into video at up to 2K across multiple aspect ratios, in durations of roughly five to fifteen seconds. Enhanced is the quality-oriented variant of that same underlying model rather than a separate architecture.
The distinguishing element is how the final 2K frame data is produced. MiniMax's model card describes a process in which an initial 768p generation is sent back through H3 together with the original multimodal context for a second pass, so the model can reconstruct detail using both its own draft and the user's instructions instead of applying a conventional post-hoc upscaler. Prompt handling also depends on H3-Context-IR, a hosted preprocessing system that rewrites free-form multimodal input into structured sections; MiniMax's own card calls it critical to output quality.
Everything else is inherited from the family. H3 accepts text, images, video and audio in a single context and generates video with native synchronized audio, without a separate audio stage.
If your workflow supplies reference material rather than text alone, the companion endpoints MiniMax H3 R2V Enhanced, MiniMax H3 R2V and MiniMax H3 (image-to-video) expose the same model through different input paths. Weights for the base checkpoint are gated and released under MiniMax's custom community license.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1h ago