MiniMaxMiniMax·🎬 Video Generation·New

MiniMax H3 Enhanced

anonymized
Try on Venice.ai ↗
Quick reference
MiniMax H3 Enhanced — TLDR
  • 🆕 Enhanced text-to-video path of MiniMax's open-weights H3 omni-modal generator
  • 📏 Clips of roughly 5–15 seconds, rendered up to 2K
  • 🔧 Refinement pass re-runs a 768p draft plus original context through H3
  • 👁️ Reads text, images, video and audio inside one unified context
  • 💬 Generates video with native synchronized audio in a single pass
  • 🎯 Prompts routed through H3-Context-IR, MiniMax's context-preprocessing layer
  • 🔒 Open weights under MiniMax's own community license agreement
  • 🏢 Released August 2026
💰 Pricing
$0.810 – $2.44
per generation
📅 On Venice since
Aug 5, 2026
-1 days ago
Provider

MiniMax is an AI company building generative models across multiple modalities, with a focus that spans both language understanding and audio creation. Their rapid release cadence in early 2026—delivering several new models within just a few months—reflects…

Read full profile →
12 models on Venice
5 video · 3 text · 3 music · 1 tts
Since Feb 12, 2026

About this model

MiniMax H3 Enhanced is the higher-fidelity text-to-video endpoint in MiniMax's H3 line, the company's general-purpose omni-modal generation family. The base sibling MiniMax H3 already turns a prompt alone into video at up to 2K across multiple aspect ratios, in durations of roughly five to fifteen seconds. Enhanced is the quality-oriented variant of that same underlying model rather than a separate architecture.

The distinguishing element is how the final 2K frame data is produced. MiniMax's model card describes a process in which an initial 768p generation is sent back through H3 together with the original multimodal context for a second pass, so the model can reconstruct detail using both its own draft and the user's instructions instead of applying a conventional post-hoc upscaler. Prompt handling also depends on H3-Context-IR, a hosted preprocessing system that rewrites free-form multimodal input into structured sections; MiniMax's own card calls it critical to output quality.

Everything else is inherited from the family. H3 accepts text, images, video and audio in a single context and generates video with native synchronized audio, without a separate audio stage.

If your workflow supplies reference material rather than text alone, the companion endpoints MiniMax H3 R2V Enhanced, MiniMax H3 R2V and MiniMax H3 (image-to-video) expose the same model through different input paths. Weights for the base checkpoint are gated and released under MiniMax's custom community license.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1h ago