MiniMaxMiniMax·🎬 Video Generation·New

MiniMax H3 Max

private
Try on Venice.ai ↗
Quick reference
MiniMax H3 Max — TLDR
  • 🆕 Max tier of MiniMax's H3 text-to-video line, released August 2026.
  • 🏢 From MiniMax, the lab behind the Hailuo video apps.
  • 🧠 H3 family is omni-modal: text, images, video, audio in one context.
  • 📏 Base H3 outputs up to 15 seconds at up to 2K.
  • 🔧 Native stereo audio: dialogue, effects and music generated with the video.
  • 🌐 Core H3 weights published under the MiniMax H3 Community License.
  • 🎯 Text-prompt entry point; a sibling Max variant covers image-to-video.
  • 📚 Max-tier-specific specifications are not detailed in MiniMax's public H3 documentation.
💰 Pricing
$0.150 – $0.720
per generation
📅 On Venice since
Aug 27, 2026
0 days ago
Provider

MiniMax is an AI company building generative models across multiple modalities, with a focus that spans both language understanding and audio creation. Their rapid release cadence in early 2026—delivering several new models within just a few months—reflects…

Read full profile →
15 models on Venice
7 video · 4 text · 3 music · 1 tts
Since Feb 12, 2026

About this model

MiniMax H3 Max is the top tier of MiniMax's H3 text-to-video family, released in August 2026 alongside its counterpart MiniMax H3 Max (image-to-video). It takes a written prompt and returns a generated clip with synchronized audio, sitting at the head of a line that began with MiniMax H3 in July 2026 and continued with MiniMax H3 Enhanced in early August.

The underlying H3 system is described by MiniMax as a general-purpose, omni-modal generative model: it accepts a unified context mixing text, images, video and audio, and produces video with native stereo audio at resolutions up to 2K and durations up to 15 seconds. Rather than layering a soundtrack on afterwards, dialogue, sound effects and music are modelled jointly with the picture. Reference-driven and first/last-frame modes exist as separate family members, including MiniMax H3 R2V Enhanced.

On MiniMax's own platform, the hosted pipeline adds two server-side stages beyond the raw weights: a context-refinement step that rewrites multimodal instructions into the representation the model consumes best, and a regeneration step that takes a 768p result plus its original context and re-renders it at native 2K. Base H3 weights are openly released under the MiniMax H3 Community License, with regional use restrictions in the license text.

Because MiniMax has not published separate documentation for the Max tier, the quality, speed or duration differences relative to the standard and Enhanced H3 text-to-video variants are best verified directly against the provider's current API reference.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 20h ago