MiniMaxMiniMax·🎬 Video Generation·New

MiniMax H3 R2V Enhanced

anonymized
Try on Venice.ai ↗
Quick reference
MiniMax H3 R2V Enhanced — TLDR
  • 🆕 Reference-to-video variant of MiniMax's H3 (Hailuo 3.0) generation family.
  • 🆕 Released August 5, 2026; newest model in this reference-to-video line.
  • 👁️ Takes reference visual material alongside a text prompt to guide output.
  • 🧠 "Enhanced" path is associated with MiniMax's multimodal context interpretation task.
  • 💬 MiniMax API docs describe an H3-Context-IR task returning an enriched video prompt.
  • 📏 Community write-ups describe H3 as generating video with natively synchronized audio.
  • 🌐 H3 is published as open weights; quantized community builds exist on Hugging Face.
  • 🔧 Siblings cover text-to-video, image-to-video and non-enhanced reference-to-video.
💰 Pricing
$0.810 – $2.44
per generation
📅 On Venice since
Aug 5, 2026
-1 days ago
Provider

MiniMax is an AI company building generative models across multiple modalities, with a focus that spans both language understanding and audio creation. Their rapid release cadence in early 2026—delivering several new models within just a few months—reflects…

Read full profile →
12 models on Venice
5 video · 3 text · 3 music · 1 tts
Since Feb 12, 2026

About this model

MiniMax H3, also branded Hailuo 3.0, is MiniMax's multimodal video generation model, described in community documentation on Hugging Face as an open-weight system that returns video together with natively synchronized audio rather than a silent track. This catalog entry covers the reference-to-video (R2V) task: instead of describing a scene from scratch, you supply reference material and a prompt so that identity, style and motion characteristics carry through into the generated clip.

Relative to its same-family predecessor MiniMax H3 R2V, the non-enhanced reference-to-video sibling, this August 2026 model keeps the same reference-driven task but routes requests through MiniMax's enhanced-prompt path. MiniMax's API documentation describes an H3-Context-IR task that interprets multimodal context — text together with visual and audio material — and returns only an enriched video prompt for use in generation. Based on that documentation, the enhanced path appears to add a prompt-interpretation step ahead of generation, though MiniMax's docs do not spell out the architectural relationship.

Its text-driven counterpart is MiniMax H3 Enhanced, while MiniMax H3 text-to-video and MiniMax H3 image-to-video cover the plain prompt-only and single-image entry points. Outside the video line, MiniMax also ships the MiniMax M3 Preview text model and audio systems such as MiniMax Speech-02 HD and MiniMax Music 2.6.

Because H3 weights are openly published, third-party quantized builds of the base model are available on Hugging Face for local inference.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 1h ago