About this model
MiniMax H3 R2V is the reference-to-video entry point of MiniMax's H3 video generation family, released in 2026. Where a text-to-video system works from a written prompt alone, a reference-to-video mode takes user-supplied material — such as images or clips of a character, product or scene — and conditions the generated footage on it, so that the subject's appearance and styling carry through the output rather than being reinvented on each run.
MiniMax positions H3 as a new-generation, open, general-purpose multimodal video model in its own platform documentation, treating multiple input modalities as one shared context instead of separate single-purpose tasks. That framing is what separates the H3 generation from MiniMax's earlier single-input video endpoints: rather than one prompt type per model, the family exposes several conditioning routes over a common backbone.
Those routes are the siblings listed here: MiniMax H3 for prompt-only generation and MiniMax H3 for animating a still, both dated to the same 2026 release. R2V is the variant to reach for when consistency with existing assets matters more than open-ended invention.
Beyond video, MiniMax ships models across other modalities, including MiniMax M3 Preview and MiniMax M2.7 for text and reasoning, MiniMax Music 2.6 for music, and MiniMax Speech-02 HD for speech. MiniMax has not published parameter counts, licensing terms or architectural details for H3 R2V, so this description stays limited to what the provider states directly.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 23h ago