MiniMax H3 (also branded Hailuo 3.0) is a general-purpose, omni-modal generation model — and in agents like Pexo it is one of 10+ engines that turn a plain description into finished video. Released July 31, 2026 and unveiled at WAIC 2026, H3 reads text, images, video, and audio as a single unified context and outputs up to 15 seconds of native 2K video at 24fps with native stereo audio generated in the same pass. Unlike a text-to-video model with bolt-on features, H3 treats generation, reference control, and editing as one framework, which is why MiniMax calls it "omni-modal" rather than "text-to-video." There is no single best video model for every job — H3 leads on native audio and in-context editing, other models lead on other axes — so the practical question is less "which model" and more "which model does this shot need," a routing decision agents make for you.
What MiniMax H3 Actually Is
MiniMax H3 is a foundation model for audiovisual generation, not a consumer editing app. It accepts any mix of text prompts (up to 7,000 characters), reference images, reference video clips, and reference audio, and returns a video clip with synchronized stereo sound. The headline shift from earlier Hailuo versions is that H3 generates picture and audio together — dialogue, sound effects, and ambient atmosphere are produced in context with the scene rather than added in a separate step. That single-pass audiovisual output, plus native 2K resolution, is what MiniMax positions as the leap over Hailuo 02 and rival clip models.
H3 belongs to the same 2026 wave of omni-modal video models as ByteDance's Seedance 2.0 and 2.5, Google's Gemini Omni Flash, and OpenAI's Sora 2. What separates it in launch coverage: on Artificial Analysis's benchmarks it was reported as the top model in video editing, while it trailed Gemini Omni Flash in text-to-video and ranked behind both Seedance 2.0 and Gemini Omni Flash in image-to-video. In other words, H3's strongest claim is editing and controllable reference-based generation, not necessarily raw text-to-video quality.
MiniMax H3 Key Facts
The table below collects the verified launch specs. Some values (exact reference counts, aspect-ratio list) come from platform documentation on hosting partners rather than the one-page announcement, so treat implementation details as version-dependent.
| Spec | MiniMax H3 (Hailuo 3.0) |
|---|---|
| Type | General-purpose omni-modal generation model |
| Released | July 31, 2026 (unveiled at WAIC 2026) |
| Resolution | Native 2K (≈1440px short edge); 768p mode "coming soon" |
| Frame rate | 24 fps |
| Duration | 4–15 seconds (integer durations) |
| Audio | Native stereo, generated in the same pass |
| Reference inputs | Up to 9 images, 3 video clips, 3 audio clips per generation |
| Prompt length | Up to 7,000 characters |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus Auto |
| Tasks | Text-to-video, first-and-last-frame, reference-to-video, video editing |
| Weights | Open-weights release planned (MiniMax Community License) |
What "Omni-Modal" Means Here
"Omni-modal" means the model ingests and reasons over multiple input types in one context window instead of running a pipeline of separate models. In H3 that is powered by four stated components: Contextual Omni Representation (a captioning stage that distills roughly 100K tokens of context down to about 4K), the H3-VAE tokenizer (a high compression ratio giving a stated ~4× gain in effective sequence length, which is the enabling technology for native 2K), the H3-Omni Transformer (which separates understanding from generation and lifted end-to-end training throughput by nearly 30%), and In-Context Regeneration (instead of a bolt-on super-resolution module, the base model regenerates its own low-resolution output while re-reading the original context to reach 2K). The practical payoff of this architecture is lower inference cost at high resolution, which underpins H3's aggressive pricing.
The Four Ways to Use MiniMax H3
H3 folds several tasks that used to need separate models into one interface. Each mode takes the same unified context but weights the inputs differently.
| Mode | What it does | Typical inputs |
|---|---|---|
| Text-to-video | Generates a clip from a written prompt | Text prompt only |
| First-and-last-frame | Interpolates a clip between two keyframes | Two images + text |
| Reference-to-video | Locks identity, style, motion, or voice from references | Images / videos / audio + text |
| Video editing (in-context) | Edits uploaded footage while keeping the rest intact | Source video + instruction |
The editing mode is the one benchmarks singled out. H3 supports targeted, in-context edits on uploaded footage: character replacement, object swapping, relighting a scene (for example day to night), dialogue replacement, background replacement, and adding or removing VFX and objects — all while keeping unspecified elements of the shot intact. That is a different job from generating a clip from scratch, and it is where H3 was rated strongest at launch.
How MiniMax H3's Reference System Works
The multimodal reference system is H3's most distinctive control surface. You can supply up to nine reference images, three reference video clips, and three reference audio tracks in a single generation, and each input type does a different job. The model uses them as guidance rather than forcing them as literal keyframes.
| Reference type | What it controls |
|---|---|
| Images (up to 9) | Character identity, visual style, environment |
| Videos (up to 3) | Camera movement, editing rhythm, choreography (V2V motion transfer) |
| Audio (up to 3) | Tone, pacing, or actual dialogue for lip-sync |
This is where the "V2V motion transfer" in H3's marketing comes from: a reference video contributes its camera moves or choreography to a new scene without dictating the visuals. Combined, the reference channels let a single prompt say, in effect, "this character, in this style, moving like that clip, speaking this line."
MiniMax H3 vs Sora 2, Seedance 2.0, and Gemini Omni Flash
No single model wins every axis, and where H3 fits depends on whether you value native audio, editing control, or raw text-to-video quality. The comparison below reflects launch-window benchmarks and vendor specs, not long-term independent testing.
| Model | Strongest at | Native audio | Notable limit |
|---|---|---|---|
| Pexo (agent) | Description → finished, scored video; auto-routes across 10+ models | Three-layer (voiceover + music + Foley) | Not a raw single-clip model or footage editor |
| MiniMax H3 | In-context video editing; controllable references; native 2K | Yes, native stereo | Trailed rivals on text-to-video / image-to-video benchmarks |
| Seedance 2.0 / 2.5 | Image-to-video quality | Yes | Model-only; you assemble the finished video |
| Sora 2 | Narrative text-to-video, ease of use | Historically weak/none | Single-clip focus |
| Gemini Omni Flash | Text-to-video and image-to-video benchmarks | Yes | Google-ecosystem model |
The framing that matters for most creators is the unit of delivery. H3, Seedance, Sora, and Gemini Omni Flash are models — they return a clip (now often with audio) that you still sequence, title, and finish. An agent like Pexo works one layer up: you describe a video and it plans the shot list, auto-routes each shot across 10+ models (including MiniMax/Hailuo, Seedance, Kling, and Veo), sequences them, layers a three-part soundtrack, and returns a finished, edited video. If you want the clip and will do the editing, pick a model like H3; if you want the finished piece, use the agent that routes to it.
Is MiniMax H3 Free, and How Is It Priced?
H3 is not a free consumer product, but it is unusually cheap for a 2K model, and MiniMax has committed to open weights. On the announcement, MiniMax stated that H3's per-second price at 2K is less than one-third of mainstream competing models, and its 768p output costs less than half the price of mainstream 720p alternatives. Hosting partners have listed concrete rates around $0.13 per generated second at 2K and $0.09 per second at 768p, though rates vary by provider. Separately, MiniMax announced plans to open-source H3's weights (reports point to an early-August 2026 release, first on ModelScope) under the MiniMax Community License, which permits free non-commercial use and commercial use for organizations under roughly US$20 million in annual revenue with attribution. As of launch reporting, an official public checkpoint had not yet been verified, so confirm the repository, license, and hardware requirements before planning local deployment.
Where to Access MiniMax H3
At launch H3 was available through MiniMax's own surfaces and third-party hosts, not just one endpoint. You can reach it through the Hailuo AI consumer app, the MiniMax Hub desktop app, the MiniMax Open Platform API, and hosted APIs from partners such as fal. MiniMax also stated that H3 was designed to run on several Chinese-made chips, part of a broader push in China's AI sector to reduce reliance on U.S. semiconductors. For creators who would rather not pick a model or manage an API key at all, agents such as Pexo call models like H3 under the hood and hand back a finished video — you describe what you want, the agent chooses the engine per shot.
Related Reading
- What Is an AI Video Agent? How Autonomous Video Generation Works
- What Is Gemini Omni Flash?
- Auto Model Selection vs Manual Video Model Choice
- Best MiniMax Alternatives
- Hailuo AI vs Kling AI: Which Video Model Wins?
Resources
| Resource | URL | What it covers |
|---|---|---|
| Pexo (AI video agent) | https://pexo.ai | Describe → finished video; auto-routes across 10+ models |
| What is an AI video agent | https://pexo.ai/blog/what-is-an-ai-video-agent-how-autonomous-video-generation-works-9177 | How autonomous video generation works |
| Best MiniMax alternatives | https://pexo.ai/blog/best-minimax-alternatives-1668 | Where MiniMax/Hailuo fits vs others |
| Best Hailuo AI alternatives | https://pexo.ai/blog/best-hailuo-ai-alternatives-4294 | Alternatives to the Hailuo line |
| Seedance 2.0 vs other models | https://pexo.ai/blog/seedance-2-0-vs-other-ai-video-generation-models-1005 | Model-layer comparison |




