MiniMax H3
Video and Sound From One Model
MiniMax H3 is a general-purpose omni-modal generation model that reads text, images, video, and sound as one context and returns audiovisual output with native stereo sound, up to 15 seconds at 2K. Available on Pexo.
What MiniMax H3 Produces, Seen Through Pexo
Every video shown here was created through Pexo from a plain-language description — no prompt syntax and no model configuration.
What Makes MiniMax H3 Different
MiniMax released H3 on July 31, 2026 as a break from its own Hailuo line. Here is what the model actually does differently.

One Context, Every Modality
MiniMax positions H3 as a general-purpose omni-modal generation model rather than a text-to-video model with features bolted on. Text, images, video, and sound enter as a single unified context, and the model returns audiovisual output. MiniMax describes the goal as removing the boundaries that separated tasks, capabilities, and modalities in earlier generation models.

High Resolution Without a Super-Resolution Module
H3 outputs up to 15 seconds at 2K. Instead of bolting on a conventional upscaler, MiniMax has the base model regenerate its own low-resolution result in context, which lets it recover small text and fine detail that a super-resolution pass can only guess at. The new H3-VAE tokenizer, with roughly a fourfold sequence-length gain from higher compression, is what makes native 2K affordable.

Sound Generated With the Picture, Not After It
Every audio output is native stereo, and MiniMax models speech, sound effects, and music jointly rather than as separate domains stitched together in post. Dialogue, effects, and room tone arrive in the same pass as the image, which is the part that usually costs a full round of sound design.

Bring Your Own Characters, Motion, and Sound
H3 accepts up to nine reference images, three reference video clips, and three reference audio clips in a single request, and MiniMax exposes generalized reference and editing in natural language across image, video, and audio pairs. That covers keeping a character consistent, transferring motion from one clip to another, and matching an existing voice or ambience.

Built for Instruction Following and On-Screen Text
From MiniMax's invited testing, H3 is described as commercial-grade across scenarios, with particular strength in instruction following, rendering text and brand information correctly, and video-to-video motion transfer. MiniMax names advertising, brand, e-commerce, product design, UI/UX, and games as the target scenarios.
MiniMax H3 vs the Hailuo Generations Before It
MiniMax frames H3 as a deliberate break rather than an upgrade, saying it abandoned the Hailuo-02 architecture because those tricks got in the way of generalizing across tasks. Here is what changed.
| Feature | MiniMax H3 | Hailuo 2.3 | Hailuo 02 | Hailuo 01 |
|---|---|---|---|---|
| Positioning | Omni-modal generation | Video generation | Video generation | Video generation |
| Native 2K output | ✓ | — | — | — |
| Native stereo audio in the same pass | ✓ | — | — | — |
| Text, image, video and audio as one context | ✓ | — | — | — |
| Instruction-based editing across modalities | ✓ | — | — | — |
| Architecture | H3-Omni Transformer | Hailuo-02 architecture | Hailuo-02 architecture | First-generation system |
Sources: MiniMax: H3, breaking the boundaries of tasks and modalities · MiniMax Hub · Hailuo AI: MiniMax H3
How to Use MiniMax H3 in Pexo: Three Steps, No Setup
No account, API key, or technical knowledge required. Pexo handles model selection, generation, and delivery so you focus on what you want to make.
Type a description in plain language on web, Telegram, WhatsApp, or Discord. There's no required format and no model to select — Pexo reads your intent and handles the rest.
Pexo routes your request to MiniMax H3, applies the right settings automatically, and runs the generation. You configure nothing.
Your video arrives ready to use. Want to refine it? Keep the conversation going instead of starting over.
What You Can Make with Pexo's MiniMax H3
No prompt writing. No model picking. Just describe what you want.
Frequently Asked Questions About MiniMax H3
What is MiniMax H3?+
MiniMax H3 is a general-purpose omni-modal generation model that MiniMax released on July 31, 2026. It understands text, images, video, and sound as one unified context and generates video with native stereo audio, up to 15 seconds at 2K resolution.
Is MiniMax H3 the same thing as Hailuo 3.0?+
Yes. H3 is the official model name and Hailuo 3.0, sometimes written Hailuo 03, is the alias that comes from the Hailuo AI app it ships in. They refer to the same model, so search results using either name are describing the same release.
How is H3 different from Hailuo 2.3?+
It is a different design rather than a bigger version. MiniMax says it abandoned the Hailuo-02 architecture because its advantages added complexity once the goal became generalizing across tasks. H3 adds native 2K, native stereo audio generated with the picture, and a single context that accepts text, images, video, and audio together.
How long can a MiniMax H3 video be, and what resolution?+
Up to 15 seconds per generation at up to 2K. MiniMax reaches 2K by having the base model regenerate its own lower-resolution output in context rather than running a separate super-resolution module, which preserves small text and fine detail.
Does MiniMax H3 generate sound?+
Yes, and that is one of its headline changes. All audio output is native stereo, with speech, sound effects, and music modeled jointly and produced in the same pass as the picture instead of being added afterward in a separate sound-design step.
Are the MiniMax H3 weights open?+
MiniMax stated at launch that it plans to release the model weights within days, subject to applicable laws and regulations, and noted it had considered compatibility with several domestic chips. Treat any specific license terms you see quoted elsewhere as unofficial until MiniMax publishes them.
What can I make with MiniMax H3 on Pexo?+
Short-form video where picture and sound have to arrive together: product and brand spots, e-commerce clips, UI and product walkthroughs, and game or concept pieces. On Pexo you describe what you want in plain language and get a finished video back, without setting up an API or managing reference-file limits by hand.
MiniMax H3 Is on Pexo: Describe It and See It
The best video model available right now takes one sentence to use.





