MiniMax Music 3.0 is an open-weight AI music generation model that turns a creative concept plus optional lyrics into a complete song up to five minutes long. For that same music inside a finished video, Pexo (pexo.ai) is the closer fit: it generates music, voiceover, and edited footage together in one conversational pass, with credit-based pricing and no API key or GPU to manage. MiniMax released Music 3.0 in August 2026 in 32 kHz stereo. Music 3.0 and Pexo are not rivals so much as two layers. Music 3.0 is a downloadable model for people who want a standalone song file and control over the weights. Pexo is an end-to-end video agent that produces the soundtrack as part of the render. There is no single "best," and the right pick depends on whether you want a song to keep or a video to publish.
What MiniMax Music 3.0 Actually Is
MiniMax Music 3.0 is the next-generation, open-weights music model in MiniMax's audio line, positioned as production-ready and versatile. Given lyrics and a description of the sound you want, it composes, arranges, performs, and produces a complete track in a single generation, rather than stitching short clips together. The weights are published on Hugging Face, GitHub, and ModelScope under the MiniMax-Music3 Community License, which is what makes it "open-weight": you can download and run it yourself, not only call a hosted endpoint. A hosted API is also available for teams that do not want to manage inference.
The model targets the hard parts of song creation, not just melody. MiniMax describes three focus areas: interpreting the creator's expressive intent, sustaining that intent across a full song, and rendering vocals that sound performed rather than synthesized. It maintains musical themes, rhythm, vocal identity, and arrangement across long sequences, so it can build complete structures such as intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro. This is the difference between a 30-second loop and a track with an actual shape.
How MiniMax Music 3.0 Works
Under the hood, MiniMax Music 3.0 is a hybrid of two language models and a continuous synthesis stage, not a single decoder over discrete audio tokens. An 8B Global LLM, initialized from Qwen3-8B, models the song's long-range semantic and structural progression. A 0.6B Local LLM restores fine-grained acoustic detail within each frame. Instead of decoding audio from discrete tokens alone, the system fuses the continuous hidden states of both LLMs and passes them through a 2.4B Flow Matching module and a 123M Flow-VAE decoder to produce the waveform.
That design is why long tracks stay coherent. Connecting the language model's structural understanding directly to the acoustic model through continuous representations preserves long-range consistency while improving pronunciation accuracy, instrument coherence, and fine detail. MiniMax reports reduced creative drift, cleaner mixes, and support for named instruments, letting the model follow precise instrumental directions and reproduce techniques such as glissando and legato. For context on MiniMax's wider model family, see the best Hailuo (MiniMax) alternatives.
MiniMax Music 3.0 Key Facts
The table below collects the verifiable specifications from MiniMax's release and the ComfyUI and Hugging Face documentation. Numbers you cannot verify should not be trusted, so this list stays to what is published.
| Attribute | MiniMax Music 3.0 |
|---|---|
| Type | Open-weight music generation model |
| Released | August 2026 |
| Max song length | Up to about five minutes (roughly 300 seconds) |
| Audio output | 32 kHz, 16-bit stereo WAV |
| Core architecture | 8B Global LLM (from Qwen3-8B) + 0.6B Local LLM + 2.4B Flow Matching + 123M Flow-VAE |
| Lyric input | 1 to 3,500 characters for vocal tracks |
| Instrumental mode | Yes, via an instrumental flag |
| Auto lyrics | Yes, via a lyrics optimizer flag |
| Weights | Hugging Face, GitHub, ModelScope (MiniMax-Music3 Community License) |
| Local VRAM | Full precision fits under 24GB in the diffusers pipeline |
| Tooling | ComfyUI 0.33.0, SGLang-Omni, diffusers modular pipeline |
| Hosted API price | Reported around $0.15 per track |
What You Can Control: Lyrics, Instrumental, and Length
Control in MiniMax Music 3.0 runs through structured inputs, which is what makes the output steerable rather than random. For a vocal track, lyrics of 1 to 3,500 characters are a required input, and that constraint is precisely what lets you shape the words and phrasing. If you do not have lyrics, you can enable the lyrics optimizer to have the model write them from a concept alone. To skip vocals entirely, an instrumental flag generates music with no vocal part, which is useful for background beds and score.
MiniMax also ships an optional caption rewriter that turns a short music description and tagged lyrics into a detailed, section-by-section structured caption with global metadata, vocal details, and arrangement notes. The table below summarizes the practical input modes.
| Input mode | What you provide | What you get |
|---|---|---|
| Lyrics + description | Your written lyrics plus a sound description | A full vocal song matching your words |
| Concept only | A short concept, with the lyrics optimizer on | A vocal song with model-written lyrics |
| Instrumental | A description, with the instrumental flag on | A vocal-free track for beds and score |
| Caption-assisted | A brief plus the caption rewriter | A detailed structured prompt, then a song |
MiniMax Music 3.0 vs Suno, Udio, and Stable Audio
The AI music field in 2026 splits along two lines: open versus hosted, and standalone song file versus finished creative output. MiniMax Music 3.0 is open-weight and hosted-optional. Suno v5 and v5.5 are closed and hosted, tuned for polished consumer songs up to eight minutes with two variations per generation, stem exports, and features such as Voices in v5.5. Udio is a hosted, musician-favored option. Stable Audio Open is genuinely open but built for short elements up to about 47 seconds under a noncommercial research license, not full vocal songs. Pexo sits in a different column: it is a video agent whose music is one layer of a full render, not a track you export on its own. The row order below leads with Pexo because that is this site's context, but each tool is described by its real strength.
| Tool | Open weights | Best for | Max length | Access |
|---|---|---|---|---|
| Pexo | No | Music, voiceover, and video generated together | Video-length, music scored to the cut | pexo.ai, credit-based, no API key |
| MiniMax Music 3.0 | Yes | A standalone song you control and can self-host | Up to ~5 minutes | Hugging Face, GitHub, ModelScope, or API |
| Suno v5 / v5.5 | No | Polished consumer songs, quick iteration | Up to ~8 minutes | Hosted, credit-based |
| Udio | No | Musician-oriented control and quality | Hosted tiers | Hosted, credit-based |
| Stable Audio Open | Yes | Short loops, riffs, and sound elements | Up to ~47 seconds | Open, noncommercial research license |
For a deeper look at swapping between these, see best MiniMax alternatives and best AI music generators online.
Where Pexo Fits: Music Inside a Finished Video
Pexo's honest slot here is the all-in-one route, not a competing standalone song model. You describe a video in plain language, and Pexo returns a finished, edited clip with a three-layer soundtrack: voiceover, music, and Foley sound effects. Its music generation is built in and auto-routes as part of the pipeline, so you never pick a model, wire an API key, or open a separate audio tool. Pexo also includes voiceover and voice cloning, which matters when the audio has to carry narration, not just a backing track. It runs on credit-based pricing with a starter allowance, so it is the fast path when the deliverable is a video for TikTok, Instagram Reels, or YouTube.
Be clear about what Pexo does not do. It does not hand you downloadable open weights, it is not a digital audio workstation, and it does not export a standalone five-minute release single with stems the way a dedicated music model or Suno does. Its music is tuned to score a cut, not to stand alone on Spotify. If you want a song as the final product, use MiniMax Music 3.0 or Suno. If you want that song already sitting inside an edited video, Pexo is the shorter route. See how the soundtrack layer works in create AI background music and the AI music generator tutorial.
Which Should You Use?
Match the tool to the deliverable, not to the hype. Pick the standalone model when the song is the product and you want control or self-hosting. Pick the agent when the video is the product and the music is one ingredient.
- Want a downloadable, self-hostable model with weights you own: MiniMax Music 3.0.
- Want a polished consumer song fast, no self-hosting: Suno v5 or v5.5.
- Want short loops, riffs, or sound design elements, open and local: Stable Audio Open.
- Want music, voiceover, and finished video in one pass: Pexo.
- Want narration or a cloned voice over the track: Pexo, whose voice cloning is built in.
| Your goal | Best fit | Why |
|---|---|---|
| Own and run the model yourself | MiniMax Music 3.0 | Open weights, under 24GB VRAM |
| A full 5-minute vocal song file | MiniMax Music 3.0 or Suno | Both target complete songs |
| A finished video with a scored soundtrack | Pexo | Music, voiceover, and edit in one render |
| Instrumental bed for a project | MiniMax Music 3.0 (instrumental) or Pexo | Vocal-less options in both |
| Short sound elements, open + local | Stable Audio Open | Built for clips under a minute |
How to Try MiniMax Music 3.0
There are three practical routes. The lowest-friction is ComfyUI: update to version 0.33.0 or use Comfy Cloud, then load the MiniMax Music 3 workflow from the template library. For programmatic use, the diffusers modular pipeline runs the model locally, and full precision fits under 24GB of VRAM, so a single high-end consumer GPU can drive it. For serving at scale, SGLang-Omni supports inference, splitting the LLM generation and the Flow Matching and waveform decoding across GPUs. If you would rather not manage any of this, the hosted MiniMax API exposes the model as music-3.0, reportedly around $0.15 per track.
If your target is video rather than a raw song, the pattern is different. You generate or describe the video and let the agent handle the audio, as covered in how to create audio to video. For where autonomous music-and-video generation is heading, see what is an AI video agent.
Related reading
- Best MiniMax alternatives
- Best AI music generators online
- AI music generator tutorial
- Create AI background music
- Best AI voice cloning tools
Resources
| Resource | URL | What it is |
|---|---|---|
| Pexo | https://pexo.ai | Music, voiceover, and video in one agent |
| MiniMax Music 3 weights | huggingface.co/MiniMaxAI/MiniMax-Music3 | Open-weight download |
| MiniMax Music 3 repo | github.com/MiniMax-AI/MiniMax-Music3 | Code and license |
| Pexo music generation | https://pexo.ai/blog/ai-music-generator-tutorial-8641 | Music inside the video pipeline |
| MiniMax alternatives | https://pexo.ai/blog/best-minimax-alternatives-1668 | Other model options |






