What Is MiniMax Music 3? Open-Weights Music Model

Priya Nair

Priya Nair

AI Infrastructure Analyst

Published: August 14, 2026
MiniMax Music 3 Release Coverage editorial cover

TLDRMiniMax Music 3 is an open-weights model generating full 5-minute songs at 32 kHz stereo, live on Hugging Face and ComfyUI since Aug 13, 2026.

MiniMax Music 3 Is an Open-Weights Model That Generates Full 5-Minute Songs

MiniMax Music 3 is an open-weights AI music generation model from MiniMax that turns lyrics and a text music description into a complete, produced song of up to five minutes in a single generation. It outputs 32 kHz, 16-bit stereo audio and went live on August 13, 2026, with weights on Hugging Face and native support in ComfyUI. MiniMax's official announcement called it a "Next-Generation Open-Weights Production-Ready & Versatile Music Model," with a cloud API listed as coming soon.

Key Takeaways

  • MiniMax Music 3 generates complete songs up to five minutes long from lyrics plus a structured music caption.
  • It shipped as open weights under the MiniMax-Music3 Community License, hosted on Hugging Face.
  • The architecture pairs an 8B Global LLM (initialized from Qwen3-8B) with a 0.6B Local LLM, plus a Flow-Matching + Flow-VAE synthesis stack.
  • Output is 32 kHz, 16-bit stereo WAV audio.
  • It runs natively in ComfyUI via an official Text to Music workflow, with a Comfy Cloud workflow available too; a hosted MiniMax cloud version was announced as "coming soon."
  • Release date was August 13, 2026 per MiniMax's official blog.

What Is MiniMax Music 3?

MiniMax Music 3 (also written Music 3.0) is a text-and-lyrics-to-music model built by MiniMax, the company behind the MiniMax M-series language models and the H3 video model. Given a creative concept and optional lyrics, the model composes, arranges, performs, and produces a finished track in one pass.

It targets the parts of music generation that short prompts struggle with: understanding a creator's expressive intent, holding that intent across an entire song, rendering instruments with physical realism, and producing vocals that sound performed rather than synthesized. The status is released, not leaked. MiniMax's official blog dated the launch August 13, 2026, and the weights are live on the MiniMaxAI/MiniMax-Music3 Hugging Face repository, with a ComfyUI-optimized distribution at Comfy-Org/MiniMax-Music-3.

The model belongs to MiniMax's music family, separate from its M-series language models and its Speech text-to-speech line. Its defining pitch is that it produces Complete Songs with Long-Range Coherence rather than the thirty-second clips typical of earlier tools.

MiniMax Music 3 at a Glance

SpecificationDetail
DeveloperMiniMax
TypeText-and-lyrics-to-music generation model
ModalityLyrics + music description → produced song audio
Song lengthUp to 5 minutes per single generation
Audio output32 kHz, 16-bit stereo WAV (local weights)
Architecture8B Global LLM (from Qwen3-8B) + 0.6B Local LLM + Flow Matching (2.4B) + Flow-VAE (123M)
Tokenizer8-layer Residual Vector Quantization (RVQ)
Pricing (weights)Free to download and self-host
Pricing (hosted MiniMax cloud API)Not yet confirmed (announced "coming soon")
AvailabilityLive: Hugging Face weights + native ComfyUI workflow (local and Comfy Cloud)
LicenseMiniMax-Music3 Community License
Release dateAugust 13, 2026

How MiniMax Music 3 Works / What Makes It Different

MiniMax Music 3 uses a hierarchical autoregressive architecture that MiniMax calls the Hybrid-LM. It separates long-range musical structure from frame-level acoustic detail across two models working together.

The Global LLM (8B), initialized from Qwen3-8B, predicts the first RVQ codebook frame by frame and models the song's long-range semantic and structural progression. The Local LLM (0.6B) predicts the remaining acoustic codebooks within each frame and restores fine-grained detail. This split lets the model keep song-level stability while resolving texture inside every frame.

The tokenizer is an 8-layer RVQ stack. The first codebook carries core semantics and structure; the remaining seven progressively encode acoustic residuals. The first layer holds 16,384 entries per an independent technical writeup of the model card.

Where most systems decode audio straight from discrete tokens, Music 3 uses Continuous Hidden-State Synthesis. It fuses the final hidden states of the Global and Local LLMs, conditions a 2.4B flow-matching module on them, and decodes through a 123M Flow-VAE to reach 32 kHz stereo audio. The synthesis path runs: fused LLM features → Flow Matching → Flow-VAE latent → Flow-VAE decoder → final audio. The Flow-VAE architecture is adapted from MiniMax Speech and retrained for the dynamic range of music. MiniMax positions this continuous path as the reason vocal articulation and instrumental texture hold up over long tracks.

Control happens through two inputs. Lyrics carry the words and can include section tags such as [Intro], [Verse], [Pre-Chorus], [Chorus], [Bridge], [Instrumental], [Solo], and [Outro]. A Structured Caption defines global metadata (genre, BPM, key, emotional progression), vocal details, and arrangement evolution across the song.

What You Can Do With MiniMax Music 3

The core workflow is generating a full track from a caption plus lyrics. ComfyUI's documentation describes generating complete songs with intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro, all steered by section tags.

Beyond straight text-to-music, the model supports instrumental-only generation and vocal-led songs with control over melody, pronunciation, and layered harmonies. The official demo page groups results across Pop, Rock, R&B, Hiphop, Blues, Bossa Nova, and more, and highlights a Regional Genres track fine-tuned for Chinese sub-genres including Zhongguofeng, Gufeng Pop, and Folk.

Now that the weights are out, hands-on signal is arriving fast. Community testers have run it locally in ComfyUI, and one reported a "Jazzy tribute" generated locally on Apple Metal, showing it working outside a single vendor's stack. If you want a hosted alternative to compare full-song generation against, Suno API covers a similar text-to-music task. For the visual and video side of MiniMax's stack, our coverage of generative model launches like Meta's Muse Image and Muse Video tracks the same open-weights-versus-hosted tension.

How MiniMax Music 3 Compares

The clearest comparison point raised at launch was Suno. Community accounts hoped Music 3 would beat Suno V5.5, and one r/StableDiffusion post summarized: "Today I have unsubscribed from Suno thanks to Minimax Music." No controlled listening test backs this.

MiniMax Music 3Suno (community reference point)
WeightsOpen weights, self-hostableHosted service
Song lengthUp to 5 minutesVaries by version
Local runNative ComfyUI workflow (local + Comfy Cloud)Not applicable
Benchmark vs. each otherNot yet confirmedNot yet confirmed

Treat the Suno framing as community expectation, not measured result. MiniMax has not published controlled benchmarks against Suno, Udio, or its own Music 2.6.

Availability: How to Access MiniMax Music 3

MiniMax Music 3 is available two ways today. First, the open weights: download from the MiniMaxAI/MiniMax-Music3 repository, or the ComfyUI-repacked files from Comfy-Org/MiniMax-Music-3, both on Hugging Face. Second, run it locally through ComfyUI's official Text to Music workflow — update ComfyUI, open Template Library > Audio, pick the MiniMax Music 3 workflow, and download the models when prompted.

MiniMax's official account confirmed the local release and said hosted cloud access is next: "Now Music 3 is open too!! Already live locally on @ComfyUI ... Cloud coming soon." (official post). ComfyUI has also published a Comfy Cloud workflow with a prompt guide (ComfyUI post).

The code is on GitHub at MiniMax-AI/MiniMax-Music3, and the model card lives at MiniMaxAI/MiniMax-Music3 on Hugging Face.

What We Don't Know Yet

Several facts remain unconfirmed in the official launch materials:

  • Hosted MiniMax cloud pricing. The MiniMax cloud API was announced as "coming soon" with no confirmed rate.
  • License scope. Weights ship under the MiniMax-Music3 Community License; permitted commercial uses need a direct read of the terms. Community claims of "fully open source" are broader than "open weights."
  • Benchmarks. No controlled comparison against Suno, Udio, or Music 2.6 has been published.
  • Local hardware floor. Community reports discuss VRAM tiers, but MiniMax has not published an official minimum spec here.

Frequently Asked Questions

What is MiniMax Music 3?

MiniMax Music 3 is an open-weights music generation model from MiniMax that composes, arranges, performs, and produces a complete song of up to five minutes from lyrics and a music description. It outputs 32 kHz, 16-bit stereo audio and runs natively in ComfyUI.

Is MiniMax Music 3 open source?

MiniMax Music 3 launched as open weights under the MiniMax-Music3 Community License, with weights live on Hugging Face. MiniMax's own framing is "open weights," and the license should be reviewed before commercial use, so "fully open source" community claims are broader than the license terms confirm.

How much does MiniMax Music 3 cost?

MiniMax Music 3 weights are free to download and run locally, so the main cost is your own GPU compute. ComfyUI also offers a Comfy Cloud workflow. A hosted MiniMax cloud version was announced as "coming soon" at launch, and official MiniMax cloud pricing was not confirmed in the launch materials covered here.

How long can MiniMax Music 3 songs be?

MiniMax Music 3 generates complete songs up to five minutes long in a single generation. It maintains theme, rhythm, vocal identity, and arrangement across that full length, producing structures like intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro.

How do I run MiniMax Music 3?

MiniMax Music 3 runs natively in ComfyUI using an official Text to Music workflow. Update ComfyUI to the latest version, download the model files from Hugging Face, load the workflow, and generate from a caption plus lyrics. A Comfy Cloud workflow with a prompt guide is also available.

MiniMax Music 3 vs Suno — which is better?

No controlled benchmark comparing MiniMax Music 3 and Suno exists in the launch materials covered here. Early community reaction framed Music 3 as a strong open-weights alternative, but quality-versus-Suno claims are unverified opinion rather than measured results.

What audio quality does MiniMax Music 3 output?

MiniMax Music 3 outputs 32 kHz, 16-bit stereo WAV audio when run locally via the open weights. It uses a continuous hidden-state synthesis path with Flow Matching and a Flow-VAE decoder rather than decoding from discrete tokens alone.

What to Watch Next

Three signals will tell you where MiniMax Music 3 lands. First, the hosted MiniMax cloud API — its price per song and rate limits will decide whether it undercuts hosted rivals. Second, the license fine print, which governs commercial use of anything you generate. Third, the first controlled listening tests against Suno and Udio, which would replace today's anecdotal quality claims with something measurable.

Building your own full-song generation? On kie.ai you can try Suno API, ElevenLabs V3, and Elevenlabs Text to Speech.

Priya Nair

About Priya Nair

Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.

View all posts by Priya Nair