What Is MiniMax H3 Max? 15-Second Video at $0.05/s
Lukas Vogel
Applied Research Editor

TLDRMiniMax H3 Max delivers 5–15-second video at 480P/768P, with official API pricing from $0.05 per second.
Meet MiniMax H3 Max, the 15-Second Video Model for Near-Real-Time Generation
MiniMax H3 Max is a video generation model jointly released by MiniMax and fal.ai, post-trained by fal.ai from MiniMax H3 open weights and optimized for fast text-to-video and image-to-video generation. It produces 5–15-second clips at 480P or 768P, and the official MiniMax Pay-as-you-go documentation lists prices of $0.05 per second at 480P and $0.08 per second at 768P. MiniMax H3 Max is available through supported API access and web-based creation tools, while its own weights and license remain unconfirmed.
Key Takeaways
- MiniMax H3 Max is fal.ai's post-trained variant of MiniMax H3, not a separately documented foundation model from MiniMax.
- The official model documentation supports Text-to-Video and Image-to-Video generation.
- Output runs from 5 seconds to 15 seconds at 480P or 768P resolution.
- Official list pricing is $0.05 per second at 480P and $0.08 per second at 768P.
- MiniMax and fal.ai position the model for faster-than-real-time video generation.
- The base MiniMax H3 is open-weight, but H3 Max's own weight-release status is not confirmed.
What Is MiniMax H3 Max?
MiniMax H3 Max is a specialized video model built on MiniMax H3. fal.ai says it post-trained MiniMax H3 for stronger prompt adherence and visual quality, then co-designed the inference stack around the model to increase speed without discarding those post-training gains. MiniMax officially endorsed that description and credited fal.ai with pushing H3 forward through post-training and inference optimization. fal.ai's launch announcement
The distinction matters because “H3 Max” can sound like an official larger or premium checkpoint in the MiniMax family. The supplied evidence instead describes it as a third-party post-trained variant based on MiniMax H3. MiniMax H3 remains the underlying open-weight model; H3 Max is the result of fal.ai adapting that model for a particular quality-latency target.
The model is released rather than merely rumored. fal.ai announced the launch on August 27, 2026, and MiniMax described H3 Max as a practical demonstration of what open weights can enable. The official MiniMax documentation now lists H3 Max alongside MiniMax H3 in its video-generation guide.
MiniMax H3 Max is fal.ai's post-trained, speed-optimized variant of MiniMax H3, not a separately documented MiniMax foundation model.
The primary problem it addresses is waiting time during video iteration. A creator can test more prompts, camera movements, compositions, and image references when a 15-second generation arrives in roughly the time required to watch it. That use case is different from maximizing resolution or producing long single-shot films.
MiniMax H3 Max at a Glance
| Specification | MiniMax H3 Max |
|---|---|
| Developer | fal.ai, using MiniMax H3 as the foundation; jointly released with MiniMax |
| Type | Video generation model |
| Modality | Video output from text prompts or image inputs |
| Generation modes | Text-to-Video (T2V) and Image-to-Video (I2V) |
| Output duration | 5–15 seconds |
| Output resolution | 480P and 768P |
| Frame rate | 24 FPS, according to a third-party product listing; official MiniMax documentation does not confirm this field |
| Prompt limit | Up to 7,000 characters |
| Context window | Not yet confirmed |
| Official API pricing | $0.05 per second at 480P; $0.08 per second at 768P |
| Availability | Official MiniMax API documentation and supported fal.ai web access |
| Weights | Not yet confirmed for H3 Max |
| License | Not yet confirmed for H3 Max |
The pricing is output-based. Input images are currently not billed under the official pricing information supplied for H3 Max. At list rates, a 15-second clip costs $0.75 at 480P or $1.20 at 768P.
How MiniMax H3 Max Works and What Makes It Different
Three useful terminology anchors describe the model's design: Prompt-Adherence Tuning, Co-Designed Inference Stack, and First-and-Last-Frame I2V.
Prompt-Adherence Tuning refers to fal.ai's post-training work on H3. The stated objective is better alignment between a written instruction and the resulting subject, action, visual style, or camera direction. The available evidence does not publish the training dataset, parameter count, architecture changes, or a reproducible training recipe.
Co-Designed Inference Stack describes the second half of the approach. fal.ai says the inference system was developed alongside post-training, with optimizations repeatedly tested against quality. MiniMax says this work required both frontier model research and deep inference and kernel optimization. The claim is therefore about the full production system, not just the neural checkpoint running in isolation.
First-and-Last-Frame I2V is the most concrete control mode. A request can use an opening image, an ending image, or both, with a text prompt describing the transition. This gives creators more control over where a short clip begins and ends than a text-only request.
The official API guide lists common aspect ratios or adaptive ratio behavior, subject to the API reference. It also documents image inputs from 256 pixels to 5,760 pixels on each dimension, with a 30-megabyte limit per image and a 64-megabyte request-body limit. Those constraints are operational details, not evidence of a larger context window.
Audio needs a confidence label. Third-party showcases and early hands-on posts describe generated audio or synchronized sound, and one product listing calls it native stereo audio. However, the official H3 Max model table supplied here confirms T2V and I2V, not a separate audio-output specification. Audio should therefore be treated as demonstrated behavior rather than a fully documented API guarantee.
fal.ai has also shown an experimental long-form checkpoint with Native Infinite Continuity support. fal says this checkpoint holds a thread across scenes instead of resetting every clip and is planned for API access the following week. That feature is experimental and should not be confused with the standard 5–15-second H3 Max endpoint. The continuity announcement
What You Can Do With MiniMax H3 Max
H3 Max is best suited to workflows where generation latency affects the number of ideas a team can test.
- Storyboard iteration: Generate several motion directions from one prompt, then refine camera movement, lighting, or subject action.
- Image animation: Turn a still image into a moving shot, with an optional final frame to define the transition.
- Short advertisements: Produce product reveals, kinetic typography, or concise social ads in 5–15-second segments.
- Motion graphics: Third-party demonstrations show kinetic type, isometric 3D builds, technical schematics, collage, cut-out, and print-inspired looks.
- Livestream experiments: Early community testing connected near-real-time generation with perpetual streams and interactive worlds.
- Rapid visual prototyping: Production teams can compare more concepts before committing to a longer edit or a higher-resolution render.
The strongest practical case is fast exploration rather than final mastering. H3 Max's official ceiling is 768P, and the supplied evidence does not establish 2K or 4K output for this variant.
How MiniMax H3 Max Compares
The most relevant comparisons are with the original MiniMax H3 and Seedance 2.5. The figures below combine official documentation with individual community tests, so they should not be read as a controlled benchmark.
| Model | Generation scope | Resolution and duration evidence | Speed evidence |
|---|---|---|---|
| MiniMax H3 Max | T2V and I2V | 480P or 768P; 5–15 seconds | Community tests measured 15-second 768P clips in about 15–18 seconds |
| MiniMax H3 | T2V, I2V, reference, and broader multimodal workflows | 768P or 2K; 4–15 seconds | A clean H3 Max speed comparison is not yet confirmed |
| Seedance 2.5 | Not yet confirmed in the supplied model specification | One same-prompt test used 15 seconds at 1080P | The test measured 4 minutes 56 seconds, compared with 18 seconds for H3 Max |
One comparison reported H3 Max producing a 15-second 768P clip in 18 seconds, while Seedance 2.5 produced a 15-second 1080P clip in 4 minutes 56 seconds. The resolutions and systems differed, so the result demonstrates a latency difference in that test, not a universal quality ranking. A separate analysis of what is Seedance 2.5 and its pricing signals provides additional context for that competitor.
The clearest evidence for H3 Max's advantage is generation latency, while comparative quality remains less settled.
Availability: How to Access MiniMax H3 Max
The official MiniMax documentation lists MiniMax-H3-Max in its video-generation guide and directs users toward Pay-as-you-go API access. The API workflow is asynchronous: a client creates a generation task, checks its status, and retrieves the resulting video. The documentation lists H3 Max for T2V and I2V, while reference generation is marked as coming soon.
fal.ai announced the model and offered web-based generations, including an early free-access promotion for T2V and I2V. Third-party creative applications also displayed H3 Max, but the most reliable access path in the supplied evidence is the official MiniMax API documentation or fal.ai's own launch surface.
For a comparable multimodal video workflow, developers can inspect MiniMax H3 (Hailuo-03), which is the base model listed in the Kie.ai catalog. That is a reference to the related model, not a claim that Kie.ai hosts H3 Max.
The current official list price is $0.05 per second at 480P and $0.08 per second at 768P. Early community posts described temporary rates of $0.025 per second and $0.04 per second, plus free daily generations, but those promotional conditions are not the current official list prices.
What We Don't Know Yet
Several important engineering details remain open:
- H3 Max's parameter count, architecture, training data, and post-training recipe have not been published in the supplied sources.
- MiniMax has confirmed that H3 is open-weight, but it has not explicitly confirmed that H3 Max weights are available.
- A single community account reported that H3 Max weights would be released. That claim remains unconfirmed.
- Native audio generation appears in demonstrations and third-party descriptions, but the official H3 Max API specification supplied here does not define it as a separate output mode.
- The reported 36× to nearly 50× speed advantage over base H3 is not a controlled benchmark. Some comparisons used different hardware, endpoints, or production stacks.
- Comparative quality, temporal consistency, human fidelity, and audio quality need broader standardized evaluation.
- The experimental long-form continuity checkpoint had not yet entered the standard API in the latest supplied update.
Frequently Asked Questions
What is MiniMax H3 Max?
MiniMax H3 Max is a video generation model jointly released by MiniMax and fal.ai, post-trained by fal.ai from MiniMax H3 open weights and optimized for fast text-to-video and image-to-video generation. It produces 5–15-second clips at 480P or 768P through supported API and web access.
Is MiniMax H3 Max open source?
MiniMax H3 Max is not confirmed to be open source or open-weight. MiniMax officially describes the underlying MiniMax H3 as open-weight, while its announcement does not state that fal.ai's H3 Max post-trained variant has released weights.
How much does MiniMax H3 Max cost?
MiniMax H3 Max costs $0.05 per second at 480P and $0.08 per second at 768P under the official MiniMax Pay-as-you-go pricing listed in the supplied documentation. A 15-second clip therefore costs $0.75 at 480P or $1.20 at 768P before discounts or account-specific terms.
How fast is MiniMax H3 Max?
MiniMax H3 Max is designed for faster-than-real-time generation, according to MiniMax and fal.ai. Early community tests measured 15-second 768P clips in approximately 15–18 seconds and 15-second 480P clips in approximately 4 seconds, but these are not standardized benchmarks.
What can MiniMax H3 Max generate?
MiniMax H3 Max generates short videos from text prompts or an image, with optional first-frame and last-frame control. The official model documentation lists text-to-video and image-to-video as supported modes, with 5–15-second output at 480P or 768P.
MiniMax H3 Max vs MiniMax H3?
MiniMax H3 Max is fal.ai's post-trained and speed-optimized variant of MiniMax H3, while MiniMax H3 is the original open-weight multimodal video model from MiniMax. H3 supports broader text, image, video, and audio workflows, whereas H3 Max officially lists text-to-video and image-to-video and generates faster.
What to Watch Next
The next useful signals are an official decision on H3 Max weight release, a stable API specification for audio and reference generation, and controlled tests that separate checkpoint quality from infrastructure speed. The experimental long-form continuity checkpoint also deserves tracking once its API availability is documented.
Building similar near-real-time video workflows? On kie.ai you can try MiniMax H3 (Hailuo-03), Wan 3.0 Video, and Kling O3.
About Lukas Vogel
Lukas reads the papers and model cards so you do not have to, focusing on reproducible claims.
View all posts by Lukas Vogel