MiniMax H3 Max: Model Introduction & Practical Guide
MiniMax H3 Max is the maximum-fidelity configuration of MiniMax's H3 video family - the full three-stage pipeline that refines your instructions, generates the audio-video clip, and rebuilds it at native 2K instead of relying on an upscale.
Here's the short version: H3 itself is an open-weight audio-video model whose public weights generate at 768p short-edge. "Max" is the catalog label for the top-quality service tier, and in MiniMax's own documentation that tier is delivered by two platform stages: H3-Context-IR (which turns messy multimodal instructions into a clean representation) and H3-Regenerate-2K (which feeds the 768p result plus the original context back through H3 to regenerate the output at native 2K). That distinction matters: this is re-generation with full context, not pixel upscaling, and the model's own documentation treats it as the path to final masters. The practical workflow most teams land on is a hybrid - iterate locally for free at 768p, then pay for the Max pipeline only on the takes that survive review.
This guide covers model overview, core features, technical specifications, how the 2K pipeline works, capability comparison, core advantages, recommended use cases, example prompts, selection recommendations, and cost/deployment notes.
Quick Facts
| Attribute | Value |
|---|---|
| Name | MiniMax H3 Max (top-quality H3 configuration) |
| Developer | MiniMax |
| Category | Audio-video generation, native 2K output |
| Family | MiniMax H3 (released July 31, 2026) |
| Pipeline | H3-Context-IR → H3-Base → H3-Regenerate-2K |
| Output | Native 2K (API), up to 15 s, 24 fps, synchronized stereo audio |
| Open-weight baseline | 768p short edge (local, free) |
| Variants | FL2VA (text + first/last frame), Ref2VA (up to 9 images + 3 videos + 3 audio) |
| Access | MiniMax platform/API; Hailuo AI app; third-party providers |
| License model | Community license (H3 family terms apply) |
Table of Contents
- Model Overview
- Core Features
- Technical Specifications
- How the 2K Pipeline Works
- Capability Comparison
- Core Advantages
- Recommended Use Cases
- Example Prompts
- Selection Recommendations
- Cost & Deployment Notes
- FAQ
- Sources & Further Reading
1. Model Overview
MiniMax released H3 on July 31, 2026, and the launch's defining move was shipping open weights for a model that generates video with synchronized stereo audio in a single pass. The public weights run locally at 768p short-edge - which is where the community story lives: ComfyUI support on day zero, sub-10GB VRAM deployments, and a fast-growing LoRA ecosystem.
The platform tier is the other half of the story. MiniMax's own documentation describes the H3 system as three modules working in sequence, and two of them exist only server-side:
- H3-Context-IR - understands and refines the text/image/video/audio instruction set into a Context Intermediate Representation. MiniMax calls it critical to final output quality.
- H3-Base - generates the audio-video result at 768p from that representation.
- H3-Regenerate-2K - takes the 768p result together with the original context and regenerates the output at native 2K, recovering real detail rather than smoothing an upscale.
That full path - refined context in, native 2K master out - is what "H3 Max" refers to in catalog listings. The naming is worth stating plainly: at review time MiniMax's public materials describe the capability through the H3-Regenerate-2K stage rather than a separately named "H3 Max" checkpoint. What you get is the same H3 system at its highest quality configuration.
2. Core Features
Native 2K output. The regeneration stage produces 2K video from the 768p base plus the original context - detail is re-rendered, not interpolated.
Context-refined prompting. H3-Context-IR normalizes multi-reference instructions before generation, which is what makes complex character/product briefs repeatable.
Synchronized stereo audio. Video and stereo sound are generated in the same pass, so Max inherits lip sync, sound effects, and ambience rather than adding them in post.
Up to 15 seconds per generation. Complete short-form deliverables in a single pass at 24 fps.
Deep reference control. Ref2VA accepts up to 9 images, 3 video clips, and 3 audio clips (≤12 files), covering character identity, product appearance, motion reference, camera language, and music.
First/last frame control. FL2VA handles text-to-video with optional first-frame and last-frame conditioning for precise transitions.
Open-weight companion workflow. The same family runs locally at 768p for free - ideal for drafting, A/B testing, and LoRA customization before spending on 2K renders.
3. Technical Specifications
| Specification | Detail |
|---|---|
| Output resolution | Native 2K (platform pipeline) |
| Local open-weight baseline | 768p short edge |
| Duration | Up to 15 seconds per generation |
| Frame rate | 24 fps |
| Audio | Synchronized stereo, co-generated |
| Reference capacity (Ref2VA) | ≤9 images + ≤3 videos + ≤3 audio (≤12 files) |
| Reference clip length | 2-15 s per clip; ≤15 s total per modality |
| Control modes | FL2VA (first/last frame), Ref2VA (omni-reference) |
| Pipeline stages | H3-Context-IR → H3-Base → H3-Regenerate-2K |
| Deployment | MiniMax API / Hailuo AI; local 768p via open weights |
Note: "native 2K" comes from the regeneration stage in the platform pipeline. Do not expect the free local weights to output 2K - that is the paid tier's job, and it is the main functional difference between running H3 yourself and buying H3 Max.
4. How the 2K Pipeline Works
Three consequences follow from this design:
- Iteration belongs upstream. Because the base generation is cheap relative to regeneration, the economical loop is: draft locally (or at 768p on the platform), select, then regenerate only the winners.
- Context quality compounds. The same context representation that shaped the 768p take is reused for the 2K pass, which is why the refined brief matters more here than in single-shot pipelines.
- Regeneration is not upscaling. A traditional upscaler guesses missing detail; H3-Regenerate-2K re-runs generation with the original instruction set, so textures, faces, and fine structure are produced rather than amplified.
5. Capability Comparison
| Dimension | H3 Max (2K pipeline) | H3 local (open weights) | Seedance 2.5 | Veo 3.1 (class) |
|---|---|---|---|---|
| Max resolution | Native 2K | 768p short edge | Platform-dependent (1080p-2K class) | Up to 4K (8 s clips) |
| Duration | 15 s | 15 s | 30 s + 2 extensions | ~8 s + extension |
| Audio | Synchronized stereo | Synchronized stereo | Joint audio-video | Native |
| Reference depth | 9 images / 3 videos / 3 audio | Same | 30 images / 10 videos / 10 audio | Up to 3 images |
| Open weights | Family weights only (768p) | Yes | No | No |
| Cost posture | Per-second API | Free (your hardware) | Per-generation credits | Premium per-second |
| Best for | Final masters | Drafts, LoRA, volume | Long narratives | 4K short shots |
Where H3 Max wins: it is the cheapest route to native 2K with synchronized stereo audio and deep reference control among the options compared here - and the same family runs free locally for everything upstream. Where it concedes: single-pass duration (15 s vs Seedance 2.5's 30 s) and maximum resolution against 4K-class competitors.
6. Core Advantages
- Regenerated 2K, not upscaled. Detail comes from a fresh generation pass with full context - the difference shows in faces, text, and fine texture.
- Refined instructions included. Context-IR is part of the pipeline, so the Max tier gets better prompt adherence than raw prompting on the same model.
- Studio-grade reference control. 12 reference files across images, video, and audio cover identity, motion, camera, and music in one generation.
- Audio-complete output. Stereo sound is generated with the video; no separate dubbing or sound-design pass for basic deliverables.
- A sane cost ladder. Free local 768p drafting plus paid 2K finishing means you only pay for takes you keep.
- Ecosystem leverage. The same family supports ComfyUI, Diffusers, SGLang/vLLM serving, and LoRA training - knowledge transfers between the free and paid tiers.
7. Recommended Use Cases
- Client deliverables: final 2K masters where a 768p draft is not acceptable.
- Brand and product films: label-accurate, identity-consistent product video with native audio.
- Character series: recurring characters held consistent through multi-image references across episodes.
- Advertising cutdowns: draft variants locally, finish only the approved edit at 2K.
- Post-production finishing: upgrading selected takes from a local H3 workflow without re-authoring the brief.
- Hybrid studio pipelines: local iteration plus paid 2K only when the shot earns it.
8. Example Prompts
1. Character-consistent 2K scene
2. Product master
3. First/last frame transition
Prompting guidance: write the brief the way Context-IR wants to read it - labeled sections for identity, scene, motion, camera, audio, and output spec. When you regenerate at 2K, resubmit the same brief so the context matches the approved 768p take.
9. Selection Recommendations
Choose H3 Max if:
- The output is a final master and 768p is not deliverable.
- Your shots depend on multi-reference consistency (characters, products, motion, music).
- You want synchronized stereo audio without a separate audio pipeline.
- You can accept 15-second single-pass clips.
Choose local H3 (open weights) if:
- You are drafting, exploring, or training LoRAs.
- Your deliverables are 768p-class or you have your own finishing pipeline.
Choose Seedance 2.5 if:
- You need 30-second single-pass narrative and timestamp-level editing.
Choose 4K-class models if:
- Resolution above 2K is a hard broadcast requirement.
10. Cost & Deployment Notes
The economics of H3 Max are best understood as a ladder, not a price:
| Path | What you pay | What you get |
|---|---|---|
| Local open weights | Hardware + electricity | 768p generation, unlimited iteration, LoRA training |
| Platform 768p generation | Per second of output | 768p with Context-IR refinement |
| Platform 768p → 2K regeneration | Per second of regeneration | Native 2K master from an approved take |
| Third-party providers | Provider rates (often discounted) | Same family, varying service quality |
Observed third-party promotional listings have quoted 2K output around ¥0.15/second, 768p around ¥0.09/second, and 768p→2K regeneration around ¥0.06/second, with reference media billed separately - treat those as market signals, not official rates. MiniMax's own pricing page governs the platform tier.
Two constraints to check before you commit:
- License terms. The H3 community license is royalty-free for research and commercial use, with specifics: entities above $20M annual revenue need separate written authorization; commercial products must display "MiniMax H3" attribution; outputs and weights cannot be used to train other non-H3 AI models.
- Regional availability. Open weights are currently unavailable for download in the EU, UK, US, and Korea while generative-video regulation evolves there. The API is globally available, and institutions in restricted regions can apply for formal authorization.
Sources & Further Reading
- MiniMax H3 open-source hub (three-stage system, licensing, FAQ) - MiniMax Design
- MiniMax-AI/MiniMax-H3 repository - GitHub
- MiniMax H3 on Hugging Face
- MiniMax API platform
- MiniMax pricing
- Hailuo AI H3 open ecosystem
Pipeline stages, licensing terms, and regional availability reflect MiniMax's published documentation as of September 2026. "H3 Max" is a catalog label for the top-quality configuration; confirm current stage names, pricing, and limits with MiniMax or your provider before production deployment.






