MiniMax H3: Model Introduction & Practical Guide
MiniMax H3 is MiniMax's open-weight video generation model - and the community reaction told the story faster than any spec sheet: within days of release, tutorials for running it on 8GB of VRAM were circulating alongside ComfyUI integrations on day zero. H3 generates video with synchronized stereo audio in a single pass, up to 15 seconds, at API output up to 2K.
Here's the short version: H3 is arguably the strongest open-weight video model available, and it comes with two variants that cover the full production surface - FL2VA (text-to-video with optional first/last frame control) and Ref2VA (full reference: up to 9 images, 3 videos, and 3 audio clips for character consistency, video editing, motion transfer, and shot continuation). The community license is royalty-free for commercial use with two conditions worth noting: companies above $20M annual revenue need separate written authorization, and commercial products must display "MiniMax H3" attribution. Open weights are currently unavailable for download in the EU, UK, US, and Korea (regulatory caution on generative video), while the API is globally available. At $0.036-0.061 per second through MiniMax's platform, the hosted pricing is aggressive.
This guide covers model overview, core features, technical specifications, capability comparison, core advantages, recommended use cases, example prompts, and selection recommendations.
Quick Facts
| Attribute | Value |
|---|---|
| Model name | MiniMax H3 |
| Developer | MiniMax |
| Category | Open-weight video generation model (with audio) |
| Output | Video + synchronized stereo audio, up to 15s |
| Resolution | API up to 2K; open weights native 768p short-edge |
| Frame rate | 24 fps |
| Variants | FL2VA (text-to-video + first/last frame), Ref2VA (full reference) |
| Reference inputs (Ref2VA) | Up to 9 images + 3 videos + 3 audio clips |
| Local deployment | ComfyUI day-0 support; runs from ~8GB VRAM |
| License | MiniMax H3 Community License - royalty-free commercial with conditions |
| API pricing | 2K video $0.061/sec; 768p $0.036/sec |
| Weight availability | Global except EU / UK / US / Korea (API available worldwide) |
Table of Contents
- Model Overview
- Core Features
- Technical Specifications
- Capability Comparison
- Core Advantages
- Recommended Use Cases
- Example Prompts
- Selection Recommendations
- FAQ
- Sources & Further Reading
1. Model Overview
The open-weight video race has a new benchmark, and H3 set it. MiniMax released the weights under a community license, and the response from the ecosystem was immediate - ComfyUI shipped day-zero support with workflow templates, Diffusers and DiffSynth-Studio pipelines followed, and the community quickly demonstrated deployment from roughly 8GB of VRAM. Partners ranging from ComfyUI's product lead to the model API's ecosystem lead publicly framed the release as a watershed for open video models.
What makes H3 significant is not just openness but completeness. Most video models pick a lane; H3 ships two variants that cover the production surface. FL2VA handles text-to-video with optional first-frame and last-frame conditioning - the standard creative workflow. Ref2VA is the full-reference workhorse: feed it up to 9 images, 3 video clips, and 3 audio clips, and it can maintain character consistency, perform video editing, transfer motion, or continue shots. That reference depth is unusual even among closed models, and unique in the open ecosystem.
Audio is co-generated: H3 produces video with synchronized stereo sound in the same pass, at up to 15 seconds and 24 fps - long enough for a complete social clip, and with sound already attached.
The licensing terms are the one area requiring attention. The community license is royalty-free for research and commercial use, with three specifics: entities above $20M annual revenue need separate written authorization; commercial products must display "MiniMax H3" prominently in their interface; and outputs/weights cannot be used to train other non-H3 AI models. Open weights are also region-restricted (not downloadable in the EU, UK, US, or Korea while those jurisdictions' generative-video rules evolve) - the API, however, is globally available with built-in safety measures, and institutions in restricted regions can apply for formal authorization.
2. Core Features
Synchronized stereo audio. Video and stereo sound are generated together in a single pass - no separate audio pipeline for dialogue, effects, or ambience.
Up to 15 seconds at 24 fps. Complete short-form clips in one generation; API output up to 2K resolution.
FL2VA variant. Text-to-video with optional first-frame and last-frame control - precise creative anchoring for transitions and scene endpoints.
Ref2VA variant. Full-reference generation: up to 9 images + 3 videos + 3 audio clips for character consistency, video editing, motion reference, and shot continuation.
Open weights with day-zero tooling. ComfyUI (with bundled templates), Diffusers, DiffSynth-Studio, and WanGP for low-VRAM machines.
Local deployment from ~8GB VRAM. Community optimizations and acceleration plugins (up to ~45% speedups reported) make it accessible on consumer hardware.
Aggressive hosted pricing. $0.061/second for 2K and $0.036/second for 768p through MiniMax's platform, with a Hailuo AI app surface for direct use.
3. Technical Specifications
| Specification | Detail |
|---|---|
| Output | Video + synchronized stereo audio |
| Max duration | 15 seconds |
| Frame rate | 24 fps |
| Resolution | API: up to 2K · Open weights: native 768p short edge |
| Variants | FL2VA (T2V + first/last frame), Ref2VA (references: 9 images / 3 videos / 3 audio) |
| Capabilities | Character consistency, video editing, motion reference, shot continuation |
| Local runtime | ComfyUI (day-0), Diffusers, DiffSynth-Studio, WanGP |
| Minimum VRAM (community builds) | ~8GB |
| API pricing | $0.061/sec (2K), $0.036/sec (768p) |
| License | Community license - royalty-free commercial; >$20M revenue needs authorization; attribution required |
| Region note | Open weights unavailable in EU/UK/US/KR (temporary); API global |
4. Capability Comparison
| Dimension | MiniMax H3 | Wan 3.0 | Veo 3.1 |
|---|---|---|---|
| Weights | Open (community license) | Closed | Closed |
| Max duration | 15s | 30s | 8s (extendable) |
| Max resolution | 2K (API) | 1080p | 4K |
| Audio | Synchronized stereo, co-generated | Yes | Native |
| Reference depth | 9 images + 3 videos + 3 audio | Multimodal input | Up to 3 images |
| Local deployment | Yes (~8GB VRAM) | No | No |
| Price posture | $0.036-0.061/sec | ¥0.3-1.2/sec | ~$0.15-0.40/sec |
Where H3 wins: it is the only model in this comparison you can run and fine-tune yourself; its reference system (Ref2VA) is the deepest available; and its price undercuts premium closed tiers by an order of magnitude at 768p. Where it trails: maximum resolution (2K API vs Veo's 4K), single-pass duration (15s vs Wan's 30s), and - for some companies - the license's revenue threshold and attribution requirements.
5. Core Advantages
- Full production surface in open weights. Both text-to-video and deep-reference workflows, self-hostable - no other release combines these.
- Audio included. Co-generated stereo sound removes the post-production sync work entirely.
- Accessible local deployment. ~8GB VRAM entry with strong community tooling lowers the floor from data center to desktop.
- Deep reference control. 9 images + 3 videos + 3 audio inputs enable character-consistent series, edits, and continuations that competitors can't match.
- Favorable licensing for most users. Royalty-free commercial use below the revenue threshold, with clear attribution terms.
- Two-tier economics. Free local generation for volume; cheap global API ($0.036/sec at 768p) for convenience and 2K output.
6. Recommended Use Cases
- Short drama and social video: 15-second clips with synchronized dialogue/effects, generated at scale locally.
- Character-consistent series: Ref2VA's multi-image references for recurring characters across episodes and campaigns.
- Video editing and continuation: modify or extend existing footage with reference-guided regeneration.
- Motion transfer: drive a target character with reference video motion.
- Local content pipelines: studios with GPU infrastructure producing without per-second API costs.
- Product and marketing video: fast, audio-complete clips for campaigns and e-commerce.
7. Example Prompts
1. Text-to-video with audio (FL2VA)
2. First/last frame transition (FL2VA)
3. Character consistency (Ref2VA)
4. Video editing (Ref2VA)
5. Shot continuation
8. Selection Recommendations
Choose MiniMax H3 if:
- You want a strong video model you can run, modify, and fine-tune locally.
- Your workflows need deep reference control (character consistency, edits, motion transfer).
- You need audio synchronized with video without a separate pipeline.
- You are cost-sensitive at volume: local generation is effectively free, and the API is priced 4-10x below premium closed tiers.
- You can comply with the community license terms (attribution; revenue threshold).
Check constraints first if:
- You are located in or deploying for the EU/UK/US/Korea markets - weights are currently unavailable there (API subscription or formal authorization are the paths).
- Your company exceeds $20M annual revenue and you plan commercial deployment - separate written authorization is required.
- You need 4K output or longer than 15-second single-pass generation.
Sources & Further Reading
- MiniMax H3 open-source hub - MiniMax Design (official)
- MiniMax official platform
- MiniMax international site
- MiniMax audio/product pricing page
License terms, regional availability, and pricing are as published at review time and are controlled by MiniMax - review the current community license before commercial deployment. Local VRAM requirements depend on quantized community builds and workflow settings.






