Frontier Models
MiniMax H3 Launches Open Multimodal Video Model at Sub-Third Pricing
Shanghai-based MiniMax introduces a 33B-parameter model that accepts text, image, video and audio inputs to produce up to 15-second 2K clips with native sound, directly challenging closed systems from ByteDance and Kuaishou.
MiniMax H3 is a 33B-parameter dense single-stream Transformer that accepts unified multimodal inputs of text, images, video, and audio to generate videos up to 15 seconds long at up to 2K resolution with native stereo sound.
MiniMax launched the H3 model on July 31, 2026.
The Shanghai-based company described H3 as a general-purpose omni-modal generation model.
H3 accepts text, images, video and audio inputs.
Background on the AI Video Generation Market
Chinese AI firm MiniMax released a new video-generation model on Friday that can process text, images, video and audio, stepping up competition in a fast-growing market led by rivals ByteDance and Kuaishou.
Closed models from ByteDance and Kuaishou have led proprietary video tools until now.
MiniMax positions H3 as an open-weights alternative that undercuts pricing.
Key Features of the MiniMax H3 Release
H3 generates videos up to 15 seconds long at up to 2K resolution with native stereo sound at 24 FPS and 32 kHz audio.
Supported output durations range from 4 to 15 seconds across various aspect ratios.
The model supports stable dialogue in 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish.
Technical Architecture and Capabilities
H3 is a 33B-parameter dense single-stream Transformer.
This architecture processes unified context across text, images, video, and audio inputs.
Early testing shows H3 is ready for commercial content creation across a wide range of use cases.
| Specification | Detail |
|---|---|
| Parameter Count | 33 billion |
| Max Video Duration | 15 seconds |
| Max Resolution | 2K |
| Frame Rate | 24 FPS |
| Audio Sample Rate | 32 kHz |
| Supported Languages | 11 |
| Input Modalities | Text, Image, Video, Audio |
| Output | Video with native stereo sound |
- Instruction following
- Accurate text and brand rendering
- V2V motion transfer
Pricing Strategy and Cost Comparison
At 2K resolution, H3 per-second price is less than one-third of mainstream models.
At 768p the price is less than half of mainstream 720p models.
Market and Stakeholder Implications
The open-weights approach allows users to download and customize the underlying system.
This accelerates the open versus proprietary AI video contest.
Content creators gain access to native multimodal audio and video generation at lower cost.
Company Statements on the Launch
Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution.MiniMax
MiniMax noted that early testing confirms readiness for commercial use cases.
Planned Open-Weights Release and Next Steps
MiniMax plans to release H3 model weights in the coming days, subject to regulations.
The release will enable further customization by developers and enterprises.
Future updates may extend duration and resolution based on user feedback.
Frequently asked
What inputs does MiniMax H3 accept?
MiniMax H3 accepts unified multimodal inputs of text, images, video, and audio.
When will MiniMax release the H3 model weights?
MiniMax plans to release H3 model weights in the coming days, subject to regulations.
Sources
- MiniMax — H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution. At 2K resolution, H3's per-second price is less than a third of mainstream models (and at 768p less than half of mainstream 720p models). Early testing shows H3 is ready for commercial content creation across a wide range of use cases, excelling at instruction following, accurate text and brand rendering, and V2V motion transfer.
- Reuters — Chinese AI firm MiniMax released a new video-generation model on Friday that can process text, images, video and audio, stepping up competition in a fast-growing market led by rivals ByteDance and Kuaishou. The Shanghai-based company said its H3 model could generate videos of up to 15 seconds in 2K resolution with native stereo sound. MiniMax said it planned to release H3's model weights within days, allowing users to download and customise the underlying system.