Tuesday, August 4, 2026

Today’s Edition

AI Intel Report

MARKETS

Frontier Models

MiniMax H3 Launches Open Multimodal Video Model at Sub-Third Pricing

Shanghai-based MiniMax introduces a 33B-parameter model that accepts text, image, video and audio inputs to produce up to 15-second 2K clips with native sound, directly challenging closed systems from ByteDance and Kuaishou.

3 MIN READ
Inside a sleek open-plan technology research facility located in Shanghai a group of anonymous engineers and technicians work at long rows of modern workstations equipped with multiple large flat-panel monitors displaying abstract colorful video frames sequences of motion graphics and waveform patterns representing audio integration all without any visible text numbers or logos. One engineer sits facing away from the viewer manipulating a keyboard and mouse connected to a high-performance computing rig with visible server racks humming in the background their posture focused on generating short video clips from combined text image and audio inputs. Nearby another technician reviews output clips on a separate monitor showing realistic 2K resolution footage of everyday scenes like urban streets and natural landscapes with synchronized sound indicators visualized through dynamic line graphs. The environment features contemporary Chinese office architecture with large windows overlooking the Shanghai skyline at dusk soft ambient lighting from overhead panels and subtle reflections on polished surfaces. Additional details include scattered external hard drives USB connected peripherals tablet devices showing static multimodal input examples such as photographs of objects and short video thumbnails all arranged neatly on desks alongside ergonomic chairs and cable management systems. The scene captures the collaborative atmosphere of advancing open-source artificial intelligence tools for video synthesis that directly compete with proprietary systems from major industry players emphasizing hardware infrastructure and human oversight in a real-world professional setting. Engineers wear standard business casual attire with no distinguishing features or faces visible to maintain anonymity. The overall composition includes layered elements like potted plants for environmental context reflective glass partitions separating work zones and background activity from other staff members engaged in similar tasks involving data processing and model testing all grounded in the practical deployment of a large-scale parameter model capable of handling diverse input modalities to output concise video segments with integrated audio. This realistic live-action view highlights the technical workspace environment associated with innovative AI development in China without any artistic embellishments or non-photographic elements ensuring a truthful representation of the news story about the launch of advanced multimodal video generation technology.
Illustration: AI Intel Report

MiniMax H3 is a 33B-parameter dense single-stream Transformer that accepts unified multimodal inputs of text, images, video, and audio to generate videos up to 15 seconds long at up to 2K resolution with native stereo sound.

MiniMax launched the H3 model on July 31, 2026.

The Shanghai-based company described H3 as a general-purpose omni-modal generation model.

H3 accepts text, images, video and audio inputs.

Background on the AI Video Generation Market

Chinese AI firm MiniMax released a new video-generation model on Friday that can process text, images, video and audio, stepping up competition in a fast-growing market led by rivals ByteDance and Kuaishou.

Closed models from ByteDance and Kuaishou have led proprietary video tools until now.

MiniMax positions H3 as an open-weights alternative that undercuts pricing.

Key Features of the MiniMax H3 Release

H3 generates videos up to 15 seconds long at up to 2K resolution with native stereo sound at 24 FPS and 32 kHz audio.

Supported output durations range from 4 to 15 seconds across various aspect ratios.

The model supports stable dialogue in 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish.

Technical Architecture and Capabilities

H3 is a 33B-parameter dense single-stream Transformer.

This architecture processes unified context across text, images, video, and audio inputs.

Early testing shows H3 is ready for commercial content creation across a wide range of use cases.

MiniMax H3 Technical Specifications
SpecificationDetail
Parameter Count33 billion
Max Video Duration15 seconds
Max Resolution2K
Frame Rate24 FPS
Audio Sample Rate32 kHz
Supported Languages11
Input ModalitiesText, Image, Video, Audio
OutputVideo with native stereo sound
  1. Instruction following
  2. Accurate text and brand rendering
  3. V2V motion transfer

Pricing Strategy and Cost Comparison

At 2K resolution, H3 per-second price is less than one-third of mainstream models.

At 768p the price is less than half of mainstream 720p models.

Market and Stakeholder Implications

The open-weights approach allows users to download and customize the underlying system.

This accelerates the open versus proprietary AI video contest.

Content creators gain access to native multimodal audio and video generation at lower cost.

Company Statements on the Launch

Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution.MiniMax

MiniMax noted that early testing confirms readiness for commercial use cases.

Planned Open-Weights Release and Next Steps

MiniMax plans to release H3 model weights in the coming days, subject to regulations.

The release will enable further customization by developers and enterprises.

Future updates may extend duration and resolution based on user feedback.

Frequently asked

What inputs does MiniMax H3 accept?

MiniMax H3 accepts unified multimodal inputs of text, images, video, and audio.

When will MiniMax release the H3 model weights?

MiniMax plans to release H3 model weights in the coming days, subject to regulations.

Sources

  1. MiniMax — H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution. At 2K resolution, H3's per-second price is less than a third of mainstream models (and at 768p less than half of mainstream 720p models). Early testing shows H3 is ready for commercial content creation across a wide range of use cases, excelling at instruction following, accurate text and brand rendering, and V2V motion transfer.
  2. Reuters — Chinese AI firm MiniMax released a new video-generation model on Friday that can process text, images, video and audio, stepping up competition in a fast-growing market led by rivals ByteDance and Kuaishou. The Shanghai-based company said its H3 model could generate videos of up to 15 seconds in 2K resolution with native stereo sound. MiniMax said it planned to release H3's model weights within days, allowing users to download and customise the underlying system.