Chinese artificial intelligence company MiniMax has officially launched its next-generation flagship multimodal model, MiniMax H3.
The company has committed to fully open-sourcing the model weights within a few days.
>>> Ceuta Border Crisis: 18 Dead as Thousands Attempt to Enter Spain
The release follows a teaser on X, where MiniMax posted a short video with the phrase "the world is multimodal" and tagged its Hailoo AI assistant platform.
It also comes shortly after the launch of MiniMax Speech 2.8, an updated voice-generation model for natural conversational audio.
Unified Multimodal Capabilities
MiniMax H3 is a unified generative model that accepts multimodal context inputs.
Users can combine text, images, video, and audio with natural language instructions to generate high-definition video output with native stereo sound.
>>> Singapore Probes Massive Attack Over Palestine Flag Display
To power the platform, MiniMax developed the H3-Omni Transformer architecture. This architecture merges text-to-image, text-to-video, image-to-video, and audio tasks into a single pre-training framework.
The design increases end-to-end training throughput by nearly 30 percent. It also utilizes an updated H3-VAE Tokenizer architecture to boost compression rates and lower high-resolution inference costs.
Cost Efficiency and Deployment
At 2K resolution, MiniMax H3 operates at less than one-third the per-second cost of comparable mainstream models. It maintains compatibility with domestic chips.
>>> UK Drought Emergency Declared as Heat and Dry Spell Worsen
The system targets commercial deployment in movie openers, game UI animations, dynamic posters, and e-commerce advertising.