Added Jul 31, 2026
MiniMax H3, released 31 July 2026, reads text, images, video and audio as one shared context and returns video with native stereo sound at up to 2K and 15 seconds. MiniMax positions it as omni-modal rather than a video model with features bolted on: text-to-video, image-to-video, editing, reference-based and audio-driven generation are all expressed as natural-language instructions over that one context instead of separate expert models. A single generation accepts up to 9 reference images, 3 reference video clips and 3 reference audio clips, 12 files in total. It runs on the MiniMax platform API as model ID MiniMax-H3 and in the consumer Hailuo app.
Why: The native stereo audio is the part that matters — most 2K video models still hand you a silent clip and leave you to score it. Generating sound in the same pass as the picture removes a whole step, and MiniMax prices the 2K tier well under the mainstream alternatives.