MiniMax H3 just landed at the top of the Artificial Analysis video leaderboard for editing, and placed top 3 in both text-to-video and image-to-video. It is the successor to the Hailuo family, and it brings a genuinely different architecture to the table: a single model that reads text, images, video clips, and audio, and outputs video with synchronized sound in one pass. MiniMax is also planning to release the weights publicly, which would make it the strongest open-weights video model by a wide margin.

What H3 actually is

MiniMax H3, also called Hailuo 3.0, is the latest flagship video model from MiniMax, the lab behind Hailuo 02 and Hailuo 2.3. The model can generate videos of up to 15 seconds in 2K resolution with native stereo sound. That is a meaningful jump from Hailuo 2.3, which topped out at roughly 10 seconds at 1080p.

H3 upgrades Hailuo's video line across the board: native 2K vs 1080p, 5-15 second clips extendable to roughly 30 seconds via an Extend tool, plus two brand-new capabilities Hailuo 2.3 lacked: omni-reference (up to 9 images, 3 video clips, 3 audio clips) and instruction-based editing, alongside one-pass synchronized audio.

The two features that actually matter

The headline numbers are nice, but the real story is in two capabilities that most video models still cannot do well.

  • Omni-reference: You can supply up to 9 reference images, up to three reference video clips, and up to three reference audio clips in one generation, and the model uses them for style, character, motion, or voice guidance without forcing them as keyframes. This is the fix for character identity drift, one of the most painful problems in multi-shot AI video production.