MiniMax H3 is a general-purpose multimodal video model that produces 4 to 15 second clips in 2K. Images, video, and audio go into a single request together — H3 reads the characters, motion, emotion, camera language, and style across all of them, then fuses them into one coherent scene with sound generated alongside the picture. Editing works the same way: point at a character, object, background, or line of dialogue and describe the change, and everything else stays intact. Built for advertising, e-commerce, brand, short drama, and game content. Ranked first for video editing on Artificial Analysis.