ArtArch Newsroom
AcademyWhat Is MiniMax H3? A Guide to the Multimodal Video Model
Learn how MiniMax H3 combines text, images, video, and audio for 5 to 15 second video generation, reference control, and precise editing.

What is MiniMax H3?
MiniMax H3 is a multimodal video model that reads text, images, video, and audio as one creative context. It can generate a new audiovisual clip or edit elements of an existing one. The supplied launch manual positions H3 as a move away from separate tools for generation, reference, and editing toward a single model that understands the relationship between subject, action, camera, sound, style, and intent.
That description is easiest to understand through a practical example. A fashion brief may contain a location reference, two talent sheets, a product image, a logo, and a sample edit. H3 can use those materials together, provided the prompt explains what each one controls. The result is a 5 to 15 second video at 24 FPS with native stereo sound.
Supplied H3 demonstration: a completed result created through the broader reference workflow.
Two creation modes cover different jobs
First and last frame mode is the direct route. It accepts no image for text to video, one image for a guided start, or two images to define both ends of the shot. When images are used, the output follows their original aspect ratio within the supported range.
Omni Reference is for briefs with several sources. It accepts up to nine images, three video clips, and three audio clips, with no more than 12 files in total. Images can define people, products, settings, or style. Video can define motion, camera work, editing rhythm, or an existing scene to modify. Audio can contribute voice tone, dialogue, music, or a timing reference, but it must accompany an image or video.
H3 generates and edits
The launch examples cover brand films, trailers, vertical drama, ecommerce, game interfaces, website motion, animation, and character videos. They also demonstrate targeted edits such as replacing a cat with a dog, adding a person to a group, changing a green-screen background, relighting a scene at night, and replacing spoken dialogue.
These examples show the kinds of instructions H3 can receive. They do not establish a guaranteed result for every input. A useful test separates what must change from what must stay fixed, then checks both after generation.
Output specifications that affect planning
H3 supports 5 to 15 second outputs at 24 FPS. Text-to-video and Omni Reference requests can use 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16. Omni Reference also has an automatic aspect-ratio option.
The manual describes a 1440p mode. From 16:9 through 9:16, the short side is 1,440 pixels. Wider formats use roughly 3.7 million pixels; the listed 21:9 example is 2976 by 1248. A 768p mode is marked for later availability.
The prompt can contain up to 7,000 characters. That capacity helps with complex briefs, but a long prompt still needs a visible hierarchy. Put identity, action, camera, sound, protected details, and final-frame requirements in separate sentences.
Keep exploring
More in Academy

MiniMax H3 Production Checklist for Commercial Video Teams
Review MiniMax H3 inputs, rights, prompt scope, identity, product accuracy, audio, format, retries, approvals, and delivery before production.

MiniMax H3 API Media Formats and URL Input Checklist
Prepare MiniMax H3 API media with supported codecs, image and audio formats, file limits, 64 MB request bodies, and URL-based inputs.

How to Plan 12 Reference Files for MiniMax H3
Assign clear identity, product, style, motion, camera, and audio roles across a MiniMax H3 mixed request without creating contradictory references.

Use MiniMax H3 for Storyboard and Visual Pitch Previews
Turn keyframes and a creative brief into a MiniMax H3 motion preview for shot order, transitions, timing, sound, and stakeholder review.

