MiniMax Hailuo H3 Max Reference to Video

MiniMax Hailuo H3 Max builds a video from your reference material. Upload up to 12 images, video clips or audio files and describe the shot — the model keeps the characters, style and motion you supplied while generating up to 15 seconds of new footage.

📄 About MiniMax Hailuo H3 Max Reference to Video
Key Features
Up to 12 reference files - images, video clips and audio combined
Keeps character identity, style and product detail across shots
5 to 15 second output, selectable per generation
480P and 768P output with seven aspect ratios
Adaptive aspect ratio follows your reference material
Prompt expansion modes: balanced or cinematic quality
💡 Use Cases
Keeping one spokesperson or model consistent across a campaign
Animating a product while preserving its exact packaging and labels
Holding a character design steady between animation shots
Rebuilding a scene in a new style from a reference frame
Matching the motion of an existing clip with new subject matter
Producing vertical social cuts from landscape reference footage
🎯 Best For
🎯 Teams who need the same character, product or style to survive across multiple generated videos
👍 Pros
Reference-driven consistency without training a custom model
Accepts images, video and audio references in one request
Fifteen second maximum is longer than most reference-to-video models
Adaptive aspect ratio removes guesswork about framing
⚠️ Considerations
Reference video and audio are capped at 15 seconds combined
Maximum resolution is 768P, so it is not a finishing-quality render
Large reference images cost extra beyond the free token allowance
📚 How to Use MiniMax Hailuo H3 Max Reference to Video
1
Upload your reference images - the character, product or style you want kept
2
Optionally add a reference video whose motion should be echoed, or audio to set the pacing
3
Describe the shot: camera movement, action and lighting, not what things look like
4
Pick duration, resolution and aspect ratio, then generate
Frequently Asked Questions
Twelve in total, counted across images, videos and audio together. Reference videos and audio must each be between 2 and 15 seconds, and their combined length cannot exceed 15 seconds.
That is what the model is built for. Supplying a clear reference image of the character is far more reliable than describing them in the prompt, and it holds up across separate generations so you can build a sequence.
Balanced stays close to the wording you wrote. Quality lets the model add cinematic detail - lighting, camera language, atmosphere - which usually looks richer but drifts further from your exact instructions.
Start at 480P while you are iterating; it costs less per second. Switch to 768P once the composition and motion are right. The price scales with both resolution and duration.
Yes. Uploading a short audio clip lets the model match the pacing and rhythm of the generated motion to your track, which is useful for music-led social cuts.

More Video Generation Models