MiniMax H3 Max Reference-to-Video (Ref): How It Works
Updated 2026-10-06
"MiniMax H3 Max ref" means the reference-to-video mode of MiniMax H3 Max: instead of only a text prompt or one start image, you give the model reference files – for example a character photo and a product shot – and it generates a new clip that keeps those subjects consistent in a scene neither reference contains. H3 Max is fal Research's post-trained version of MiniMax's open-weight H3 video model. It outputs 480p or 768p clips of 5–15 seconds at 24 fps with native audio, and its reference endpoint was added on fal after the model's launch.
What reference-to-video does
Text-to-video invents everything from words. Image-to-video animates one picture, usually as the first frame. Reference-to-video sits in between: the references tell the model who or what must appear, and the prompt tells it what happens and where. Typical uses:
- Keeping the same character across several clips for a short story or ad.
- Placing a product into a new setting without shooting it there.
- Combining two subjects – say a person and a pet – in one shot.
In fal's own comparison of H3 and H3 Max, the reference test used two image references dropped into a location that neither image showed.
MiniMax H3 Max ref: key specs
| MiniMax H3 Max | MiniMax H3 | |
|---|---|---|
| Built by | fal Research, post-trained from H3 | MiniMax (open weights) |
| Modes on fal | Text, image and reference to video | Text, image and reference to video |
| Resolution | 480p, 768p (both native) | 480p, 768p native; 2K and 4K upscaled |
| Length | 5–15 seconds, 24 fps | 5–15 seconds, 24 fps |
| Reference files | Up to 12 in total, no per-type caps published | Up to 12: max 9 images, 3 videos, 3 audio |
| Audio | Native audio and lip-sync | Native audio and lip-sync |
| Price on fal | $0.05/s at 480p, $0.08/s at 768p | $0.05/s at 480p, $0.06/s at 768p, more for 2K/4K |
Specs and prices are from fal's published comparison (August 31, 2026). Check the model page before large jobs, since endpoints and prices change.
How to use H3 Max reference-to-video
- Pick clean references. One subject per image, well lit, with the face or product clearly visible. Remove cluttered backgrounds when you can.
- Refer to each file in the prompt. If the interface numbers uploads, say which is which – "the woman from image 1 holds the bottle from image 2".
- Describe one continuous shot. Subject, action in order, camera move, setting and light, then sound. Timed beats ("0–2 s …, 2–5 s …") help the model keep the order.
- Start at 480p and 5 seconds. It costs about $0.25 on fal at 480p, so test motion and likeness cheaply before rendering at 768p.
- Check consistency. Compare the face, logo or product shape against your references, and regenerate rather than accept drift.
H3 Max vs H3 for references
Choose H3 Max when you want fast turnaround at 768p and strong prompt adherence – fal says it post-trained the model for that. Choose plain H3 when you need 2K or 4K delivery, or a larger mixed pack of references with video and audio files, since H3 publishes per-type limits for those. At 768p, H3 is the cheaper of the two on fal.
Try MiniMax H3 online
Our MiniMax H3 video generator runs H3 Max text-to-video and image-to-video in the browser, so you can test the model's motion and prompt adherence before paying for reference runs elsewhere. Multi-file reference upload isn't available on our page yet.
FAQ
What does "ref" mean in MiniMax H3 Max ref?
It's short for reference-to-video: generating a clip guided by reference files (images, and on H3 also video and audio) plus a text prompt.
How many references can MiniMax H3 Max use?
fal lists up to 12 reference files in total for H3 Max without per-type caps. Plain H3 lists up to 12, with at most 9 images, 3 videos and 3 audio files.
How much does H3 Max reference-to-video cost?
On fal, H3 Max costs $0.05 per second at 480p and $0.08 per second at 768p. A 10-second 768p clip is about $0.80.
Can MiniMax H3 Max make 2K or 4K video?
No. H3 Max outputs 480p or 768p. For 2K or 4K, use plain MiniMax H3, which upscales from a 768p render.
Is H3 Max an official MiniMax model?
It was trained by fal Research on top of MiniMax's open-weight H3 base and is served on fal. MiniMax makes the base H3 model.
GenModelHub is independent and not affiliated with MiniMax or fal.
Try the MiniMax H3 AI Video Generator
Run the MiniMax H3 video model online – from a text prompt, in about a minute.
Generate a video