The workflow below splits every generation into three passes with one job each. It is assembled from MiniMax's official H3 documentation and examples, the published API price table, and community-tested workarounds. What it is not: a memoir. There are no personal war stories here, only the mechanics and the math.

Pass 1: Draft cheap, ask one question
The rule: each generation answers exactly one question. "Does this composition work?" is a question. "Does this composition work, is the lighting cinematic, and does the character look right" is three questions, and a failed result teaches nothing about which one broke.
Draft settings:
- MiniMax H3 at 768p, or Hailuo 2.3 Fast if the only question is motion
- 4–6 seconds. Temporal coherence is strongest early in a clip; if it breaks at 6 seconds it will break at 15
- Prompt optimizer OFF. The optimizer rewrites intent; drafts should be literal
- Audio ignored entirely at this stage
The draft prompt skeleton is deliberately minimal:
\`\`\`
[subject], [one action], [location], [time of day].
Static tripod shot. No camera movement.
\`\`\`
If the framing is wrong here, no amount of 2K rendering saves it. Re-roll the draft until the composition is right, then move on.
The budget logic: at $0.08/second, a 6-second draft costs $0.48. A 15-second 2K final costs $1.95. Every mistake caught in Pass 1 is a $1.95 lesson a $0.48 test would have taught.
Pass 2: Lock the identity
The classic failure mode: shot 1 has your character, shot 3 has her cousin. The fix is mechanical. Give the model one reference image per job, and write the identity into the prompt as a contract.

*The reference map.* Number assets in upload order and assign each exactly one responsibility:
\`\`\`
Overall mood, lighting and film look: reference image 1
Character: reference image 2
Product: reference image 3
End-card logo: reference image 4
\`\`\`
In the prompt: \`整体氛围参考图1;人物参考图2;产品参考图3;品牌logo参考图4。\`
*The identity-lock clause.* After the map, pin the attributes in words. Concrete beats poetic. The model cannot hold "ethereal vibe" across shots, but it holds a list:
\`\`\`
保持身份一致:高马尾、左耳银色耳坠、深青色立领长衫、腰间悬白玉扣,全程不变。
\`\`\`
*Three constraints worth treating as law:*
1. One asset, one job. Two images that both claim to define the same face will conflict, and the result is drift. MiniMax's own reference documentation warns about this.
2. Front-facing, well-lit reference images. The model copies what it can see clearly.
3. First/last-frame mode and reference mode are mutually exclusive on H3. One per generation, never mixed. This is a hard API constraint.
Test the lock with a 6-second clip and scrub the middle 2 seconds, where testers report coherence peaks. If the face holds there, it holds.
Pass 3: Render with one camera move
H3 takes natural-language camera direction. The vocabulary confirmed in MiniMax's official examples: dolly in/out, pan left/right, tilt up/down, tracking shot, crane up/down, orbit, handheld, aerial, rack focus. The most useful entry may be \`static tripod shot\`, which kills the micro-movement the model adds when the camera is unspecified.

*One move per shot.* "Slow push-in while panning right as she turns" is three moves. The model will pick one at random or blend them into mush. Write one. If the scene needs a second move, generate a second clip and cut, the way the industry has done for a century.
For multi-shot generations, use shot order rather than timestamps. This pattern comes directly from MiniMax's official H3 examples:
\`\`\`
镜头1:近景固定机位,她站在观景窗前,舰队灯光扫过侧脸。
镜头2:切至中景缓慢推近,最后一艘舰船跃迁,强光爆闪。
镜头3:切至特写,她闭眼再睁开,只有低频轰鸣与金属应力声。
\`\`\`
Shot-distance policy, also from the official examples: faces belong in close-up or medium shots. Wide shots are for backs, profiles, or empty environment only. Faces in wide shots are where identity goes to die.
*Constrain the soundscape.* Silence is not the default on H3. Leave audio undescribed and the model invents dialogue, sometimes gibberish — the community calls this the "H3 gibberish problem." For a speech-free clip, say so explicitly:
\`\`\`
声音只用环境风声,没有音乐、没有人声、保持无字幕。
\`\`\`
Fill every second of the soundscape or the model improvises. For intentional dialogue, keep a single language and use the format \`角色说话:'...'\`.
The full prompt, assembled
Every pass uses the same skeleton; Pass 3 fills every slot:
\`\`\`
[precise subject] → [action details] → [scene/environment] →
[lighting & color tone] → [camera movement] → [visual style] →
[image quality] → [constraints]
\`\`\`
A product-demo example, ready to adapt:
\`\`\`
整体氛围、场景和胶片质感参考图1;人物资产参考图2;产品资产参考图3;品牌 ending logo 参考图4。
镜头1:影棚中景缓慢推近,人物右手自然拿起产品并将标签朝向镜头。
镜头2:切至产品特写,标签清晰可读。
镜头3:切至品牌 logo 定版。
保持人物面部与产品包装一致,不生成额外字幕或水印。
\`\`\`
Five known failure modes and the documented fixes
1. *Character drift. Fix: use the identity-lock clause in reference mode, or use a first-frame pin in first-frame mode—never both. Keep clips to 4–6 seconds and verify against the middle 2 seconds.
2. Camera ignores direction. Fix: one move per shot. Cut between moves; never blend.
3. Gibberish audio. Fix: constrain the soundscape explicitly. Silence must be requested; it is not the default.
4. Reference conflicts. Fix: one job per asset. A few purposeful references beat filling all 9 image, 3 video and 3 audio slots.
5. Bad text, bad hands.* Fix: keep on-screen text minimal and actions simple. H3 recovers small text better than the 2.x line, but it is still the weakest link.
The meta-rule: failed generations still cost credits. Preflight the prompt — no contradictions, no overloaded shots — before rendering long.
What this desk does not do
It does not replace taste. The desk buys consistency and cost control; the idea still has to be yours. And H3 is not always the right tool: for shots led by body motion or physics, Hailuo 2.3 remains the flagship. A reasonable split is Pass 1 on 2.3, finals on H3.
The full-loop arithmetic at published pricing: three 6-second 768p drafts ($1.44), one 6-second lock test ($0.48), one 15-second 2K final ($1.95). Under $4 from blank prompt to finished clip, provided no pass is skipped.

How this tutorial was put together
Workflow mechanics, pricing, and camera vocabulary were verified against MiniMax's H3 launch materials (WAIC, July 31 2026), the Hugging Face MiniMax H3 model blog, MiniMax platform API documentation, and community references. Pricing is MiniMax's published API pricing as of October 2026. Consumer-app credit costs on hailuoai.video change frequently; check the live table there before budgeting. Nothing in this tutorial is based on the author's personal generation history.
Sources
- [1] MiniMax H3 launch coverage and model blog (Hugging Face)
- [2] MiniMax platform API pricing and H3 reference documentation
- [3] Community references: H3 usage guides and the “H3 gibberish problem” discussion