MiniMax H3 Official Handbook: A Three-Part Formula That Turns Prompts from Scene Descriptions into Asset Instructions
truck left + pan right. If you don't want background music, close the prompt with "non-diegetic music: N/A." These are the rules MiniMax learned the hard way, and they're all in the official handbook.- MiniMax's H3 handbook is built around a three-part formula: reference material description + core creative concept + visual process description.
- The most useful rules are the counterintuitive ones. For an orbit shot, don't write "orbit shot" — write
truck left + pan right. - The final section, titled "The Best Way to Write Prompts," offers one piece of advice: don't write them yourself.
The handbook's value is a three-part formula
For the H3 video model it released in late July, MiniMax published a handbook that distills prompt writing into one formula, three elements, and six pitfalls, with dozens of examples pairing original prompts, reference images, and final outputs. It never argues for how powerful the model is; it just hands you copyable writing patterns.
Most AI video prompt guides stop at "be clear about your subject, scene, and action" — technically true, practically useless. This handbook gives you the specific, hard-won rules: for an orbit shot, don't write the words "orbit shot"; if you don't want background music, you have to close with a specific phrase. You could tinker for a month and never stumble onto them.
The whole handbook hinges on this formula:
Complete prompt = reference material description + core creative concept + visual process description
First, know what H3 can take in and put out
Before writing a single word, know the model's capability envelope — it dictates how you write the prompt. The 7,000-character limit, for instance, explains why official examples routinely run thousands of characters — MiniMax genuinely wants you to fill the space.
| Item | H3 Spec |
|---|---|
| Output length | 4–15 seconds |
| Output frame rate | 24 FPS |
| Output audio | Every result includes audio, native stereo |
| Output aspect ratio | First/last frame mode: follows the original aspect ratio of the input image Text-to-video mode: follows the user-specified ratio within 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 All-in-one reference mode: follows the user-specified ratio within the range above, or lets H3 decide in auto mode |
| Resolution | 768p mode (available now): when the output ratio is between 16:9 and 9:16, the short edge is 768px; otherwise total area is ~1M pixels, e.g., 1536×672 at 21:9. 768p results can be upscaled to 1440p 1440p mode (recommended by MiniMax 👍): when the ratio is between 16:9 and 9:16, the short edge is 1440px; otherwise total area is ~3.7M pixels, e.g., 2976×1248 at 21:9 |
| First/last frame input | 0–2 images; dimensions [256, 5760] on each side; aspect ratio 5:2 to 2:5 With no image input, it falls back to text-to-video mode |
| All-in-one reference input | Images ≤ 9, dimensions [256, 5760] Videos ≤ 3, each 2–15 seconds, total ≤ 15 seconds, dimensions [256, 5760], aspect ratio 5:2 to 2:5 Audio ≤ 3 clips, but audio must be paired with image or video input — audio alone isn't allowed, each 2–15 seconds, total ≤ 15 seconds Mixed input cap: 12 files total; with no input, falls back to text-to-video mode |
| Input formats | Video: H.264/AVC, H.265/HEVC; audio track inside video: AAC, MP3 Image: JPG, JPEG, PNG, WEBP, HEIC, HEIF Audio: WAV, MP3 |
| File size | Video: 50MB each; image: 30MB each; audio: 15MB each (no total limit — limits are per file); API request body: 64MB (recommended to pass assets via URL) |
| Prompt limit | No more than 7,000 characters |
| Language support | Multilingual prompt input and output TTS precisely covers 11 languages, including Chinese, English, Japanese, Korean, French, German, Spanish, and more Additional languages are explorable — Arabic, Thai, Indonesian, Hindi, and 40+ languages total Suitable for global brand campaigns, multilingual ads, cross-language shorts, and localized product stories |
On the input side, there are two very different entry points. Pick the wrong one and the best prompt in the world won't save you:
The three capability pillars the handbook defines for H3
Knowing what the model is actually good at tells you where to point your prompt. The handbook opens with an official categorization, quoted directly:
| Three capabilities | Official description |
|---|---|
| Native multimodal understanding & generation | Supports text, image, audio, and video inputs; understands the people, motion, sound, emotion, camera work, style, and intent across different materials; naturally fuses multiple references into one integrated process from understanding source material to generating complete audiovisual content. |
| Precise multimodal editing & control | Supports multi-dimensional editing of people, objects, scenes, sound, and rhythm, with fine-grained instruction following. Creators can iterate on existing content, more reliably realizing adjustments to picture, sound, and detail. |
| Commercial-grade content generation | Aimed at real commercial scenarios in film, advertising, brands, e-commerce, and games. It handles subtitles, brand messaging, creative effects, product displays, UI/UX motion demos, game visuals, and stylized expression, helping creators and content teams move faster through concept validation, visual pitches, and production. |
The first two pillars each break down into four sub-capabilities, and the handbook's examples are organized along those lines:
| Native multimodal understanding & generation | Precise all-round editing & modification |
|---|---|
| Multi-asset joint reference Text, image, audio, and video combine freely to inform generation |
Character & object editing Replace, remove, or add people and objects; flexibly adjust the visual focus |
| Character, action & camera reference Reference the protagonist's look, body movement, camera motion, and composition |
Scene & visual effects editing Replace backgrounds, adjust lighting and atmosphere, and add or modify VFX |
| Voice transfer & cloning Reference an audio clip's timbre to seamlessly update a character's voice |
Voice, dialogue & timbre modification Replace dialogue, transfer timbre, adjust voice and related sound content |
| Atmosphere & editing comprehension Reference the original video's visual style, narrative rhythm, sound atmosphere, and overall editing feel |
High-precision instruction control Accurately execute local edits and refinements while keeping unedited content stable |
Part one: assign a job to every asset
This is the biggest mental shift in the whole handbook. In the past, the first thing you did was describe the picture. Now the first thing you do is go through every asset you uploaded and say what it's for.
Two things to cover: the label (in upload order, @image1, @audio2, @video3) and the job description.
MiniMax lists the "jobs" as a vocabulary of 13 labels. The list itself is copy-paste material; scan it when writing part one:
@image1 provides the appearance of character XX (if there are multiple characters, make sure the reference is unambiguous), @video2 provides the action reference.
Two practical extras. If a specific trait in your reference really must survive, say so explicitly — consistency improves. For voice cloning, if the dialogue or lyrics have to stay intact, MiniMax strongly recommends pasting the full lyrics into the prompt, e.g., "@audio1 serves as the voice-cloning source for the target video, with the exact lyrics: 'ABCDEFGHIJKLMN'."
In practice, the "voice reference" label works like this: give an audio clip as the timbre reference, give it a line to speak, and the character delivers that line in that voice.
The last label, "video editing," covers modifying an existing clip. The simplest prompt is plain language: "replace the cat in the video with a dog."
If you didn't upload anything, skip this section.
Part two: one sentence to frame the whole video
With asset assignments done, establish what this video is. The core creative concept should lock the piece in one sentence. MiniMax requires four things in it: subject (who/what), location (where), event (what's happening), and genre/style (live-action / animation / cinematic / commercial...), plus an optional special camera move (aerial / one continuous shot / slow motion).
A young woman in Hanfu (@image1) practices swordplay in a courtyard where cherry blossoms drift through the air — classical Chinese aesthetic, cinematic, one continuous shot.
Two hard rules hide in this section, rarely mentioned elsewhere.
H3 will cut by default. If you don't say "one continuous shot," it will cut on its own. You have to be explicit.
For an orbit shot, don't write "orbit shot." MiniMax recommends truck left + pan right, or the reverse, truck right + pan left.
MiniMax also lists specific options for "genre/style" and "cutting style" that you can pick from directly:
Genre/style: realistic / animation / cinematic / ad / documentary...
Specific styles: cyberpunk / neon aesthetics / graffiti style, etc.
Special camera moves: aerial / one continuous shot / slow motion...
· Cuts happen by default
· Camera style: be specific. For orbit shots, use truck left + pan right
or truck right + pan left rather than saying "orbit shot"
· Cutting style:
· Standard cuts
· Dissolves (fade)
· Cuts synced to the rhythm
· Quick cuts
If any element links directly to a reference asset, feel free to emphasize it with @image1 / @audio1 / @video1.
A truck is the whole camera sliding sideways; a pan is the camera staying put while the lens rotates left or right. When you circle someone with your phone, your feet truck sideways and your wrist pans to keep them centered — stack the two and it reads as an orbit. The official advice: decompose the result into two concrete actions, which the model understands better.
One more text-related rule: if you want specific text, a logo, a headline, or button copy to appear in the video, you must write the exact string into the prompt. For example, "The phone screen displays the title: 'AI Video Creation' with the button text: 'Start Now'." In the handbook's full-of-text motion examples, every character is hard-coded into the prompt:
Part three: write what you want — and don't want — on a timeline
Once the overall tone is locked, you finally write what happens in every second. This section is organized into timeline blocks, with each cut acting as a timestamp, and each block covering six items.
0–3s: Wide shot. The woman (@image1) walks slowly into the cherry-blossom courtyard (@image2) from screen left, background blurred, no dialogue, just footsteps. 3–8s: Cut to a medium shot of the woman (@image1). She draws her sword (@image3), slowly assumes a stance, and cherry petals fall from the tree. Camera pushes in. 8–12s: Cut to a close-up. A flash of the blade, slow motion, petals scatter from the force of the sword. Non-diegetic music: N/A
That last line is how you write "don't want." If you want no background music, saying "no BGM" doesn't work. You have to write this:
Non-diegetic music: N/A Or in English format: "non_diegetic_music": N/A
It's the layer of music the characters can't hear and only the audience does — i.e., the score. The tense strings in a horror movie are non-diegetic; the protagonist can't hear them. The opposite is a sound source that exists in the scene itself, like a radio the character switches on. To stop H3 from adding a score, you have to name it.
Two more notes: favor concrete, visible images over metaphors — H3 is built for explicit visual instructions, not abstract similes; and if there's dialogue, write the exact lines, don't just say "she says something cutting."
What does this section look like at full power? One handbook prompt asks for four simultaneous changes, and every one is specified to the letter:
If dialogue and shots are out of sync, the lip-sync breaks
But there's a storyboard trap the handbook calls out specifically. It adds a note after this rule: "This one needs stressing — a lot of lip-sync problems come from exactly this."
Dialogue length has to match shot length. Don't stuff a long speech into a 3-second shot. MiniMax's exact words: "it will seriously hurt the result."
Three related rules:
If dialogue spans multiple shots, say which shots it spans. H3 can handle J-cuts and L-cuts, but you have to spell them out.
Sound and picture deliberately switch at different moments: the dialogue continues into the next shot before the picture changes, or the next scene's audio arrives before the image catches up. Like hearing someone call your name from the doorway a beat before you see them — audio first, picture second.
Make clear who is speaking and whether they're on screen. The handbook gives a complex example:
A voice-over begins: Wake up, wake up. Then it cuts to a medium close-up of a middle-aged woman, owner of the voice, who continues: It's time to go to school!
When you cut, say what shot size you're cutting to and which earlier character is the subject — cross-shot consistency stays much better.
Three modes, three prompt-writing styles
The same formula changes depending on how you feed it. The handbook splits it into three modes:
| Mode | Writing notes |
|---|---|
| Multimodal reference image + video + audio together | Give every asset a role: @image1 locks the face, @video1 locks the action, @audio1 sets the mood |
| Image-to-video | With a single image, say whether it's the first frame or the last frame; with both first and last frames, H3 won't add cuts on its own — it only fills in the motion, light, and sound between the two frames |
| Pure text generation | Be more specific overall (subject appearance, scene details, action must all be explicit); lean on the layered style of "wide shot to establish space + medium shot to carry action + close-up for detail" |
That middle row is the trap: a lot of people send first/last frames hoping H3 will "perform" between them, and it just obediently fills in the motion with no shot changes. That's by design, not a failure.
Six most common pitfalls, as listed by MiniMax itself
With the writing rules covered, here's the handbook's own summary of mistakes. This table belongs next to your keyboard:
| What you wrote | How to fix it |
|---|---|
| One long paragraph, no structure | Break it into the three-part formula |
| Uploaded assets but never said what they're for | Add "@image1 is an XX reference" |
| Want music but wrote "no BGM" | Those contradict each other. Drop one or split it per scene |
| Want one continuous shot but wrote many storyboards | Keep the whole text as one narrative; delete the 【Shot N】 structure |
| Want a consistent lead face but didn't upload an image | You must upload a character reference and label it "character reference" |
| Prompt too short and no reference material | Write at least subject appearance + scene details + action + style |
Two thousand-character full examples you can copy
Rules only get you so far. At the end of the handbook, MiniMax provides two fully written examples — complete prompts over a thousand characters each, strictly following the three-part structure.
Example one: the on-screen-lyrics MV
Trap music video aesthetic. The most instructive part is a dedicated "groove rules" section that binds the beat to specific on-screen actions:
hi-hat roll → rapid micro-shakes, frame skips, text splitting into fragments snare → text suddenly enlarges, hard cut, character's shoulders press down 808 bass hit → low-end presses the frame, brief image distortion, text stretches vertically or compresses horizontally vocal keyword → lip-sync, jaw movement, head-bobbing, hand gestures push forward
Example two: hand-drawn effects
A live-action tram car fused with hand-drawn glowing animation. The most instructive part is the shape-continuity requirement: one hand-drawn line changes form a dozen times across 15 seconds (ticket → paper swallow → caterpillar → arrow → little sailboat → mini tram → snail → umbrella → small fish → sunset cloud sea), and MiniMax requires each transformation to retain traces of the previous form — the ticket's dashed border, the swallow's wings, the caterpillar's dots — so the audience feels it's one thing continuously transforming, not a new character appearing.
It also specifies the camera's rhythm: the camera always trails the hand-drawn animation by half a beat, letting the animation run first and the camera hesitantly chase after it.
Both thousand-character prompts share a trait: a huge share of the space goes to "don't want" — the prohibitions are more detailed than the requests. The hand-drawn one lists "no polished 3D rendering, no ad-style tidy composition, no subtitles, logos, or background music, no giant eyes or split mouths, no new character appearing out of nowhere after the previous form vanishes" — closing off the ways it could go wrong before they happen.
