Tool Guide · Xiaohu Explains

MiniMax H3 Official Handbook: A Three-Part Formula That Turns Prompts from Scene Descriptions into Asset Instructions

For an orbit shot, don't write "orbit shot" — write truck left + pan right. If you don't want background music, close the prompt with "non-diegetic music: N/A." These are the rules MiniMax learned the hard way, and they're all in the official handbook.
TL;DR
  • MiniMax's H3 handbook is built around a three-part formula: reference material description + core creative concept + visual process description.
  • The most useful rules are the counterintuitive ones. For an orbit shot, don't write "orbit shot" — write truck left + pan right.
  • The final section, titled "The Best Way to Write Prompts," offers one piece of advice: don't write them yourself.
This article draws from MiniMax's official H3 handbook (marked as continuously updated; we worked from the August 7 version). Prompts and rules are as published by MiniMax. Pricing references come from MiniMax's July 31 launch post; the handbook itself doesn't list prices.
Opening

The handbook's value is a three-part formula

For the H3 video model it released in late July, MiniMax published a handbook that distills prompt writing into one formula, three elements, and six pitfalls, with dozens of examples pairing original prompts, reference images, and final outputs. It never argues for how powerful the model is; it just hands you copyable writing patterns.

Most AI video prompt guides stop at "be clear about your subject, scene, and action" — technically true, practically useless. This handbook gives you the specific, hard-won rules: for an orbit shot, don't write the words "orbit shot"; if you don't want background music, you have to close with a specific phrase. You could tinker for a month and never stumble onto them.

The whole handbook hinges on this formula:

Reference material description What each asset does + Core creative concept One sentence to frame the whole video + Visual process description What happens every second Three parts — none can be skipped
The official master formula for prompts. Each of the next sections breaks down one part.
Direct quote · Master formula
Complete prompt = reference material description + core creative concept + visual process description
Prerequisites

First, know what H3 can take in and put out

Before writing a single word, know the model's capability envelope — it dictates how you write the prompt. The 7,000-character limit, for instance, explains why official examples routinely run thousands of characters — MiniMax genuinely wants you to fill the space.

ItemH3 Spec
Output length4–15 seconds
Output frame rate24 FPS
Output audioEvery result includes audio, native stereo
Output aspect ratioFirst/last frame mode: follows the original aspect ratio of the input image
Text-to-video mode: follows the user-specified ratio within 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
All-in-one reference mode: follows the user-specified ratio within the range above, or lets H3 decide in auto mode
Resolution768p mode (available now): when the output ratio is between 16:9 and 9:16, the short edge is 768px; otherwise total area is ~1M pixels, e.g., 1536×672 at 21:9. 768p results can be upscaled to 1440p
1440p mode (recommended by MiniMax 👍): when the ratio is between 16:9 and 9:16, the short edge is 1440px; otherwise total area is ~3.7M pixels, e.g., 2976×1248 at 21:9
First/last frame input0–2 images; dimensions [256, 5760] on each side; aspect ratio 5:2 to 2:5
With no image input, it falls back to text-to-video mode
All-in-one reference inputImages ≤ 9, dimensions [256, 5760]
Videos ≤ 3, each 2–15 seconds, total ≤ 15 seconds, dimensions [256, 5760], aspect ratio 5:2 to 2:5
Audio ≤ 3 clips, but audio must be paired with image or video input — audio alone isn't allowed, each 2–15 seconds, total ≤ 15 seconds
Mixed input cap: 12 files total; with no input, falls back to text-to-video mode
Input formatsVideo: H.264/AVC, H.265/HEVC; audio track inside video: AAC, MP3
Image: JPG, JPEG, PNG, WEBP, HEIC, HEIF
Audio: WAV, MP3
File sizeVideo: 50MB each; image: 30MB each; audio: 15MB each (no total limit — limits are per file); API request body: 64MB (recommended to pass assets via URL)
Prompt limitNo more than 7,000 characters
Language supportMultilingual prompt input and output
TTS precisely covers 11 languages, including Chinese, English, Japanese, Korean, French, German, Spanish, and more
Additional languages are explorable — Arabic, Thai, Indonesian, Hindi, and 40+ languages total
Suitable for global brand campaigns, multilingual ads, cross-language shorts, and localized product stories

On the input side, there are two very different entry points. Pick the wrong one and the best prompt in the world won't save you:

First / last frame entry 0–2 images 0 images = pure text-to-video 2 images = first frame + last frame No video, no audio All-in-one reference entry Images ≤ 9 Videos ≤ 3 Audio ≤ 3 2–15s each Total ≤ 15s Can't be sent alone Mixed input: 12 files max
The audio rule is the easiest to trip over: audio must accompany an image or video. Illustration by this site, based on the handbook's spec table.

The three capability pillars the handbook defines for H3

Knowing what the model is actually good at tells you where to point your prompt. The handbook opens with an official categorization, quoted directly:

Three capabilitiesOfficial description
Native multimodal
understanding & generation
Supports text, image, audio, and video inputs; understands the people, motion, sound, emotion, camera work, style, and intent across different materials; naturally fuses multiple references into one integrated process from understanding source material to generating complete audiovisual content.
Precise multimodal
editing & control
Supports multi-dimensional editing of people, objects, scenes, sound, and rhythm, with fine-grained instruction following. Creators can iterate on existing content, more reliably realizing adjustments to picture, sound, and detail.
Commercial-grade
content generation
Aimed at real commercial scenarios in film, advertising, brands, e-commerce, and games. It handles subtitles, brand messaging, creative effects, product displays, UI/UX motion demos, game visuals, and stylized expression, helping creators and content teams move faster through concept validation, visual pitches, and production.

The first two pillars each break down into four sub-capabilities, and the handbook's examples are organized along those lines:

Native multimodal understanding & generationPrecise all-round editing & modification
Multi-asset joint reference
Text, image, audio, and video combine freely to inform generation
Character & object editing
Replace, remove, or add people and objects; flexibly adjust the visual focus
Character, action & camera reference
Reference the protagonist's look, body movement, camera motion, and composition
Scene & visual effects editing
Replace backgrounds, adjust lighting and atmosphere, and add or modify VFX
Voice transfer & cloning
Reference an audio clip's timbre to seamlessly update a character's voice
Voice, dialogue & timbre modification
Replace dialogue, transfer timbre, adjust voice and related sound content
Atmosphere & editing comprehension
Reference the original video's visual style, narrative rhythm, sound atmosphere, and overall editing feel
High-precision instruction control
Accurately execute local edits and refinements while keeping unedited content stable
Formula part one

Part one: assign a job to every asset

This is the biggest mental shift in the whole handbook. In the past, the first thing you did was describe the picture. Now the first thing you do is go through every asset you uploaded and say what it's for.

Two things to cover: the label (in upload order, @image1, @audio2, @video3) and the job description.

@image1 @video1 @audio1 Character reference Lock this face Action reference Perform this motion Voice reference Use this person's voice
Three assets, three assignments, all in the same prompt. Illustration by this site.

MiniMax lists the "jobs" as a vocabulary of 13 labels. The list itself is copy-paste material; scan it when writing part one:

Character reference lock face/look Object reference lock object Scene reference lock scene Keyframe be explicit about first/last Voice reference lock timbre Storyboard generate per the boards Style reference match this style Composition reference match this layout Audio reuse use the audio as-is Partial audio reuse use only a segment/track Action reference lock the motion Camera reference lock the camera move Video editing add/remove/modify in a video
Direct quote · Part one example
@image1 provides the appearance of character XX (if there are multiple characters, make sure the reference is unambiguous), @video2 provides the action reference.

Two practical extras. If a specific trait in your reference really must survive, say so explicitly — consistency improves. For voice cloning, if the dialogue or lyrics have to stay intact, MiniMax strongly recommends pasting the full lyrics into the prompt, e.g., "@audio1 serves as the voice-cloning source for the target video, with the exact lyrics: 'ABCDEFGHIJKLMN'."

In practice, the "voice reference" label works like this: give an audio clip as the timbre reference, give it a line to speak, and the character delivers that line in that voice.

Voice cloning example. The prompt is a single line: "The character speaks: Follow the wind, live free. Leave worries behind, enjoy the moment. Voice reference: audio 1." Source: MiniMax official handbook

The last label, "video editing," covers modifying an existing clip. The simplest prompt is plain language: "replace the cat in the video with a dog."

The full prompt: "Replace the cat in the video with a dog." Source: MiniMax official handbook

If you didn't upload anything, skip this section.

Formula part two

Part two: one sentence to frame the whole video

With asset assignments done, establish what this video is. The core creative concept should lock the piece in one sentence. MiniMax requires four things in it: subject (who/what), location (where), event (what's happening), and genre/style (live-action / animation / cinematic / commercial...), plus an optional special camera move (aerial / one continuous shot / slow motion).

Direct quote · Part two example
A young woman in Hanfu (@image1) practices swordplay in a courtyard where cherry blossoms drift through the air — classical Chinese aesthetic, cinematic, one continuous shot.
The subject is the woman in Hanfu (@image1 locks her face), the location is the cherry-blossom courtyard, the event is sword practice, the genre is classical-Chinese style with cinematic texture, and the camera move is added at the end. All four elements are present.

Two hard rules hide in this section, rarely mentioned elsewhere.

Rule one

H3 will cut by default. If you don't say "one continuous shot," it will cut on its own. You have to be explicit.

Rule two

For an orbit shot, don't write "orbit shot." MiniMax recommends truck left + pan right, or the reverse, truck right + pan left.

MiniMax also lists specific options for "genre/style" and "cutting style" that you can pick from directly:

Direct quote · Style and cutting options
Genre/style: realistic / animation / cinematic / ad / documentary...
Specific styles: cyberpunk / neon aesthetics / graffiti style, etc.

Special camera moves: aerial / one continuous shot / slow motion...
  · Cuts happen by default
  · Camera style: be specific. For orbit shots, use truck left + pan right
    or truck right + pan left rather than saying "orbit shot"
  · Cutting style:
    · Standard cuts
    · Dissolves (fade)
    · Cuts synced to the rhythm
    · Quick cuts

If any element links directly to a reference asset, feel free to emphasize it with @image1 / @audio1 / @video1.
What truck and pan mean

A truck is the whole camera sliding sideways; a pan is the camera staying put while the lens rotates left or right. When you circle someone with your phone, your feet truck sideways and your wrist pans to keep them centered — stack the two and it reads as an orbit. The official advice: decompose the result into two concrete actions, which the model understands better.

One more text-related rule: if you want specific text, a logo, a headline, or button copy to appear in the video, you must write the exact string into the prompt. For example, "The phone screen displays the title: 'AI Video Creation' with the button text: 'Start Now'." In the handbook's full-of-text motion examples, every character is hard-coded into the prompt:

Motion-graphics example. The prompt specifies the brand name and slogan that must appear, plus "No other readable text may appear in the frame." Source: MiniMax official handbook
Formula part three

Part three: write what you want — and don't want — on a timeline

Once the overall tone is locked, you finally write what happens in every second. This section is organized into timeline blocks, with each cut acting as a timestamp, and each block covering six items.

0–3s 3–8s 8–12s …one block per cut Each block covers six items: Shot size Content Camera Action Dialogue Sound FX After the "want" list, write a separate "don't want" section — the most common case is killing the background music
How to structure the visual process description. Illustration by this site, drawn from the handbook.
Direct quote · Part three example
0–3s: Wide shot. The woman (@image1) walks slowly into the cherry-blossom courtyard (@image2) from screen left, background blurred, no dialogue, just footsteps.
3–8s: Cut to a medium shot of the woman (@image1). She draws her sword (@image3), slowly assumes a stance, and cherry petals fall from the tree. Camera pushes in.
8–12s: Cut to a close-up. A flash of the blade, slow motion, petals scatter from the force of the sword.
Non-diegetic music: N/A

That last line is how you write "don't want." If you want no background music, saying "no BGM" doesn't work. You have to write this:

Direct quote · Turning off background music
Non-diegetic music: N/A

Or in English format:
"non_diegetic_music": N/A
What "non-diegetic music" means

It's the layer of music the characters can't hear and only the audience does — i.e., the score. The tense strings in a horror movie are non-diegetic; the protagonist can't hear them. The opposite is a sound source that exists in the scene itself, like a radio the character switches on. To stop H3 from adding a score, you have to name it.

Two more notes: favor concrete, visible images over metaphors — H3 is built for explicit visual instructions, not abstract similes; and if there's dialogue, write the exact lines, don't just say "she says something cutting."

What does this section look like at full power? One handbook prompt asks for four simultaneous changes, and every one is specified to the letter:

The prompt asks: replace the canned drink at the start with Coca-Cola; change the glowing "FamilyMart" sign in the background to "Hu Hui"; swap every snack in the plastic bag at the end for Coke cans; and change the final line of dialogue from "I bought some snacks" to "I bought a bunch of Coke." Source: MiniMax official handbook
Storyboards

If dialogue and shots are out of sync, the lip-sync breaks

But there's a storyboard trap the handbook calls out specifically. It adds a note after this rule: "This one needs stressing — a lot of lip-sync problems come from exactly this."

Core

Dialogue length has to match shot length. Don't stuff a long speech into a 3-second shot. MiniMax's exact words: "it will seriously hurt the result."

❌ A whole speech squeezed into a 3-second shot Shot 1 · 3s Dialogue far exceeds what this shot can hold → lip-sync breaks ✓ Dialogue broken into pieces, each matched to a shot Shot 1 · 3s Shot 2 · 5s Shot 3 · 3s Dialogue length = shot length
Illustration by this site

Three related rules:

If dialogue spans multiple shots, say which shots it spans. H3 can handle J-cuts and L-cuts, but you have to spell them out.

What J-cut / L-cut are

Sound and picture deliberately switch at different moments: the dialogue continues into the next shot before the picture changes, or the next scene's audio arrives before the image catches up. Like hearing someone call your name from the doorway a beat before you see them — audio first, picture second.

Make clear who is speaking and whether they're on screen. The handbook gives a complex example:

Direct quote · How to write dialogue across shots
A voice-over begins: Wake up, wake up. Then it cuts to a medium close-up of a middle-aged woman, owner of the voice, who continues: It's time to go to school!

When you cut, say what shot size you're cutting to and which earlier character is the subject — cross-shot consistency stays much better.

Variants

Three modes, three prompt-writing styles

The same formula changes depending on how you feed it. The handbook splits it into three modes:

ModeWriting notes
Multimodal reference
image + video + audio together
Give every asset a role: @image1 locks the face, @video1 locks the action, @audio1 sets the mood
Image-to-videoWith a single image, say whether it's the first frame or the last frame; with both first and last frames, H3 won't add cuts on its own — it only fills in the motion, light, and sound between the two frames
Pure text generationBe more specific overall (subject appearance, scene details, action must all be explicit); lean on the layered style of "wide shot to establish space + medium shot to carry action + close-up for detail"

That middle row is the trap: a lot of people send first/last frames hoping H3 will "perform" between them, and it just obediently fills in the motion with no shot changes. That's by design, not a failure.

The error log

Six most common pitfalls, as listed by MiniMax itself

With the writing rules covered, here's the handbook's own summary of mistakes. This table belongs next to your keyboard:

What you wroteHow to fix it
One long paragraph, no structureBreak it into the three-part formula
Uploaded assets but never said what they're forAdd "@image1 is an XX reference"
Want music but wrote "no BGM"Those contradict each other. Drop one or split it per scene
Want one continuous shot but wrote many storyboardsKeep the whole text as one narrative; delete the 【Shot N】 structure
Want a consistent lead face but didn't upload an imageYou must upload a character reference and label it "character reference"
Prompt too short and no reference materialWrite at least subject appearance + scene details + action + style
Model texts

Two thousand-character full examples you can copy

Rules only get you so far. At the end of the handbook, MiniMax provides two fully written examples — complete prompts over a thousand characters each, strictly following the three-part structure.

Example one: the on-screen-lyrics MV

Trap music video aesthetic. The most instructive part is a dedicated "groove rules" section that binds the beat to specific on-screen actions:

Direct quote · Groove rules
hi-hat roll → rapid micro-shakes, frame skips, text splitting into fragments
snare → text suddenly enlarges, hard cut, character's shoulders press down
808 bass hit → low-end presses the frame, brief image distortion, text stretches vertically or compresses horizontally
vocal keyword → lip-sync, jaw movement, head-bobbing, hand gestures push forward
It doesn't vaguely say "sync to the beat." It maps each drum sound to one concrete on-screen action.
The final output generated from the full on-screen-lyrics MV prompt. Source: MiniMax official handbook

Example two: hand-drawn effects

A live-action tram car fused with hand-drawn glowing animation. The most instructive part is the shape-continuity requirement: one hand-drawn line changes form a dozen times across 15 seconds (ticket → paper swallow → caterpillar → arrow → little sailboat → mini tram → snail → umbrella → small fish → sunset cloud sea), and MiniMax requires each transformation to retain traces of the previous form — the ticket's dashed border, the swallow's wings, the caterpillar's dots — so the audience feels it's one thing continuously transforming, not a new character appearing.

It also specifies the camera's rhythm: the camera always trails the hand-drawn animation by half a beat, letting the animation run first and the camera hesitantly chase after it.

The final output generated from the full hand-drawn effects prompt. Source: MiniMax official handbook
Common thread

Both thousand-character prompts share a trait: a huge share of the space goes to "don't want" — the prohibitions are more detailed than the requests. The hand-drawn one lists "no polished 3D rendering, no ad-style tidy composition, no subtitles, logos, or background music, no giant eyes or split mouths, no new character appearing out of nowhere after the previous form vanishes" — closing off the ways it could go wrong before they happen.