ByteDance's Seedance 2.5 prompt guide: Copy-ready templates for videos up to 30 seconds
How to assign roles across 50 assets, structure a 30-second video, or change just one part of an existing clip—the official guide turns every task into a reusable template.
- ByteDance has published its own guide to writing Seedance 2.5 prompts. It covers one-line generation, asset orchestration, long-form sequencing, video editing, and extension—with official templates and every example collected on this page for easy copying.
- The guide follows one rule: Do not make the model guess. Learn its three recurring patterns, and you understand half the system.
- It also spells out what the model cannot do. For several types of content, ByteDance itself says not to expect a perfect first-pass result.
What Seedance 2.5 is
Just two days after releasing Seedance 2.5, ByteDance published an official prompt guide covering everything from one-line generation to orchestrating 50 reference assets, structuring a 30-second video, and changing a single part of an existing clip. Every task comes with a template you can copy directly. For anyone creating videos with Jimeng AI or the API, this is the first comprehensive official guide to the new features. Until now, users had to piece together techniques from community posts.
Here is the quick primer. Seedance 2.5 is a video creation model released by ByteDance's Seed team on July 31. It doubles the maximum length of a single generated video from 15 seconds to 30 seconds, raises the number of reference assets that can be combined in one job from 15 to 50, and adds editing and extension for existing videos. It became available in Jimeng AI and Doubao Pro on launch day. Those expanded capabilities demand a different kind of prompt: Assets need numbered roles, and longer videos need scene-by-scene structure. That is exactly what this guide addresses.
ByteDance highlights four major advances over 2.0:
- The maximum duration of a single clip rises to 30 seconds, allowing more coherent shot construction, emotional development, and narrative progression without stitching clips together in post-production.
- High-fidelity temporal extension can naturally expand short footage while preserving characters, scenes, and camera continuity. This speeds up production and improves quality, especially for narrative-heavy work such as TV commercials, short dramas, and brand films.
- A single job can accept up to 50 omni-modal reference assets in any mix of images, video, and audio. Creators can provide a full cast, location set, reference shots, and brand music at once, allowing the model to interpret them together and produce a result closer to the intended concept.
- Consistency across complex assets has also improved. Detailed assets such as professional 3D blockouts, product packaging, and brand visual identity systems can be carried into the output with much higher fidelity, preserving earlier design and shot-planning work.
- New localized editing tools can precisely change a background, product, person, or other element while keeping the overall frame, camera work, and pacing intact—making it practical to create once and deliver multiple versions.
- Generated footage can now be edited and refined for production, reducing rework in post. Typical uses include localizing ads by swapping people, products, or copy; reusing e-commerce footage across multiple SKUs; and producing channel-specific versions of brand content.
- Native support for more than 10 languages lets creators describe ideas in their own language without translating them first.
- Instruction following has also improved, with more precise control over complex camera choreography, emotional turns, and layered scene descriptions for global creative and commercial work.
Our full launch breakdown, demo videos, and access links are here:
Read the full guide and every technique points to the same principle: Specify everything the model would otherwise have to guess. Three recurring patterns run through the document: "Use only... Do not use..."; "At the end, the frame shows..."; and "X corresponds to @ImageN." You will see them repeatedly in the sections below.
A prompt has six parts—use only what you need
The basic formula is simple: subject + action or event is required. It tells the model who or what is doing what. Four other parts are optional: scene and environment, visual style, camera movement and cuts, and sound. Omit anything the task does not need. Summarize the main action first, add detail only to critical movements, and do not describe the same action twice. Generation settings such as duration and aspect ratio belong in the generation interface or API, not in the prompt.
<Subject> performs <primary action or event> in <scene and environment>. The visuals feature <visual style>. Use <shot size, camera angle, camera movement, or cuts>. Audio includes <dialogue, ambience, sound effects, or music>.
A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf. Soft morning light enters through the window. The wet clay has a delicate sheen, and the workbench remains tidy. Begin with a medium shot of the wheel-throwing process, slowly push in toward the cup's surface texture, then cut to a frontal view of the shelf. Retain the low hum of the pottery wheel, the friction of clay, and subtle indoor ambience.
Four symbols separate music, sound effects, dialogue, and subtitles
Natural language works throughout the prompt. When you need to distinguish music, sound effects, dialogue, and subtitles, use these symbols:
| Content | Symbol | Example |
|---|---|---|
| Music | () | (Soft, measured piano music plays in the background) |
| Sound effect | <> | <A bell rings in the distance> |
| Dialogue | {} | {Hello, welcome back} |
| Subtitle | 【】 | 【Chapter 1: Departure】 |
To control subtitles and sound, state exactly what to keep and what to exclude:
No background music. Keep only the characters' dialogue, ambience, and action sound effects. No subtitles. No audio at all.
Reinforce the intended dialogue language
When dialogue is not in Chinese, name the language before the line:
The girl says softly in Japanese: {もう大丈夫です}If English text is spoken in Chinese, or you need a regional accent, reinforce the instruction with this formula:
Dialogue Language + Regional Variety or Accent + Delivery Style + Speaker + {Dialogue}
Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren't coming.}
Dialogue language: authentic Los Angeles English. The young man says in natural Los Angeles vernacular: {No way, you actually made it.}You can load 50 assets, but the reliable range is much smaller
A single job can combine up to 50 reference assets. Each of the four asset types has an input limit, but the guide also gives a much smaller recommended range. The limit defines what the model accepts; the recommended range is where it tends to stay reliable. Stability usually falls as the number of assets rises:
| Asset type | Input range | Recommended range |
|---|---|---|
| Images | Up to 30 images, each no larger than 4K | 1 to 8 subjects is ideal for subject images |
| Video | Up to 10 clips, no more than 30 seconds total | 1 to 5 subjects, with each clip lasting 5 to 10 seconds |
| Audio | Up to 10 clips, no more than 30 seconds total | Keep only dialogue, voice, ambience, or music directly relevant to the task |
| Video editing | Edit with a source video and reference images | Source video under 20 seconds and 1 to 5 reference images |
You can still go beyond the recommended range: subject images can cover 9 to 12 subjects, subject audio or video can cover 6 to 10, and editing jobs can use 6 to 8 reference images. One more detail improves stability: When the same subject needs several angles, use one image per angle instead of combining the views into a collage. Declare them like this:
@Image 1 defines the front view of the same folding desk lamp. @Image 2 defines the left-side structure of the same folding desk lamp. @Image 3 defines the right-side structure of the same folding desk lamp. @Image 4 defines the rear structure of the same folding desk lamp. All four images define one folding desk lamp. The output must contain only one lamp throughout.
Define every asset: What to use and what to reject
Once the assets are uploaded, state what each one contributes. Three rules are non-negotiable: Put the mapping directly in the prompt; do not rely only on text labels embedded in the images; and never make the model guess which asset belongs to which person, prop, or scene. You do not need a "do not use" clause for every asset. Add one only when a person, background, or composition could leak into the final video by mistake.
@Image 1 defines <subject>'s <appearance, clothing, structure, or material>. @Video 1 defines <motion, camera movement, or pacing>. @Audio 1 defines <character or sound type>'s <voice, dialogue, ambience, or music>. <Subject> completes <primary action or event> in <scene>. The visuals feature <visual style>, with <camera treatment>.
View the official example: Potter with reference assets
@Image 1 defines the ceramic artist's facial features, hairstyle, and dark green apron. Do not use the image background. @Image 2 defines the wooden workbench, window placement, and morning light of the pottery studio. Do not use the people in the image. @Video 1 defines the pacing of throwing clay with both hands, lifting the cup, and placing it down. Do not use the person's identity, clothing, or scene from the video. The ceramic artist finishes a pale blue cup in the pottery studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf. Begin with a medium shot of the wheel-throwing process, then slowly push in toward the cup's surface texture. Retain the sound of the wheel, the friction of clay, and indoor ambience.
One more rule: If a reference video already defines the action, camera movement, and sequence precisely, state what to inherit instead of narrating every movement again. Repetition can conflict with the asset itself.
How to organize dozens of assets: Assign them scene by scene
Once the asset count grows, the prompt's real job is to define the relationships among people, props, scenes, actions, and sounds. Cramming every asset into one sentence does not help. ByteDance proposes a five-step method designed to make the model select the right assets in each scene. Getting every asset on screen at once is never the goal:
First, give every person, product, and prop a unique name and bind it to an asset:
<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing. <Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing. <Prop A> corresponds to @Image 3. Use only the structure, material, and color. <Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.
Second, sort larger asset sets by type—people, props, scenes, actions, and sounds each get their own group:
[Characters] <Conservator> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing. <Registrar> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing. <Exhibition Installer> corresponds to @Image 3. Use only the appearance, hairstyle, and clothing. <Guide> corresponds to @Image 4. Use only the appearance, hairstyle, and clothing. Do not interchange the four characters' appearances, clothing, actions, positions, or dialogue. [Props] <Sample Case> corresponds to @Image 5 and belongs only to <Conservator>. <Record Board> corresponds to @Image 6 and belongs only to <Registrar>. [Scenes] <Conservation Lab> references @Image 7. Use only the space, materials, and lighting. <Gallery> references @Image 8. Use only the space, materials, and lighting. [Motion and Audio] @Video 1 defines the motion of <Conservator> opening <Sample Case>. Do not use the person or scene from the video. @Audio 1 defines <Guide>'s voice and specified dialogue.
Third, when the same character uses several assets across multiple scenes, build a single consolidated profile:
[Subject Profile: Conservator] Appearance and clothing: @Image 1. Fixed prop: <Sample Case> from @Image 5. Locations: <Conservation Lab> and <Gallery>. Motion references: the case-opening motion from @Video 1 and the sample-placement motion from @Video 2. Do not use: other characters' clothing. Do not give this character <Record Board> or guide equipment.
Fourth, call assets scene by scene. Each scene names only the assets it uses, the event that occurs, and the required end state:
Scene 1 | Inspection in the Conservation Lab Use: <Conservator>, <Sample Case>, <Conservation Lab>, and the case-opening motion from @Video 1. Event: <Conservator> opens <Sample Case> at the workbench and inspects the sample inside. End state: <Conservator> remains on the inner side of the workbench. <Sample Case> stays beside the conservator's right hand, which is on the left side of the frame. Scene 2 | Registration in the Gallery Use: <Registrar>, <Record Board>, and <Gallery>. Event: <Registrar> checks the number on <Record Board> beside the display case. End state: <Registrar> still holds <Record Board> with both hands. No other character enters the display-case area.
