ByteDance公式「Seedance 2.5」プロンプトガイド:1文の基本構文から30秒動画まで、全テンプレートを収録
50件の参照素材にどう役割を割り当てるか、30秒動画をどう分割するか、映像の一部だけをどう直すか。タスク別の推奨記法が、コピー可能なテンプレートとして公式にまとめられました。
- ByteDanceがSeedance 2.5のプロンプト記法を公式解説。一文生成、素材の割り当て、長尺動画の分割、編集、延長まで、タスクごとに公式テンプレートがあります。全例をこのページに収録しているので、そのまま使えます。
- ガイド全体が教えているのは、実は一つだけです。モデルに推測させないこと。全編を貫く3つの構文を覚えれば、半分は理解したも同然です。
- 最後には、モデルでは難しいことも率直に明記されています。公式が、直接生成を推奨しない項目もあります。
まず知っておきたいSeedance 2.5の概要
ByteDanceは、Seedance 2.5の公開からわずか2日後に、公式プロンプトガイドを公開しました。一文での生成から、50件の参照素材の割り当て、30秒動画の分割、既存映像の一部だけを変える方法まで、タスクごとにそのまま使えるテンプレートを用意しています。即夢AIやAPIで動画を制作する人にとって、これは初の体系的な公式記法マニュアルです。これまでは、新機能の使い方をコミュニティー投稿から探るしかありませんでした。
初めて触れる人のために、まず概要を押さえましょう。Seedance 2.5はByteDanceのSeedチームが7月31日に公開した動画制作モデルです。1本の動画を、前世代の15秒から倍の30秒まで直接生成できます。1回に組み合わせられる参照素材も15件から50件へ増え、既存動画の編集と延長にも新たに対応しました。公開当日から即夢AIとDoubao Proで利用できます。機能が増えたことで、プロンプトの書き方も変わりました。素材には番号を付けて役割を割り当て、長尺動画は区間ごとに記述する必要があります。このガイドは、まさにその課題を解決するものです。
2.0と比べ、公式が挙げる主な進化は4つです。
- 1区間の上限を30秒へ拡大。カットのつながりが自然になり、後編集でつなぎ合わせなくても、感情の起伏や物語の進行を一つの動画として描けます。
- 高忠実度の時間的延長にも対応します。人物、シーン、カメラの一貫性を維持したまま、短い素材をより長い動画へ自然に延ばせるため、制作効率と完成品質が向上します。テレビCM、ショートドラマ、ブランドストーリーなど、物語の完成度が重視される用途に特に適しています。
- 1回に入力できる上限を、50件のオムニモーダル素材(画像/動画/音声を自由に組み合わせ可能)へ引き上げました。登場人物一式、シーン一式、参照カット、ブランド音楽を一度に入力し、モデルが総合的に理解したうえで、意図に沿った動画を出力できます。
- 複雑な素材を安定して維持する能力も向上しました。プロ向けの3Dホワイトモデル、商品パッケージ、ブランドVIなど、複雑なアセットも元の形を保って反映できます。事前に設計した内容やカメラワークを、そのまま再利用できます。
- 動画の部分編集機能を追加しました。映像全体、カメラ、テンポを変えずに、背景、商品、人物などの一部だけを正確に変更できます。「1回制作し、複数版を納品する」運用が可能です。
- 生成結果を修正、反復、納品できるようになり、後工程の手戻りを減らせます。代表例は、海外広告で人種、商品、コピーをすばやく差し替えるローカライズ、EC素材のSKU別一括展開、ブランドコンテンツのチャネル別バージョン制作です。
- 10以上の言語にネイティブ対応します。世界中のクリエイターが、翻訳や追加の言語変換を介さず、母語で制作意図を記述できます。
- 指示の理解力と制御力も向上しました。複雑なカメラワーク、感情の転換、多層的なシーン記述にも、より正確に従います。世界市場向けの制作と商用利用を支えます。
公開内容の詳しい解説、デモ動画、利用方法はこちらの記事にまとめています。
ガイド全体を読むと、すべてのテクニックが一つの原則に集約されていると分かります。モデルが推測する余地を、すべて明記して塞ぐことです。「〜だけを採用し、〜は採用しない」「終了時に画面で確認できるのは〜」「Xは@画像Nに対応する」という3つの頻出構文が、全編を貫いています。以降の各節でも繰り返し登場します。
プロンプトを構成する6要素、必要なものだけ使う
基本式では、主体+動作または出来事が必須です。「誰または何が」「何をしているか」を示します。シーン環境、ビジュアルスタイル、カメラワークとカット、音声の4要素は任意です。不要なら省きます。動作はまず主要な流れを簡潔にまとめ、重要な部分だけ詳しく書きます。同じ動作を繰り返し記述しないでください。生成パラメータ(長さ、アスペクト比)はプロンプトに書かず、生成画面またはAPIで設定します。
<Subject> performs <primary action or event> in <scene and environment>. The visuals feature <visual style>. Use <shot size, camera angle, camera movement, or cuts>. Audio includes <dialogue, ambience, sound effects, or music>.
A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf. Soft morning light enters through the window. The wet clay has a delicate sheen, and the workbench remains tidy. Begin with a medium shot of the wheel-throwing process, slowly push in toward the cup's surface texture, then cut to a frontal view of the shelf. Retain the low hum of the pottery wheel, the friction of clay, and subtle indoor ambience.
音声と文字を指定する4つの記号
プロンプトは基本的に自然言語で書けます。音楽、効果音、セリフ、字幕をさらに区別したい場合は、次の記号を使います。
| 内容 | 記号 | 例 |
|---|---|---|
| 音楽 | () | (穏やかなテンポのピアノ曲が流れる) |
| 効果音 | <> | <遠くから鐘の音が聞こえる> |
| セリフ | {} | {こんにちは、おかえりなさい} |
| 字幕 | 【】 | 【第1章:旅立ち】 |
字幕や音声を制御するときは、残すものと不要なものを明記します。
No background music. Keep only the characters' dialogue, ambience, and action sound effects. No subtitles. No audio at all.
セリフの言語を安定させる書き方
中国語以外のセリフを使う場合は、セリフの前に言語を明記します。
The girl says softly in Japanese: {もう大丈夫です}英語の文章なのにモデルが中国語で話す場合や、地域アクセントを指定したい場合は、次の式で補強します。
Dialogue Language + Regional Variety or Accent + Delivery Style + Speaker + {Dialogue}
Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren't coming.}
Dialogue language: authentic Los Angeles English. The young man says in natural Los Angeles vernacular: {No way, you actually made it.}50件入力できても、安定域はずっと小さい
1回に最大50件の参照素材を組み合わせられます。4種類の素材にはそれぞれ入力上限がありますが、公式はそれより小さい「推奨範囲」も示しています。上限は機能上の限界であり、推奨範囲こそ安定域です。素材が増えるほど、安定性は下がりやすくなります。
| 素材タイプ | 入力範囲 | 推奨範囲(安定域) |
|---|---|---|
| 画像 | 最大30枚、1枚4K以下 | 主体画像は1〜8主体を推奨 |
| 動画 | 最大10本、合計30秒以下 | 1〜5主体、1本5〜10秒 |
| 音声 | 最大10本、合計30秒以下 | タスクに直接関係するセリフ、声質、環境音、音楽のみ |
| 動画編集 | 動画と参照画像を併用 | 元動画20秒以内、参照画像1〜5枚 |
推奨範囲を超えて試すこともできます。主体画像は9〜12主体、主体の音声・動画は6〜10主体、編集用の参照画像は6〜8枚まで拡張できます。安定性を高めるもう一つのポイントは、同じ主体を複数の角度から示す場合、複数の角度を1枚のコラージュにせず、1枚に1視点を載せることです。宣言は次のように書きます。
@Image 1 defines the front view of the same folding desk lamp. @Image 2 defines the left-side structure of the same folding desk lamp. @Image 3 defines the right-side structure of the same folding desk lamp. @Image 4 defines the rear structure of the same folding desk lamp. All four images define one folding desk lamp. The output must contain only one lamp throughout.
素材ごとに明記する:何を採用し、何を採用しないか
素材をアップロードしたら、各素材が何を提供するかを一つずつ宣言します。鉄則は3つです。対応関係を必ずプロンプト本文に書く。画像内に貼った文字ラベルだけに頼らない。どの素材がどの人物、小道具、シーンに対応するかをモデルに推測させない。「採用しない」はすべての素材に書く必要はありません。素材内の人物、背景、構図が誤って動画へ入り込みそうな場合だけ明記します。
@Image 1 defines <subject>'s <appearance, clothing, structure, or material>. @Video 1 defines <motion, camera movement, or pacing>. @Audio 1 defines <character or sound type>'s <voice, dialogue, ambience, or music>. <Subject> completes <primary action or event> in <scene>. The visuals feature <visual style>, with <camera treatment>.
公式例を見る(陶芸家・素材版)
@Image 1 defines the ceramic artist's facial features, hairstyle, and dark green apron. Do not use the image background. @Image 2 defines the wooden workbench, window placement, and morning light of the pottery studio. Do not use the people in the image. @Video 1 defines the pacing of throwing clay with both hands, lifting the cup, and placing it down. Do not use the person's identity, clothing, or scene from the video. The ceramic artist finishes a pale blue cup in the pottery studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf. Begin with a medium shot of the wheel-throwing process, then slowly push in toward the cup's surface texture. Retain the sound of the wheel, the friction of clay, and indoor ambience.
もう一つ注意があります。参照動画によって動作、カメラワーク、順序がすでに正確に決まっている場合は、継承する内容だけを説明してください。動作を一つずつ書き直す必要はありません。重複記述は、素材自体と競合するおそれがあります。
数十件の素材を整理する:シーン別に割り当てる
素材が増えると、プロンプトの中心は人物、小道具、シーン、動作、音声の対応関係を明記することへ移ります。全素材を一文へ詰め込んでも意味はありません。公式は5段階の整理法を示しています。最終目標は、各シーンで正しい素材を選ばせることです。全素材を同時に登場させることではありません。
第1段階では、人物、商品、小道具を個別に命名し、それぞれの素材と結び付けます。
<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing. <Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing. <Prop A> corresponds to @Image 3. Use only the structure, material, and color. <Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.
第2段階では、素材が増えたらタイプ別に整理します。人物、小道具、シーン、動作と音声を、それぞれ別のグループに分けます。
[Characters] <Conservator> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing. <Registrar> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing. <Exhibition Installer> corresponds to @Image 3. Use only the appearance, hairstyle, and clothing. <Guide> corresponds to @Image 4. Use only the appearance, hairstyle, and clothing. Do not interchange the four characters' appearances, clothing, actions, positions, or dialogue. [Props] <Sample Case> corresponds to @Image 5 and belongs only to <Conservator>. <Record Board> corresponds to @Image 6 and belongs only to <Registrar>. [Scenes] <Conservation Lab> references @Image 7. Use only the space, materials, and lighting. <Gallery> references @Image 8. Use only the space, materials, and lighting. [Motion and Audio] @Video 1 defines the motion of <Conservator> opening <Sample Case>. Do not use the person or scene from the video. @Audio 1 defines <Guide>'s voice and specified dialogue.
第3段階では、同じ人物を複数のシーンで複数素材から参照する場合、その人物の情報を一つの設定にまとめます。
[Subject Profile: Conservator] Appearance and clothing: @Image 1. Fixed prop: <Sample Case> from @Image 5. Locations: <Conservation Lab> and <Gallery>. Motion references: the case-opening motion from @Video 1 and the sample-placement motion from @Video 2. Do not use: other characters' clothing. Do not give this character <Record Board> or guide equipment.
第4段階では、シーンごとに使用する素材、発生する出来事、終了状態だけを指定します。
Scene 1 | Inspection in the Conservation Lab Use: <Conservator>, <Sample Case>, <Conservation Lab>, and the case-opening motion from @Video 1. Event: <Conservator> opens <Sample Case> at the workbench and inspects the sample inside. End state: <Conservator> remains on the inner side of the workbench. <Sample Case> stays beside the conservator's right hand, which is on the left side of the frame. Scene 2 | Registration in the Gallery Use: <Registrar>, <Record Board>, and <Gallery>. Event: <Registrar> checks the number on <Record Board> beside the display case. End state: <Registrar> still holds <Record Board> with both hands. No other character enters the display-case area.
30秒動画:段階に分け、各段階では一つだけ変える
出来事の多い物語は連続する段階に分けます。各段階には、主要な状態変化を一つだけ割り当てます。その段階の終了時に、画面で直接確認できる状態も明記してください。花束を誰が持っているか、はさみをどちら側へ戻したか、といった情報です。「終了時」の行は、この記法全体の要です。モデルが確認できる画面上のアンカーになります。
[Generation Goal] Generate a <video type>. The central subject is <subject>, and the primary event is <story summary>. [Stage 1] Initial state: <initial state of characters, props, and scene>. Primary event: <one primary action or event>. End state: <character positions, prop ownership, or visible scene state>. [Stage 2] Continue from the previous stage: <state that must remain unchanged>. Primary event: <one primary action or event>. End state: <observable state>. [Stage 3] Primary event: <closing event>. End state: <final visible state>. [Maintain Consistency] Keep <character identity, number of characters, clothing, prop ownership, spatial direction, and audio relationships> consistent.
公式例を見る(花店の注文包装工程)
[Generation Goal] Generate an instructional video showing a flower shop's order-packing process. <Florist> and <Store Assistant> arrange, wrap, and hand off a bouquet together. [Stage 1] Initial state: <Florist> stands behind the workbench. Loose flower stems, scissors, and wrapping paper lie on the tabletop. Primary event: <Florist> arranges the stems and trims them to length. End state: <Florist> holds the bouquet in the left hand, and the scissors are back on the right side of the workbench. [Stage 2] Continue from the previous stage: both characters retain the same identities and clothing, and <Florist> still holds the bouquet. Primary event: <Store Assistant> unfolds the wrapping paper. <Florist> places the bouquet inside and ties it with a green ribbon. End state: the wrapped bouquet lies flat in the center of the workbench, with the ribbon bow facing the camera. [Stage 3] Primary event: <Store Assistant> picks up the bouquet and places it on the pickup shelf. End state: the bouquet is centered on the pickup shelf, and both characters stand behind the workbench inspecting the finished order. [Maintain Consistency] Keep <Florist> and <Store Assistant>'s identities, clothing, workbench orientation, scissors position, and bouquet ownership consistent.
