Product launch · XiaoHu's take

ByteDance launches Seedance 2.5 with 30-second videos and up to 50 reference assets per prompt

Jimeng AI and Doubao Pro rolled it out the same day. The Volcano Engine Ark API is not yet available, and pricing remains undisclosed.

The one-minute read
  • A single generation used to top out at 15 seconds—enough for one shot. Building a full scene meant generating separate clips and stitching them together. A 30-second generation can now hold an entire scene.
  • The reference limit has jumped from 15 assets to 50. One concert demo assigns distinct roles to all 18 reference images.
  • Prompts have changed too: instructions are split into timed segments, and every asset is numbered.
⚑ This article draws on ByteDance Seed's launch blog, the Seedance 2.5 project page, and two posts from official accounts published that day. Every demo video is a handpicked finished result. Doubao Pro plan requirements come from PChome. Claims about Jimeng AI and Dreamina's "long video mode" and 1-second timestamp precision come from the platforms themselves and do not fully match the model blog; each is attributed where it appears.
The launch

Twice the length, over three times the references, and timestamp control

ByteDance Seed launched its latest video-creation model, Seedance 2.5, on July 31. For video creators, the upgrade goes beyond image quality: a single generation can contain an entire scene, dozens of reference assets can be directed at once, and edits can target a specific moment in the timeline.

The launch opened with a 4-minute, 22-second short film generated entirely by Seedance 2.5, with no live-action footage. The same film appeared in a ByteDance Seed (@ByteDanceSeed_) post that day. One obvious question remains: 4 minutes and 22 seconds is far beyond the 30-second single-generation limit, but the launch page never explains how the clips were joined.

A 4-minute, 22-second creative short generated entirely by Seedance 2.5, with visuals and audio created together. Source: Seedance 2.5 launch blog.

Seedance 2.0 launched on February 12, 2026, a little over five months ago. The underlying architecture has not changed: it still uses the same multimodal system to jointly generate audio and video from text, images, audio, and video. The progress comes in three areas: how long one generation can run, how many references it can use at once, and how precisely a finished clip can be edited. That third point needs some context. Seedance 2.0 could already revise selected clips, characters, actions, or plot beats. Seedance 2.5 adds timestamp control and improves green-screen editing, viewpoint and camera-motion control, and reference-based editing.

Max generation length 15 sec · Seedance 2.0 30 sec · Seedance 2.5 Reference limit per generation 15 assets 9 images + 3 videos + 3 audio clips 50 assets 30 images + 10 videos + 10 audio clips Edit precision (equal bars; not a quantity) 2.0: clip / character / action / plot 2.5: adds timestamp control Better green-screen / camera / reference editing
Three changes from Seedance 2.0 (2026-02-12) to 2.5 (2026-07-31), based on figures in the two launch blogs. The first two rows are proportional within each metric and should not be compared horizontally across rows. The third describes capabilities, so its bars are intentionally equal in length.
30 sec
Maximum per generation (15 sec before)
50 assets
Maximum references per generation (15 before)
4:22
Opening film length (joining method undisclosed)
Storytelling

Thirty seconds is enough for a complete scene

Doubling the duration may sound like a simple jump from 15 seconds to 30. The real claim is that 30 seconds can now hold a structured sequence of connected shots—not just the same image stretched to twice the length.

The demo follows a female singer taking the stage. The camera pushes through a gap in heavy red curtains into a warmly lit backstage dressing room, where she adjusts her earpiece with her back to the lens. A crew member tells her it is time to go on. She turns toward the camera and begins singing City Pop. The camera pulls backward as she passes through the curtains and down a backstage corridor, interacting naturally with a dancer while a staff member hands her a microphone. They step onto the stage together. The camera circles behind them, gradually revealing the red-and-black set, LED screens, tracking spotlights, smoke, and reflective floor. Finally, it pulls back to a wide view of the arena, packed with spectators, signs, and glow sticks.

A continuous 30-second take from the dressing room to a wide arena view, with no cuts. It used text only, with no reference assets—the launch page labels it T2V, or text-to-video. Source: Seedance 2.5 launch blog.
Through curtain Push forward Dressing room Called to stage She turns Starts singing Camera retreats Follows her Backstage Circle backstage Stage design Unfolds Pull back Arena wide shot Glow-stick crowd 00:00 00:05 00:10 00:15 00:20 00:25 00:30 Seedance 2.0 single-generation limit Cards show story order; widths are not durations
Five beats from the singer demo, arranged in story order. The launch page gives the sequence of scenes but no second-by-second timing, so the card widths do not represent duration. The bracket below marks where the previous 15-second limit would end.
Full prompt used for the singer sequence
一镜到底手持稳定器跟拍,镜头从红色厚重幕布缝隙缓缓推进,进入暖色后台化妆间。年轻女歌手背对镜头整理耳机,工作人员提醒她准备登台。她回头看向镜头,开始唱起 City Pop。镜头后退跟拍,歌手穿过幕布进后台通道,与舞伴自然互动,一位工作人员把麦克风递给她。随后歌手与舞伴登上舞台,镜头绕至背面,红黑舞美、LED 屏、追光、烟雾与反光地板逐渐展开。镜头最终拉远至体育馆全景,呈现满场观众、灯牌、荧光棒与欢呼声,营造年轻自由的演唱会高潮氛围。
The prompt follows the camera from beginning to end. It starts with the setup—a continuous take tracked on a handheld stabilizer—then describes each event in sequence. It is already more than ten times as detailed as "a female singer performing at a concert."

The subject and background stay consistent through a full camera orbit

The longer a generated video runs, the more likely its subject is to mutate: a face changes after a turn, clothing patterns stop matching, or the background drifts. The Peking Opera demo is built to stress exactly that problem. As the lead spins and flicks her water sleeves, the camera completes an orbit around her. The characters and background stay consistent throughout, while the sleeves trace arcs that respond plausibly to gravity.

Peking Opera demo: an orbiting camera move, swinging water sleeves, and three characters entering in sequence. This one used reference assets—the launch page labels it R2V, meaning video generated from reference images, videos, or audio. The prompt assigns separate images to the overlord, Consort Yu, the martial performer, and the setting. Source: Seedance 2.5 launch blog.

Taking aim at the familiar AI look

AI video often has a recognizable kind of falseness: waxy skin, vacant eyes, excessive saturation, and a greasy sheen over the entire frame. This release adjusts object materials, skin texture and gaze, lighting, and saturation to move closer to live-action footage. The same update also targets another common failure mode—unwanted subtitles or music appearing from nowhere. Neither improvement comes with a before-and-after demo, so the size of the gain can only be judged through hands-on testing.

Extension

Beyond 30 seconds, official sources disagree on extension limits

Thirty seconds is obviously not enough for a complete film. Longer sequences rely on extension: the model continues from an existing video while preserving its characters, setting, visual style, and sound. That removes several manual steps—splitting the sequence into shots, repeatedly generating and stitching clips, and repairing the joins.

Extension demo: a boy runs through a subway car carrying a soccer ball. When the train stops and the doors open, he races outside; the protagonist follows him into the street, eventually catches up, sees the upset boy look up, calms down, and pats his head. The original instruction reads: "延长视频,衔接 @视频 1 画面内容和主体继续生成一个 30 秒的视频,保持人物主体、场景、画面风格、声音音效一致。" Source: Seedance 2.5 launch blog.

Three official sources give different answers about the maximum length, with no explanation of how they relate. The launch blog says the model supports "multiple rounds of extension" for "several minutes" of coherent content. The Seedance 2.5 project page says it "supports two video extensions." Jimeng AI and its international counterpart Dreamina separately advertise a platform-level "long video mode" that runs up to 3 minutes.

Single generation · blog and project page 30 sec Project page: "supports two extensions" ≈90 sec 30×3, inferred from "two" Jimeng AI / Dreamina: "long video mode" 3 min Launch blog: "multiple rounds," "minutes" No number, so no bar can be drawn 060 sec 120 sec180 sec Dark = official figure; light = our estimate
The three claims shown alongside the model's per-generation limit. The horizontal axis is proportional in seconds, with a scale below. The 90-second bar is our inference from "supports two extensions"; no official source states a 90-second limit. The launch blog gives no number for "several minutes," so that row has no bar.

My reading is that the model itself offers 30 seconds plus two extensions, while 3 minutes is a separate platform feature built by Jimeng AI and its international counterpart Dreamina. A Dreamina AI (@dreamina_ai) launch post lists it separately as "Long video mode: up to 3 minutes." Official sources never clarify which claim encompasses the others.

Prompting

Prompts now read like second-by-second shot lists

Duration, references, and editing are the three headline changes. But when the demo prompts are placed side by side, another shift appears: instructions now need timed segments, and every asset needs a number.

The breakfast demo uses this prompt:

Full editing prompt for the breakfast demo
编辑 @视频 1,保持人物、动作及画风不变,仅调整运镜。15 秒分段运镜设计:0–4 秒,微型 FPV 贴锅穿行,跟随弹起的吐司后横甩至咖啡;4–7 秒,推近并沿锅边横移,跟随煎蛋翻起落回;7–11 秒,急速升至顶视角,匀速下压扫过餐盘和钥匙;11–15 秒,手持近镜跟随双手快速横甩,最后推近早餐,再拉回双人中景。全程连贯稳定。
The prompt does two things: it edits the existing @Video 1 while changing only the camera movement. The characters, actions, and visual style remain unchanged.
00:00–00:04 Mini FPV Skim the pan Follow toast Whip to coffee 00:04–00:07 Push in Track pan rim Follow egg Flip and land 00:07–00:11 Rise overhead Press downward Sweep plate And keys 00:11–00:15 Handheld close-up Follow hands Push to food End on two-shot 0s4s7s 11s15s Prompt fixes 4 boundaries; widths match duration
The four-part structure of the breakfast prompt. Card widths and the scale below are proportional to time. Every timestamp and camera move comes directly from the prompt.
Camera-editing demo: the same breakfast scene, with characters and actions preserved while the camera is redesigned around the four segments above. Source: Seedance 2.5 launch blog.

The Peking Opera prompt follows the same pattern: "0-5 seconds: open on a close-up of the overlord in @Image 2... 6-10 seconds: orbit steadily around a medium shot of Consort Yu in @Image 1... 11-20 seconds: the martial performer in @Image 3 enters and performs an aerial backflip." Both prompts share two traits: timed segments and numbered assets.

Timestamp control

The prompt is effectively a shot list, complete with a timeline, an asset list, and a fixed event for each segment. The launch page calls this "precise timestamp control." During generation, each time range can specify the story beat, viewpoint, camera movement, and pacing. After generation, a particular clip can be selected and its characters, actions, sound, or plot changed while maintaining continuity on either side of the edit.

The barrier to operating the tools is lower—you no longer need to know how to shoot, edit, or light a scene. The barrier to explaining what you want is higher. You need to plan the shots in your head and describe them accurately in text. If a generation fails, the usual approach is to reroll the entire clip. Once prompts become precise to the second, success depends on one more variable: whether you clearly specified what should happen when.

The launch blog does not quantify timestamp precision. A post from Dreamina AI (@dreamina_ai) claims "precise timestamp control down to 1 second." That is the platform's own wording.

References

One prompt gives each of 18 reference images a distinct role

More detailed instructions require more source material. The per-generation reference limit has climbed from 9 images + 3 videos + 3 audio clips clips—15 assets total—to 30 images + 10 videos + 10 audio clips clips, or 50 in all. The image allowance alone has risen from 9 to 30.

Fifty slots can feel abstract. The concert demo uses 18 of them, assigning each image a specific job in a single prompt.

@Image 1 Concert hall setting ↑ Rear stage Choir @Images 11 – 14 Orchestra (sides and rear) @Images 6 – 10 Pianist @Image 2 Cellist @Image 3 Violinist @Image 4 Lead singer (moves forward) @Image 5 Front of stage Audience @Images 15 – 18 18 images total; 17 map to roles and zones Prompt gives zones; block positions are our layout
How the 18 reference images are divided in the concert demo, arranged by the numbering in the prompt. The prompt only says who stands in the center, at the sides, or in the rear. The exact block placement is our visualization, not a precise stage plan.
Full prompt for the concert demo
30 秒音乐会片段,16:9 横屏,电影感写实,真实音乐厅光影,暖金色舞台灯,正式古典音乐会氛围。场景@图一,钢琴家参考@图二,大提琴参考@图三,小提琴参考@图四,主唱参考@图五,管弦等其他乐队参考@图六到图十,合唱团参考图十一到图十四,观众席参考@图十五到图十八。主唱舞台中央走向前沿,钢琴家在钢琴旁,乐队分布两侧与后方,合唱团在舞台后方。开场高位俯拍音乐厅全景,钢琴家落键,主唱进入追光演唱,镜头自然带过小提琴、大提琴与管弦乐队,小提琴明亮,大提琴温暖。后段合唱团加入,主唱与第一排观众短暂眼神交流,观众微笑点头。结尾镜头后移,演唱结束,观众跟随鼓掌。
The prompt does three jobs: it assigns 18 images to individual performers and areas, specifies where everyone stands, and choreographs the camera. The first five images are named individually. The other 13 are grouped by number for the orchestra, choir, and audience.
Ensemble demo with multiple people in one frame: several characters' appearances and voices are reproduced at once, with their individual traits remaining stable throughout. Source: Seedance 2.5 launch blog.

Beyond the higher limit, the model adds or improves several reference modes: white-model reference, motion reference, and creative reference. Motion and creative reference are named but not explained or demonstrated. White-model reference gets a full demo, so it deserves a closer look.

Workflow

White-model references and green-screen editing connect 3D and on-set workflows

White-model reference is a reference feature, while green-screen editing is an editing feature. The project page groups them under "built for professional video creation" because both aim at the same target: connecting the model to tools and assets already used in professional production. The terminology can be opaque, so here is what each means.

White-model references: build the structure, then render it

A white model is an untextured gray or white 3D model containing only shapes and positions. Film, advertising, and game teams already use this step: they block out a scene in 3D software, fixing the camera, character positions, composition, and motion paths before production. That process is called previs, short for previsualization. Seedance 2.5 can use the white model directly as a structural reference. Characters remain where they were placed, the camera follows the planned orbit, and objects move along the specified paths, making complex composition and staging more predictable.

Think of it this way

Before filming, you block out the action on a table with toy bricks and figures. Previously, the director still had to recreate that layout on set. Now the layout itself can be handed to the model as the blueprint for the shot.

The same system also handles lighting. From the white model's spatial information, it can derive physically plausible light direction, color temperature, intensity, and cast shadows.

Your white model: shapes, camera, paths Camera Subject Motion path Camera orbit Model-rendered final video Light direction · temperature · intensity Shadow falls away from light Same geometry: camera, path, and subject stay intact Model adds materials, lighting, and shadows Illustration, not official media; see demo below
The two layers of white-model reference, illustrated by this site. The creator defines the camera, paths, and subject on the left. The generated result preserves that geometry while adding materials, lighting, and shadows.
Full prompt for the white-model demo
参考 @白模 1 的运镜、镜头节奏、景别变化、主体轨迹和镜头调度。参考 @图片 2 的角色外形、场景、材质、光照、色彩和童话氛围,将白模渲染为梦幻温暖、儿童幻想感、3D 动画短片。剧情依次为:幻想天空飞行→云海神兽伴飞→俯冲入海→鳐鱼海底穿梭→镜面时空裂缝→宇宙摘星→幻化回房间→爸爸盖被→合上绘本定格。
Each asset controls a different layer. The white model supplies the camera and motion paths. A second reference image supplies appearance, materials, lighting, and mood. The story is then chained together with arrows.
White-model reference demo: a gray previs sequence becomes a childlike fantasy 3D animated short while retaining the camera choreography blocked out in the white model. Source: Seedance 2.5 launch blog.

Green screen: replace the setting and recalculate how the clothes move

Green-screen editing works with another kind of existing footage: an actor filmed against a green backdrop. The model replaces the green screen with a realistic setting and can also swap obstacles, clothing, and supporting characters. This goes beyond cutting out the subject and pasting in a background. The direction of flowing fabric, the state of the hair, walking rhythm, and the way light falls across the body are recalculated to follow the physics of the new environment, helping the subject and setting fit together.

Green-screen editing demo: the same soccer-training footage is rebuilt in three settings—outdoor practice, with obstacles replaced by rocks, bricks, tires, and crates; a lounge, where friends encourage the player; and an international match, where training poles become defenders and a goalkeeper before the lead scores. Source: Seedance 2.5 launch blog.

White models and green screens point to the same idea: structures prepared in 3D software and footage shot against green screen can feed directly into the next production stage. In its launch post that day, Dreamina AI (@dreamina_ai) paired the release with Maya and Blender plugins aimed squarely at film-grade production pipelines.

Deployment

It is already in classrooms; industrial use is still a list of possibilities

Where is it actually being used? ByteDance offers one education demo and one industrial demo, but the evidence behind them is not equally strong.

Education: turn the world behind a textbook passage into a scene

The model is already part of "Doubao Classroom" in the Doubao AIXUE app. It turns the historical background, characters, and events behind a text into visual scenes, moving lessons beyond static explanation. It can also help teachers produce instructional videos, turning abstract topics such as scientific principles, historical events, and experiments into animated demonstrations. That lowers the cost of creating teaching materials and makes them easier to customize.

A Doubao Classroom scene: an impressionistic view of Lin'an during the Southern Song dynasty. Children run while reciting, "Suddenly I turn—and there he is, where the lantern light is dim." The camera tilts up to Xin Qiji looking back, with a man standing among the distant lights. Source: Seedance 2.5 launch blog.

Industry, robotics, and autonomous driving: generate training footage

Industrial use is different. Generated video could train robots to perceive scenes and manipulate objects. It could also support industrial simulation, process training, and equipment demonstrations. For autonomous driving, it could simulate rare conditions such as extreme weather or complex road situations, providing additional test samples. But every claim in this section is framed as something the model "can" do. No deployed customer, volume, or timeline is provided.

Industrial demo: a white-model car assembly sequence is rendered as a polished, photorealistic production clip. It uses the same white-model capability. Camera movement, composition, shot scale, component placement, assembly order, and motion paths come from the white model; materials, lighting, color, and reflections come from another image. Source: Seedance 2.5 launch blog.
Availability

Where you can use it

The launch blog says Seedance 2.5 is "rolling out progressively." Products in China and overseas moved on the same day.

ChannelStatusAccess / requirements
Jimeng AILaunched that dayWeb → Video generation → Select Seedance 2.5. The platform also offers a separate "long video mode" of up to 3 minutes, according to the Dreamina AI (@dreamina_ai) launch post
Doubao ProLaunched that dayVideo generation → Select Seedance 2.5. PChome reports that Enhanced and Premium plan users receive priority access to the "Creative Video" skill, which can work with models including Seedream 5.0 Pro
Volcano Engine Ark APINot yet availableThe launch blog says it "will also be available soon." It was not open on launch day, and pricing was not announced
Dreamina (international)Launched that dayAvailable to Dreamina subscribers, initially in Southeast Asia, the Middle East, Africa, Europe, and South America, with other regions to follow; Maya and Blender plugins launched alongside it
The image model paired with it in Doubao Pro
ByteDance launches Seedream 5.0 Pro: split one image into more than 10 independent layers and generate dense charts in real scenes
The same team's image-side model launched on July 8. Video needs reference images, and generating those images is part of the same workflow.
🧰 Quick start · Seedance 2.5
PriceNo official pricing. Jimeng AI and Doubao Pro use their own subscription plans. Dreamina deducts credits per generation and is available only to subscribers. The Volcano Engine Ark API is not yet available
Learning curveSelect the model on the web—no deployment required. Good results still require shot-list prompts divided by the second and labeled with asset numbers
Four complete demo prompts you can reuse: the singer's continuous take, breakfast camera editing, assigning roles to 18 concert images, and rendering a fairytale short from a white model. The pattern is to divide the prompt into timed segments, number every asset, and separate camera choreography from appearance references.
Source
One-shot filmmaking, references at will | Seedance 2.5 officially launchedByteDance Seed·seed.bytedance.com·2026-07-31
Site notes
All 10 demo videos come from the Seedance 2.5 launch blog. The original files had unusually high bitrates—up to 51 Mbps at 1080p, totaling 811MB—so this site transcoded them to 126MB and self-hosted the results without changing the visual content. Comparative figures for Seedance 2.0, including 15 reference assets and a 15-second limit, come from its launch blog. "Two video extensions" and "built for professional video creation" come from the Seedance 2.5 project page. "Long video mode: up to 3 minutes" and timestamp precision down to 1 second come from Dreamina's international launch post and are platform-level claims. Doubao Pro plan requirements come from PChome. Seedance 2.0's launch blog confirms that the previous generation already supported video editing and extension. The 90-second row in the chart is this site's inference from "two extensions"; no official source gives that figure. The mapping between the five story beats and the timeline, the two-layer white-model diagram, and the concert-stage layout are all illustrations created by this site.