Skip to main content

Get 2× Credits on All Paid Plans and Credit Packs

Guide
2026/08/03

Seedance 2.5 Prompt Guide: The Official Formula ByteDance Uses (2026)

Seedance 2.5's API opens August 7. This is the official prompt formula, the reference-role template, the sound and subtitle syntax, and the material limits — all from ByteDance's own practice guide.

Xiaosanpao, author at Tryonr — AI virtual try-on

Author

Xiaosanpao

Table of Contents

18

Seedance 2.5 Prompt Guide: The Official Formula ByteDance Uses (2026)

Seedance 2.5 Prompt Guide: The Official Formula ByteDance Uses

I've generated over 200 Seedance videos. For most of that time I had a problem I couldn't explain.

The same prompt that produced a clean 5-second clip would fall apart at 15. Characters drifted. The camera wandered off. The lighting changed halfway through for no reason at all.

I assumed it was the model. It wasn't. It was me. I was writing 5-second prompts and stretching them.

Then ByteDance published its enterprise practice guide for Seedance 2.5 — the document its own commercial teams use — and it explained every one of those failures on the first read. Not with vague advice about "being descriptive." With an actual formula, a template for assigning jobs to your reference material, and hard limits on how much you can feed the model before it starts diluting what matters.

Seedance 2.5's API opens August 7, 2026. Here's what's in that guide, most useful parts first.

Why Seedance 2.5 Prompts Break If You Write Them Like 2.0

Two things changed in 2.5, and both of them break old prompting habits.

The clip got twice as long. Seedance 2.0 topped out at 15 seconds. Seedance 2.5 generates 30 seconds in a single shot with no stitching. That sounds like a pure win, and it is — but a prompt written as one flowing sentence doesn't have enough scaffolding to hold 30 seconds together. The model fills the gap with its own choices, and those choices drift.

The reference budget exploded. Seedance 2.0 accepted 9 images and 3 audio/video clips. Seedance 2.5 accepts 50 references: 30 images, 10 videos, 10 audio clips. More room is only an advantage if you tell the model what each file is for. Upload eight images with no instructions and the model has to guess which one is the character, which is the background, and which is just style.

Everything below is about those two problems.

The Official Seedance 2.5 Prompt Formula

ByteDance's guide gives one formula. Six parts, and you can drop any part you don't need:

Subject + Action or event + Setting and environment + Visual style + Camera work + Sound

Here's what each part is actually for:

  1. Subject + action — who or what is doing what. This is the floor. Summarize the main process first, then add detail only to the moments that matter. Don't describe the same action three different ways.
  2. Setting and environment — place, time of day, weather, spatial relationships, background state.
  3. Visual style — light, color, material, texture, overall mood.
  4. Camera work — shot size, camera position, movement, focus target, how cuts connect.
  5. Sound — dialogue, voice character, ambient sound, effects, music.

And the template that formula turns into:

<Subject> does <main action or event> in <setting and environment>.
The image is <visual style>.
The camera uses <shot size, position, movement, or cut>.
Sound includes <dialogue, ambience, effects, or music>.

A worked example, adapted from the official guide:

A ceramicist finishes a pale blue cup in her studio at dawn, lifting it off
the wheel and setting it at the center of a wooden rack.
Soft morning light comes through the window; the wet clay has a fine sheen
and the workbench stays tidy.
The camera records the throwing in a medium shot, pushes slowly into the
surface texture of the cup, then cuts to the front of the rack.
Keep the low hum of the wheel, the friction of the clay, and light room tone.

Notice what isn't in there: no resolution, no frame rate, no aspect ratio. Generation parameters don't belong in the prompt. Set those on the generation page or in the API call. Writing them into the prompt just adds noise the model has to ignore.

How to Write a 30-Second Prompt: Split It Into Beats

Seedance 2.5 30-Second Single Shot Prompt Structure Shown as a Continuous Film Strip of Timed Beats

This is the single highest-leverage change you can make.

Don't write 30 seconds as one idea. Write it as timed beats with explicit timestamps. Every long-form example in ByteDance's guide is written this way — the science explainers, the ad spots, the multi-room tracking shots. All of them.

The pattern:

0-3s   [what's on screen] [what the camera does] [dialogue or sound]
3-8s   ...
8-14s  ...

Here's a condensed science explainer, adapted from the official guide:

Create a science video explaining how the Moon formed. Cinematic quality,
strong visual impact, scientifically accurate.

0-2s: Deep space wide shot. Earth's curve and blue atmosphere fill the lower
frame; a dark red body hangs in the distance. Camera nearly static.

2-3s: Cut to a closer view. A huge body, its surface cracked with orange
lava, approaches Earth's limb. Camera static, the body feels heavy and close.
Narration: "4.5 billion years ago, Earth was struck by a body the size of Mars."

3-5s: Impact at the limb. A white-orange flash erupts from the contact point,
debris and fire spray outward along the surface. Camera pushes in with the
blast.

5-8s: The superheated material gathers into a glowing molten sphere, debris
and dust orbiting around it. Camera pulls back slowly.

Three things make this work:

  • Timestamps force pacing. Without them the model decides how long each idea lasts, and it decides badly.
  • Each beat names the camera. Not just what's in the shot — what the camera is doing during that beat.
  • Dialogue is attached to a beat, not floating at the end of the prompt.

If you take one thing from this guide, take this one.

The timed-beats prompt above, generated. Each timestamp becomes a beat you can see.

Tell the Model What Each Reference Is For

Seedance 2.5 Reference Roles Interface Showing Images Being Assigned to Subject Scene and Motion Slots

You reference material with the @ symbol — @image1, @video1, @audio1. That part everyone knows.

What almost nobody does is the part that actually matters: every reference needs a sentence saying what it provides. The guide is blunt about this. Don't write "reference @image1" with no head or tail. And don't rely on labels written inside the image, or expect the model to work out on its own which file maps to which character.

The template:

@image1 provides <subject>'s <appearance, clothing, structure, or material>.
@video1 provides <the action, camera movement, or rhythm>.
@audio1 provides <the voice, dialogue, ambience, or music> for <who>.

<Subject> does <main action> in <setting>.
The image is <visual style>; the camera uses <camera work>.

And the same thing filled in:

@image1 provides the ceramicist's face, hair, and dark green apron.
Do not use the background from this image.
@image2 provides the studio's wooden bench, window position, and morning
light. Do not use the person in this image.
@video1 provides the rhythm of the hands throwing, lifting, and placing the
cup. Do not use the identity, clothing, or location from this video.

Those "do not use" lines are doing real work. A reference image carries everything in it — subject, background, framing, color. If you only want the jacket, say so, or you'll get the room it was photographed in.

One more case worth memorizing. When several images are different angles of the same object, say that explicitly:

@image1 defines the front of the folding lamp.
@image2 defines the left side of the same folding lamp.
@image3 defines the right side of the same folding lamp.
All three images define one single folding lamp. There is only ever one
folding lamp in the finished video.

Skip that last line and you can end up with three lamps.

The Syntax for Music, Sound Effects, Dialogue, and Subtitles

Plain natural language works fine. But when you need to separate music from sound effects from dialogue from on-screen text, Seedance has dedicated characters for it. This mapping comes straight from ByteDance’s official practice guide:

What you wantCharacterExample
Music( )(soft piano plays in the background)
Sound effect< ><a bell rings in the distance>
Dialogue{ }{Hello, welcome back}
On-screen subtitle【 】【Chapter One: Departure】

You can also turn sound off explicitly, which is more reliable than hoping:

No background music. Keep only dialogue, ambience, and action sound effects.
No subtitles.
No sound at all.

For non-English dialogue, name the language before the line. Seedance 2.5 natively supports more than ten languages — Chinese, English, Spanish, Indonesian, Malay, Thai, Arabic, Portuguese, Vietnamese, Japanese, Korean — and the guide's recommended formula is:

Language + regional variant or accent + delivery + speaker + {line}

Language: American English. The girl says, in a natural, conversational
American accent: {I thought you weren't coming.}

Language: Los Angeles American English. The young man says, casually:
{No way, you actually made it.}

This is also the fix for the most common complaint I see — you write an English line and the model delivers it in another language. Name the language explicitly and the problem goes away.

Negative prompts work too. The official interior-design example ends with a full negative block:

Negative Prompt: Do not change the room layout. Do not move furniture.
Do not add or remove objects. Do not change the camera movement.
No people, no text, no logo, no flickering.

Why Your Character Drifts: The Reference Ratio Rule

Seedance 2.5 Reference Ratio Showing Dozens of Image and Audio References Converging Into a Single Consistent Output

This is the section that answered my original problem, so I'll give it the space it deserves.

Seedance 2.5 accepts 50 references. That is a ceiling, not a target. ByteDance's own evaluation found that past a certain point extra material makes results worse, not better — key features get diluted and noise increases.

The numbers:

If your subject references are…RecommendedStretch
Audio or video≤ 5 subjects6–10 if you must
Images≤ 8 subjects9–12 if you must

Clip length matters as much as clip count. Keep individual video and audio references to 5–10 seconds. That's the sweet spot for recognition and stability. A 2-second clip doesn't carry enough information for the model to lock onto. A 30-second clip dilutes the features you care about and drags in noise.

Single view vs multi view:

  • At 5 subject images or fewer, either works.
  • Past 5 subjects, single-view images are more stable.
  • If you need multiple angles, upload them as separate images. Do not upload one image containing a grid of angles.

And when you're over budget, drop things in this order:

Core characters > key products and props > setting and environment > overall style

Style is the first thing to cut. It's also the thing people protect first, which is exactly backwards.

For completeness, the hard limits:

CountSizeFormatDuration
Images0–30≤30MB, up to 4Kjpeg, png, webp, bmp, tiff, gif, heic, heif
Video0–10≤200MB, 480p–4Kmp4, mov2–30s each, ≤30s total
Audio0–10≤15MBmp3, wav2–30s each, ≤30s total

Video and audio budgets are counted separately — 30 seconds of each, not 30 seconds combined. And you can mix modalities seven ways: images only, video only, audio only (new in 2.5), or any combination of the three. When you use all three, you can also designate which reference image serves as the first or last frame.

8 Prompts You Can Copy

All eight are adapted from ByteDance's official practice guide. Swap the subjects and locations to fit your project.

1. Wordless 30-second narrative

A 30-second wordless visual story with the quiet elegance of the French
countryside, on a summer afternoon in Provence. A young woman in an off-white
linen dress finishes a handwritten letter at an old wooden table inside a
stone cottage, walks outside, follows a gravel path through a sunflower field,
and reaches a vintage yellow postbox at the edge of the village, slipping the
letter inside. The camera opens on a shallow-focus close-up of the tabletop,
moves through the doorway as light gives way to shadow, then rises and pulls
back into a wide golden valley. Ambient sound only — cicadas, wind, pen on
paper, and the soft click of the postbox closing.

Why it works: the whole route is described as one continuous path, and the camera has an explicit arc — close-up, through the door, rise and pull back. That's what keeps 30 seconds from wandering.

2. Product commercial, studio into lifestyle

A 30-second vertical furniture commercial. The product is a pale sage modular
sofa with rounded lines, low deep cushions, an off-white throw, and brown
cushions.

First 15 seconds — pure white seamless studio. The sofa is shown complete;
the camera pushes in slowly, tracks horizontally, then pans across the
silhouette before pulling back to a locked wide shot. Even soft lighting,
clean white background, no light sweeps, no white flashes.

Last 15 seconds — a warm-lit living room in the evening. A woman in off-white
knitwear reads on the sofa, the throw over her legs; she sets down a coffee
cup, a cat jumps up, a child runs in and leans on her shoulder. A floor lamp
comes on; the three of them settle quietly. The camera pulls back slowly and
holds, leaving room for a logo.

Why it works: two halves with an explicit hinge. The model isn't guessing where the tone changes.

3. Fashion film with multiple references

Generate a 30-second one-take video. Six models enter one after another; the
camera travels continuously and the styling changes on the beat.
@image1 wearing @image2 in the setting of @image3, sound reference @audio1.
@image4 wears @image5, in the hat from @image6, walking out with a coffee,
background @image7. @image8 poses in the setting of @image9, wearing @image10
and @image11, sound reference @audio2.

Why it works: every reference has a stated job. Nothing is left to inference.

4. 3D animated ad, beat by beat

3D animated commercial. Bright, translucent color; the fruit and juice should
feel intensely refreshing. High-end commercial animation with a touch of
exaggerated humor. The desert lizard character is cute and expressive.

0-3s: A desert baking under the sun. The air distorts with heat. A horned
lizard lies flat on scorching sand, tongue out, eyes glazed. Sound: rolling
heat, a faint dry crackle.
3-6s: The lizard stops. Its nose twitches. Buried in the sand is a cold,
plump grapefruit beaded with condensation. Sound: a bright discovery chime.
6-8s: The lizard dives at it, hugs it with both arms, face pressed to the
peel, blissful. Hold for one second. Sound: a thud, then half a second of
silence.
8-11s: The peel splits. The flesh inside is luminous. Juice doesn't trickle
out — it erupts. Sound: a crisp bite, then an exaggerated burst.

Why it works: sound is specified per beat. That's what makes the comic timing land.

5. Keyframe-driven concept film

Cinematic brand concept film. @image1 is the first frame. The image trembles
slightly; the camera pushes in toward tree shadows racing past a window,
faster and faster, then cuts abruptly to @image2 where the speed drops away
and the camera drifts forward along a stream. The camera descends underwater
— bubbles in the audio — and a group of orange jellyfish drifts past the lens
@image3. The camera pulls back slowly.

Why it works: first frame is declared, and each cut is anchored to a specific reference.

6. Edit: remove a subject

Video edit: remove everyone from @video1 except the lead.

That's the whole prompt. Editing commands are supposed to be short.

7. Edit: relight the whole film

Change the lighting in @video1 from harsh midday overhead light to
low side light at sunset.

Why it works: it names the current state and the target state. "Make it prettier" gets you nothing.

8. Edit: change a performer's age and expression

Keep the composition, camera position, lighting, and performance rhythm of
@video1. Only rewrite the lead actress's appearance and expression: age her
naturally from her mid-twenties to sixty, letting the restraint in her eyes
soften, a tear cross the corner of her eye, the corner of her mouth lift,
and finally break into a smile through the tears. One continuous take, no
cuts, no flicker; features shift with age without drifting.

Why it works: it locks what must not change before it says what should. That's the edit pattern in one line — and there's a lot more to it, which I've written up separately in the Seedance 2.5 video editing guide.

Seedance 2.5 Prompt FAQ

How many references does Seedance 2.5 support? Fifty: 30 images, 10 videos, and 10 audio clips. Individual videos and audio clips run 2–30 seconds, with a 30-second total for video and a separate 30-second total for audio.

Does Seedance 2.5 support negative prompts? Yes. ByteDance's own interior-renovation example ends with a full negative block covering layout, furniture positions, camera movement, and unwanted artifacts.

What languages does Seedance 2.5 support? More than ten natively, including Chinese, English, Spanish, Indonesian, Malay, Thai, Arabic, Portuguese, Vietnamese, Japanese, and Korean. Name the language before the dialogue line for reliable delivery.

Can Seedance 2.5 edit a video I already have? Yes. Feed the finished video in as @video1, describe the change, and add a line locking everything else. The output keeps the original aspect ratio and roughly the original duration.

How long can a Seedance 2.5 clip be? Thirty seconds in a single shot, with no stitching — double Seedance 2.0's 15-second ceiling.

How much does Seedance 2.5 cost? There's no free tier for the model itself, and access opens August 7, 2026. I've broken down the pricing and what actually changed from 2.0 in Seedance 2.5 vs 2.0.

The Bottom Line

Three things separate a Seedance 2.5 prompt that works from one that doesn't.

Write 30 seconds as timed beats, not as one sentence. Timestamps, and a camera instruction inside each beat.

Give every reference a job. One line per file saying what it provides — and what it doesn't.

Stay under the ratio. Five audio/video subjects, eight image subjects, 5–10 seconds per clip. Fifty references is a ceiling, not a goal.

Seedance 2.5 opens on August 7. You don't have to wait to practice — the formula, the reference-role template, and the syntax table all work on Seedance 2.0 today, and you can generate with Seedance 2.0 on Tryonr right now with no watermark. Get the habits right on 2.0 and you'll be productive on 2.5 the day it opens.

Still working on 2.0? Start with the Seedance 2.0 prompt guide. Want the full picture of the model itself? That's on the Seedance 2.5 model page.

Prompts in this guide are adapted from ByteDance's official Seedance 2.5 practice guide.

More Posts

Get e-commerce photography tips & Tryonr updates — join 300+ sellers

Join the community

Be the first to hear about new AI features, best practices for product photos, and real case studies from other online sellers.