Seedance 2.5 Reference-to-Video API

bytedance/seedance-2.5/reference-to-video

Seedance 2.5 (Reference-to-Video) generates audio-synchronized videos up to 30 seconds long from text prompts and a mix of image, video, and optional audio references. Assign each reference a clear role to carry subjects, composition, visual style, action, camera movement, pacing, or sound into a new scene, with support for up to 50 references in one generation.

Input

932/20000
Total references0/50 · remaining 50
0/30 · remaining 30
0/10 · remaining 10
0/10 · remaining 10

Add at least one reference image or video. Audio cannot be submitted alone.

Output

Idle

Your generated video will appear here

Configure the required inputs, resolution, and duration, then run the task.

5 sec × $0.315/sec = $1.58

Continue with

Examples

Use Reference Image 1 only for the fictional man's face, short wavy dark hair, navy wool coat, gray scarf, and body proportions. Use Reference Video 1 only for the single turn-and-catch motion, natural weight shift, cloth timing, and slow waist-height camera movement; do not copy the woman, her clothing, or the train platform. Use Reference Audio 1 only for rain intensity and action timing. Create one continuous five-second shot in a quiet covered ferry walkway at night. The man stands alone holding one plain cream envelope in his right hand. A brief gust loosens the envelope; he turns once, catches it against his chest, then becomes still. Preserve his referenced identity and outfit throughout. Synchronized audio: soft rain on the roof, one paper flutter, and distant water ambience. No cuts, no extra people, no dialogue, no writing on the envelope, no readable text, no logos, no products, no advertising, no watermark.

Seedance 2.5 Reference-to-Video

Seedance 2.5 Reference-to-Video combines a text prompt with image, video, and optional audio references to generate a new scene with synchronized sound. Use the prompt to assign each asset a specific role while the model carries subjects, composition, visual style, action, camera movement, pacing, and sound from those references into the generated video.

Why Choose This?

  • Establish appearance with imagesUse image references to guide a character, product, environment, composition, or broader art direction.

  • Communicate motion with videoUse reference video when action, camera movement, blocking, pace, or shot rhythm is easier to show than describe.

  • Guide rhythm and sound with audioUse optional audio to indicate ambience, rhythm, voice character, or the sonic mood you want the scene to follow.

  • Assign every reference a clear roleState which asset controls identity, style, action, camera, or sound instead of asking the model to infer every relationship.

  • Scale complex reference setsCombine up to 30 images, 10 videos, and 10 audio files when one scene needs several distinct sources of direction.

Parameters

ParameterRequirementDescription
promptRequired

String. Defines how each reference contributes to the new scene; 1–20,000 characters after trimming.

reference_image_urlsConditional

String array of up to 30 public image URLs. Uses images to guide identity, appearance, composition, environment, or style.

reference_video_urlsConditional

String array of up to 10 public video URLs. Uses videos to guide action, camera movement, blocking, pace, or shot rhythm.

reference_audio_urlsOptional

String array of up to 10 public audio URLs. Uses audio to guide ambience, rhythm, voice character, or sound direction.

durationRequired

Integer. Sets output length from 4 through 30 seconds, inclusive; the Playground preselects 5 seconds.

resolutionRequired

String. Sets output resolution and must be sent explicitly; the Playground preselects 720p.

720p480p
aspect_ratioOptional

String. Controls output framing; the Playground preselects and explicitly sends auto.

auto16:99:161:121:94:33:4
generate_audioOptional

Boolean. Requests a generated audio track; the Playground preselects true and explicitly sends either value.

truefalse

How to Use

  1. Decide what each asset contributesLabel the creative job first: character identity, product appearance, environment, motion, camera path, or sound direction.

  2. Establish the visual anchorStart with at least one image or video that defines the visible world; add audio when it should guide rhythm or sound.

  3. Add a motion reference when neededUse a video reference for the dancer's movement and camera timing, not as an undefined instruction for the entire result.

  4. Connect every role in the promptWrite the relationship plainly: use the first image for the performer, the video for choreography, and the audio for tempo.

  5. Resolve conflicts before generationRemove references that disagree on identity, camera direction, lighting, or timing before configuring the output.

  6. Configure the outputSet duration, resolution, aspect ratio, and generated audio after the reference relationships are clear.

  7. Generate and reviewRun the request, inspect which reference roles carried into the result, then simplify or relabel competing inputs for the next pass.

Pricing

Without reference video, bill output seconds. With any reference video, output seconds plus every separately floored video duration all use the with-video rate.

UsageRateDetails
480p, no reference video28 credits/output sec ($0.140/output sec)A 5 second image-only task costs 140 credits ($0.700).
720p, no reference video63 credits/output sec ($0.315/output sec)A 5 second image-only task costs 315 credits ($1.58).
480p, with reference video17 credits/billing sec ($0.085/billing sec)A 5 second output plus 2.9 and 3.8 second videos bills 5 + 2 + 3 = 10 seconds, or 170 credits ($0.850).
720p, with reference video38 credits/billing sec ($0.190/billing sec)The same 10 billing seconds cost 380 credits ($1.90); unknown video duration produces only a defensible minimum.

Best Use Cases

  • Recurring character campaign clipsCombine character sheets, wardrobe images, and a scene prompt to create new performance clips for a recurring campaign subject.

  • Choreography and camera studiesPair a character image with a movement video to create a shot study for performance, blocking, camera path, or rhythm.

  • Product campaign variationsCombine product, environment, and style images with new action direction to create campaign video variations.

  • Performance and music conceptsPair performer imagery, choreography video, and an audio reference to create a music or stage-performance concept clip.

Pro Tips

  • Write the role beside each asset in your working prompt: character, wardrobe, product, environment, motion, camera, voice, or rhythm.
  • Replace 'use all references' with a relationship: keep the subject from image one, follow the motion in video one, and use audio one only for tempo.
  • Remove references that compete on identity, lighting, camera direction, or timing before you add more detail to the prompt.

Notes

  • At least one reference image or video establishes the visual input; reference audio can support that visual set.
  • The three reference arrays accept no more than 50 items in total across their individual limits.
  • After upload, every media value must resolve to a public, directly downloadable HTTP(S) URL.

Seedance 2.5 Reference To Video API — Frequently asked questions

What is the Seedance 2.5 Reference-to-Video API?

Seedance 2.5 is developed by ByteDance Seed. The Reference-to-Video API combines a text prompt with at least one image or video reference to asynchronously generate a new video. You can add more image, video, or audio references, up to 50 files in total, and request synchronized audio.

How do I call the Seedance 2.5 Reference-to-Video API?

Send POST /api/generate/submit with a Bearer API key, set model to seedance-2.5/reference-to-video, and place every generation field inside input. A successful submission returns task_id immediately; the API tab and linked documentation include runnable examples.

Open the complete API documentation
How much does the Seedance 2.5 Reference-to-Video API cost?

Without reference video, 480p costs 28 credits/output second and 720p costs 63 credits/output second. With video, output seconds plus each separately floored reference duration use 17 or 38 credits/billing second. For example, a 5-second 480p output with 2.9- and 3.8-second references uses 10 billing seconds and costs 170 credits ($0.850).

What inputs does the Seedance 2.5 Reference-to-Video API accept?

The input object accepts up to 30 images, 10 videos, and 10 audio files, with no more than 50 references in total. Use at least one image or video; audio may accompany those visual references.

How do I get the generated video?

Poll GET /api/generate/status/{task_id} with task_id. On finished, read data.files[].file_url; on failed, stop and read the error. You can also provide callback_url for the terminal result.

Which Seedance 2.5 endpoint should I choose?

Choose Reference-to-Video when different assets need explicit appearance, motion, camera, or sound roles. Use Text-to-Video for language-only creation, or Image-to-Video for a required start frame and optional end frame.