Seedance 1.5 Pro Text to Video API

bytedance/seedance-v1.5-pro/text-to-video

Seedance 1.5 Pro Text to Video turns written scenes into videos with synchronized audio and cinematic camera movement. It follows your scene, dialogue and action instructions to build cohesive short clips with expressive characters and coordinated sound.

Input
889/2500
OutputReady
720p · 8 seconds · with audio · 64 credits = $0.320

Examples

Four-second realistic wildlife macro video. Vertical composition, locked camera, tropical rainforest daylight. TWO SEPARATE broad wet leaves are fully visible, one at upper left and one at lower right, separated by a small clear AIR GAP. Exactly ONE red-eyed green tree frog with orange toes crouches at the very tip of the LEFT leaf, facing the RIGHT leaf. The main event is a SINGLE HOP ACROSS THE GAP: within the first second it pushes off with both hind legs; its entire body and all four feet visibly leave the left leaf and travel through the air; it lands once on the separate right leaf by the second second. The right leaf briefly bends under the landing, releasing a few droplets. It then sits still for the remaining time. Show the complete takeoff, airborne body and landing in the same shot. No crawling along a leaf, no extra frog, no transformation. Natural quiet rain with one soft wet-leaf tap at landing; no music, speech, text or logos.

A single continuous four-second cinematic widescreen science-fiction shot inside an unoccupied spacecraft cargo bay in zero gravity. Exactly one ordinary silver open-ended wrench floats freely at the center, slowly translating a few centimeters and rotating smoothly about its long axis. Its rigid metal shape and both ends remain unchanged. A slow short lateral camera slide produces clear parallax between strapped cargo rails in the foreground and a large round porthole behind. The blue curved limb of Earth is visible through the porthole, soft reflected blue light across brushed metal and warm practical lights. All cargo is securely strapped down. Only a quiet steady interior ventilation hum, no wrench impact sound, no explosive effects, no people, speech, music, writing, logos or subtitles.

Seedance 1.5 Pro Text to Video

Seedance 1.5 Pro Text to Video is developed by ByteDance’s Seed team for joint audio-video generation. It interprets written descriptions of scenes, dialogue and movement to create short videos with coordinated voices, environmental sound and camera motion. Creators can establish the setting, direct character actions and choose fixed or moving framing for short dramas, advertising concepts and social storytelling.

Capabilities

  • Scenes from written directionDescribe who is in the scene, what they do and how the camera follows them.

  • Sound and motion togetherGenerate dialogue, ambient sound and action cues together with the video using Generate Audio.

  • Expressive camera movementDescribe a tracking shot, slow push-in or close-up to connect the subject’s action with the visual storytelling.

  • Framing under your controlChoose a frame ratio, resolution and duration, then use Fixed Lens when the scene calls for a stationary camera.

Parameters

ParameterRequirementDescription
promptRequired

Describe the scene, actions, camera movement and sound in 3–2,500 Unicode characters after trimming surrounding whitespace.

aspect_ratioRequired

Choose 1:1, 21:9, 4:3, 3:4, 16:9 or 9:16 for the output frame.

1:121:94:33:416:99:16
resolutionOptional

Choose 480p, 720p or 1080p. Vidgo defaults to 720p when omitted.

Default720p480p1080p
durationRequired

Set the video duration to the integer 4, 8 or 12 seconds.

4812
fixed_lensOptional

Set true to keep the camera fixed, or false to allow camera movement.

truefalse
generate_audioOptional

Set true for synchronized audio or false for a silent video. Vidgo defaults to true when omitted.

Defaulttruefalse

How to use

  1. Describe the sceneWrite the subject, setting and action sequence in Prompt.

  2. Direct movement and soundDescribe the action, camera movement and dialogue or ambient sound you want to hear.

  3. Choose the output settingsSet Aspect Ratio, Resolution and Duration, then choose Generate Audio and Fixed Lens.

  4. Run and reviewReview the displayed cost and select Run. Preview the completed clip or download the video.

Pricing

Each video is billed by resolution, duration and audio setting. One credit equals $0.005. Both endpoints use the same rates.

UsageRateDetails
480p · 4 seconds · silent9 credits / $0.045Per video
720p · 4 seconds · silent16 credits / $0.080Per video
480p · 8 seconds · silent; 480p · 4 seconds · with audio18 credits / $0.090Per video
480p · 12 seconds · silent21 credits / $0.105Per video
720p · 8 seconds · silent; 720p · 4 seconds · with audio32 credits / $0.160Per video
480p · 8 seconds · with audio36 credits / $0.180Per video
1080p · 4 seconds · with audio or silent40 credits / $0.200Per video
480p · 12 seconds · with audio; 720p · 12 seconds · silent42 credits / $0.210Per video
720p · 8 seconds · with audio64 credits / $0.320Per video
1080p · 8 seconds · with audio or silent70 credits / $0.350Per video
720p · 12 seconds · with audio84 credits / $0.420Per video
1080p · 12 seconds · with audio or silent100 credits / $0.500Per video

Use cases

  • Short drama scenesDescribe a conversation, its setting and the characters’ reactions within one clip.

  • Advertising conceptsTurn a written product scenario into a motion study with coordinated sound.

  • Social storytellingChoose a vertical or square frame and build a focused visual moment around one action.

Prompt guidance

  • Write the intended action in chronological order and keep the camera direction consistent throughout the clip.
  • Put spoken words in quotation marks and identify who speaks each line.
  • Describe the environment’s sound as well as the visible action when generating audio.

Related models

Seedance 1.5 Pro Text to Video API frequently asked questions

What is the Seedance 1.5 Pro Text to Video API?

Seedance 1.5 Pro Text to Video API is ByteDance’s model interface for generating videos from written scene descriptions. It combines expressive movement, coordinated sound and camera direction in short clips. Its joint audio-video generation follows scene, dialogue and action instructions, with fixed-camera control for steady framing. You can call it programmatically or try it in the Playground tab.

Can Seedance 1.5 Pro Text to Video generate synchronized audio?

Set Generate Audio to true and describe the dialogue, ambient sound or action cues in Prompt. Audio and video are generated together, allowing sound to follow the scene’s movement.

How should dialogue be written for Seedance 1.5 Pro Text to Video?

Identify the speaker, put each spoken line in quotation marks and describe the tone and pace. The model can generate voices in multiple languages and dialects with lip movement aligned to speech.

How does Seedance 1.5 Pro Text to Video follow an action sequence?

Establish the character and setting, then describe the action in chronological order. Connect each movement to the same subject and use a camera instruction that follows the scene’s progression.

Can Seedance 1.5 Pro Text to Video create 12-second scenes?

Yes. Set Duration to 12 for a longer action or dialogue beat. You can also choose 4 or 8 seconds and match the prompt’s action sequence to that length.

How does Fixed Lens work in Seedance 1.5 Pro Text to Video?

Set Fixed Lens to true for steady framing. Describe movement through the subject’s actions within that frame; set it to false when you want the camera to follow a described move.

Can Seedance 1.5 Pro Text to Video generate 1080p video?

Yes. Select 1080p in Resolution and choose a 4, 8 or 12-second duration. The same duration options also apply to 480p and 720p output.