FLUX 3 Image to Video API

blackforestlabs/flux-3/image-to-video

FLUX 3 Image to Video turns one start image and a text prompt into 5–20 seconds of high-fidelity video, with realistic physical motion, camera control, and optional synchronized native audio. It faithfully preserves the source subject, clothing detail, and lighting composition while smoothly expanding action, camera moves, and soundscape from your prompt.

Input
152/20000
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 5 sec × 34/sec = $0.850 (170 credits)
Continue with

Examples

The subject turns their head slightly toward the camera with a calm expression while a warm breeze gently moves their collar. Preserve the same face, clothing, window light, and background. Camera: locked medium close-up. Native audio: soft outdoor breeze and quiet room tone. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

A sudden countryside gust lifts the cream sheet and the faded shirt on the line; they billow and slap, then settle. Preserve the same posts, yard, dusk light, and camera. Camera: locked wide shot. Native audio: fabric slap, dry grass wind, distant insects. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Steam rises from the tea cup in slow curling ribbons and drifts toward the open window; dust motes shimmer in the morning light. Preserve the same cup, book, windowsill, and lighting. Camera: locked close shot. Native audio: a faint breeze, a distant kettle tick, quiet room tone. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

FLUX 3 Image to Video

FLUX 3 Image to Video uses a single still frame as the visual starting point and a text prompt to create a continuous clip with optional native audio. It carries forward subject identity, material texture, and ambient lighting from the opening image, with 5–20 second duration, 8 aspect ratios, and 720p / 1080p output.

Why Choose This?

  • Physics-driven motion from a stillAnchor on the subject, lighting, and composition of the start frame, then inject realistic physical motion that brings product art, portraits, or scene stills to life.

  • Native audio-visual syncAutomatically infer and match ambient effects and dynamic soundscapes, with sound toggling native audio on or off.

  • Faithful subject and composition carryoverParse facial structure, wardrobe texture, and spatial lighting so identity and framing stay stable as motion unfolds.

  • Prompt-directed action and cameraDescribe a head turn, fabric sway, push-in, or orbit so the existing image develops around a clear momentum plan.

  • 5–20 second high-fidelity takesGenerate continuous action from 5 to 20 seconds for micro-expressions, product reveals, and narrative blocking.

  • Flexible framing and clarity tiersChoose among 8 aspect ratios including auto, plus 720p and 1080p resolution for multi-channel delivery.

Parameters

ParameterRequirementDescription
promptRequired

String with minimum length 1. Directs subject action, camera movement, visual change, and sound after the start frame.

image_urlsRequired

String array containing exactly one publicly accessible image URL used as the opening frame and visual anchor.

durationOptional

Integer. Sets output length from 5–20 seconds; the Playground preselects 5 seconds.

Default5
resolutionOptional

String. Sets output resolution to 720p or 1080p; the Playground preselects 720p.

Default720p1080p
aspect_ratioOptional

String. Controls framing, supporting auto and common landscape/portrait ratios; the Playground preselects auto.

Defaultauto21:92:116:94:31:13:49:16
soundOptional

Boolean. Controls whether native audio is generated; the Playground preselects true.

Defaulttruefalse

How to Use

  1. Provide a start image URLSupply one publicly accessible image URL with a clear subject and composition in image_urls as the opening baseline.

  2. Describe action, camera, and soundIn prompt, write the subject motion, camera move, and ambient or action audio that should develop from the start frame.

  3. Choose a durationPick an integer length between 5 and 20 seconds, with 5 seconds selected by default, to match the main action and pacing.

  4. Set resolution and framingChoose 720p or 1080p and an aspect ratio for your delivery format, or keep auto to follow the start frame.

  5. Confirm the sound settingKeep sound as true for synchronized native audio, or set it to false for picture-only output.

  6. Review the cost and runReview the cost shown on the Run button, complete the required inputs and prompt, then click Run.

  7. Preview and download the videoWhen the task finishes, preview the video and sound in the output panel, then download the result.

Pricing

Billed by output video seconds and resolution tier, with synchronized native audio included in the result. 1 credit = $0.005.

UsageRateDetails
720p34 credits/sec ($0.17/sec)Default 5s at 720p is 170 credits ($0.85).
1080p58 credits/sec ($0.29/sec)5s at 1080p is 290 credits ($1.45).

Best Use Cases

  • Animated product visualsUse a product still as the start frame and describe camera movement and material reflections for brand presentation footage.

  • Portrait and character animationStart from a portrait and design head turns, expressions, or clothing motion so static characters enter the shot naturally.

  • Concept art and illustration motionAdd wind, drifting clouds, and shifting light to scene art, with matching natural ambient sound.

  • Social vertical contentTurn selected photography into 9:16 short videos with native audio for multi-platform publishing.

Pro Tips

  • Choose a start image with a clear subject outline, lighting, and composition to establish the visual starting point.
  • Describe action that continues from what is already visible, such as “The person raises their cup as the camera slowly moves closer.”
  • State facial features, clothing, or setting details to carry forward in one sentence, then specify camera motion separately.
  • When describing momentum, include start, finish, and pacing—for example, “slowly turns, then holds a gaze toward camera.”
  • Tie sound cues to visible events and setting, such as “Audio: soft outdoor breeze and natural ambient sound.”

Usage notes

  • FLUX 3 Image to Video generates video from one start image and a text prompt, using prompt and image_urls as its main inputs. Duration, resolution, aspect ratio, and sound configure the output.
  • When sound is true, native synchronized audio is included; you can also add ambient or action-sound cues in the prompt.
  • Use publicly accessible HTTP(S) URLs for API media inputs so the service can retrieve the files.
  • Save the task_id returned by an API submission to query progress and retrieve the result.

FLUX 3 Image to Video API frequently asked questions

What is the FLUX 3 Image to Video API?

FLUX 3 Image to Video is a Black Forest Labs image-to-video model. It takes one start image and a text prompt to generate 5–20 seconds of high-fidelity continuous action video, with optional synchronized native audio that automatically matches ambient effects and dynamic soundscapes. Anchored on the still’s subject, lighting, and composition, it smoothly injects realistic physical motion while faithfully preserving the opening frame’s framing, clothing detail, and illumination. You can call it programmatically or try it from the playground above.

How does FLUX 3 Image to Video keep facial and clothing consistency from a still?

The model anchors facial structure, wardrobe texture, and spatial lighting from the start image. Use a clear, detailed source still, emphasize appearance traits to carry forward in the prompt, then separately describe actions such as a head turn or raised hand to guide motion without rewriting identity.

How should I describe action momentum in a FLUX 3 Image to Video prompt?

Continue from what is already visible in the frame and state start, finish, and pacing—for example, “slowly turns, then holds a gaze toward camera as a breeze moves the collar.” Add camera push-ins or orbits so physical inertia and framing change share one timeline.

Does FLUX 3 Image to Video automatically generate matching ambient audio?

Yes. When sound is true, the model infers ambient effects and dynamic soundscapes from the start-frame setting and action cues in the prompt. You can also append specific sound notes at the end of the prompt, such as outdoor breeze or room tone.

How should a still’s framing match FLUX 3 Image to Video aspect_ratio?

Set aspect_ratio to auto to follow the start image’s width-to-height ratio. Choosing a fixed ratio such as 16:9 or 9:16 adapts the composition while keeping the subject centered for your target publishing format.

How does the sound field control native audio in FLUX 3 Image to Video?

sound defaults to true and includes synchronized native audio in the result; set it to false for picture-only output. With audio enabled, add dialogue, ambient, or action-sound cues in the prompt so the soundtrack aligns with visible motion.

When should I pair FLUX 3 Image to Video with end-frame control?

This endpoint focuses on evolving action and camera from a single opening still. If you need to anchor both the start and closing compositions, use FLUX 3 First Last Frame to Video with two ordered frames for dual-anchor transition generation.