FLUX 3 First Last Frame to Video API

blackforestlabs/flux-3/first-last-frame-to-video

FLUX 3 First Last Frame to Video anchors both opening and closing still frames to generate 5–20 seconds of coherent transitional video with optional synchronized native audio. Direct motion trajectories, camera movement, and audio between two precise visual compositions.

Input
124/20000
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 5 sec × 34/sec = $0.850 (170 credits)
Continue with

Examples

Begin exactly on the first frame and finish on the last frame. One continuous 5-second locked close-up: the folded paper crane's wings slowly uncrease and lift until they match the end still. Preserve the same crane, paper color, table, window light, and camera. Native audio: dry paper flex, a faint wooden-table creak, quiet room tone. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Begin on the first frame and finish on the last frame. One continuous 5-second studio shot: the dancer rises from the low opening stance into the tall closing pose with one smooth turn. Preserve the same dancer, practice clothes, wooden floor, and soft side light. Camera: locked full-body wide shot. Native audio: bare feet on wood, steady breath, room tone. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Begin on the first frame and finish on the last frame. One continuous 5-second locked wide shot of the same rooftop garden: dawn blue light warms into gold morning sun, plant leaves lift gently in a rising breeze. Preserve the same pots, railing, skyline, and camera. Native audio: morning birds, a distant city hum, soft wind. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

FLUX 3 First Last Frame to Video

FLUX 3 First Last Frame to Video provides precise boundary control over both the start and culmination of a shot. Supplying ordered start and end frames establishes visual targets, allowing the model to compute plausible physical motion, subject transitions, and camera movement over 5–20 seconds with native synchronized audio.

Why Choose This?

  • Dual anchor composition controlFirmly pins down both opening pose and closing composition, guiding the video toward your intended final frame.

  • Smooth physical interpolationLeverages deep motion dynamics to generate organic displacement, material deformations, and fluid transitions between two still baselines.

  • Prompt-guided transition choreographyUse text instructions to guide intermediate details such as camera pacing, turning points, and progressive lighting changes across the clip.

  • Native synchronized sound synthesisGenerates audio that follows the arc of movement, synchronizing footfalls, ambient acoustic shifts, and music to visual progression.

  • Configurable 5–20 second pacingAllows fine-tuning generation length from 5 to 20 seconds, accommodating both brisk visual cuts and slow, contemplative morphs.

  • Pristine 720p and 1080p deliveryMaintains crisp sharpness and textural depth from opening frame to closing frame across 720p and 1080p resolution choices.

Parameters

ParameterRequirementDescription
promptRequired

String. Directs motion, camera trajectory, transitional behavior, and native audio connecting the two frames.

image_urlsRequired

String array containing exactly 2 public image URLs. Item 0 is the start frame, and item 1 is the end frame.

durationOptional

Integer. Sets output video duration in seconds from 5 to 20; the Playground preselects 5 seconds.

Default5
resolutionOptional

String. Sets output resolution to 720p or 1080p; the Playground preselects 720p.

Default720p1080p
aspect_ratioOptional

String. Controls framing ratio, supporting auto and standard formats; the Playground preselects auto.

Defaultauto21:92:116:94:31:13:49:16
soundOptional

Boolean. Controls whether native synchronized audio is generated alongside video; default is true.

Defaulttruefalse

How to Use

  1. Supply start and end image URLsPrepare publicly accessible URLs for the opening still and closing still, ordered in image_urls as [start_url, end_url].

  2. Describe transition and camera behaviorIn prompt, detail the movement trajectory connecting both states, specifying camera maneuvers and audio cues.

  3. Set duration and timingSelect an output duration between 5 and 20 seconds depending on whether the transition requires rapid or gradual pacing.

  4. Choose resolution and aspect ratioSelect 720p or 1080p resolution and configure an appropriate aspect ratio or leave set to auto.

  5. Configure audio synthesisLeave sound set to true to synthesize matching transitional audio, or toggle to false for silent output.

  6. Generate and review transitionConfirm the estimated credits, click Run, and inspect the continuous motion linking the two images before downloading.

Pricing

Billed per output second by resolution tier, including native synchronized audio. 1 credit = $0.005.

UsageRateDetails
720p34 credits / sec ($0.17 / sec)Default 5s at 720p is 170 credits ($0.85).
1080p58 credits / sec ($0.29 / sec)5s at 1080p is 290 credits ($1.45).

Best Use Cases

  • Storyboard scene transitionsConnect discrete storyboards by computing natural continuous camera movements between key compositions.

  • Product transformation revealsIllustrate product state changes, unboxing, or mechanical assembly by setting closed and open stills as boundary anchors.

  • Character posture choreographyDefine start and end stances for characters, allowing the model to synthesize balanced physical locomotion.

  • Seamless looping animationsPass identical opening and closing frames to create continuous, cyclic dynamic loops.

Pro Tips

  • Keep subject appearance, wardrobe, and ambient perspective logically consistent across both stills for seamless interpolation.
  • Focus prompt instructions on how the subject transitions, such as 'subject stands up smoothly and steps toward the window'.
  • For large compositional shifts between frames, set duration to 8 seconds or longer to allow natural deceleration and acceleration.
  • To generate looping clips, set the exact same image URL in both slots and instruct subtle cyclic movement in prompt.
  • Highlight audio shifts across the motion in the prompt, such as 'starts with subtle rustling and ends on a solid closing thud'.

Notes

  • FLUX 3 First Last Frame to Video requires an image_urls array containing exactly two public image URLs in chronological sequence.
  • Frame 0 is locked to the opening still, while the final frame is locked to the end still, with intermediate frames synthesized.
  • After submission via API, record the returned task_id to poll progress and download the completed asset.
  • Rendered videos are output in standard MP4 format and can be dropped directly into timelines to link neighboring shots.

FLUX 3 First Last Frame to Video API frequently asked questions

What is the FLUX 3 First Last Frame to Video API?

FLUX 3 First Last Frame to Video is a Black Forest Labs model for synthesizing continuous video transitions between two designated still frames. Using user-supplied start and end stills as fixed composition targets, it generates 5 to 20 second video clips up to 1080p resolution with optional synchronized native audio. Built with deep motion interpolation capabilities, it preserves subject identity, spatial perspective, and surface textures while computing physically coherent camera trajectories and movement between both poles. You can call it programmatically or try it from the playground above.

How are images ordered in FLUX 3 First Last Frame to Video?

The image_urls array must contain exactly two items in chronological sequence: index 0 serves as the opening frame, and index 1 serves as the closing destination frame. The model generates forward motion from the former to the latter.

How should prompts guide the transition between frames?

Prompts should focus on how the subject moves and how the camera transitions between states. For example, 'subject stands up smoothly and steps to the desk while the camera dollies right' gives the interpolation engine explicit guidance for plausible motion paths.

How do I create seamless looping animations with FLUX 3 First Last Frame to Video?

Supply the exact same image URL as both index 0 and index 1 in image_urls, and describe a continuous cyclic action in prompt (such as a 360-degree orbit or subtle atmospheric motion). The clip will conclude seamlessly where it began.

Does FLUX 3 First Last Frame to Video generate transition audio?

Yes. When sound is true, the model evaluates visual changes across both frames alongside prompt instructions to synthesize matching dynamic sound effects and ambient progression.

What duration is recommended for wide differences between frames?

If there is significant displacement or a complex posture shift between the start and end images, setting duration to 8–15 seconds gives the physics engine sufficient time to render natural acceleration and deceleration.

Which endpoint should I use to direct more than two keyframe points?

This endpoint specializes in dual-anchor interpolation. To position up to 10 keyframes along a 24 fps timeline, use the FLUX 3 Keyframes to Video endpoint for comprehensive multi-shot storyboard directing.