FLUX 3 Text to Video API

blackforestlabs/flux-3/text-to-video

FLUX 3 Text to Video turns text prompts into 5–20 second high-fidelity video with physically accurate motion, optional synchronized native audio, and eight framing options. It sustains action and camera continuity across the take while aligning dialogue, effects, and ambient sound with on-screen events.

Input
168/20000
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 5 sec × 34/sec = $0.850 (170 credits)
Continue with

Examples

One continuous 5-second cinematic shot. A violinist performs in a sunlit stone cathedral, dust motes drifting through tall shafts of morning light. Camera: slow forward dolly at chest height, no cuts. Lighting: warm volumetric sunbeams against cool stone. Native audio: a rich solo violin melody echoing in the stone acoustic space, faint bow texture. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

One continuous 5-second ultra-wide macro shot. Clear water cascades over mossy basalt steps in a rainforest, droplets scattering off wet stone. Camera: low locked shot two centimeters above the waterline. Lighting: soft canopy-filtered daylight with sparkling highlights. Native audio: close-miked water rush, droplet ticks, distant bird calls. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Vertical 9:16, one continuous 5-second shot. A single tumbleweed rolls across a cracked desert playa at dusk, casting a long shadow. Camera: low tracking shot beside the tumbleweed, steady pace. Lighting: low amber sun with purple sky. Native audio: dry stems scraping hardpan, a steady desert wind, one distant coyote call. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

FLUX 3 Text to Video

FLUX 3 Text to Video generates cinematic video and synchronized native audio directly from text prompts. With accurate physical motion simulation, it supports 5–20 second continuous takes, eight aspect ratios, and 720p / 1080p output for professional storytelling and high-quality production.

Why Choose This?

  • Accurate physical motion simulationRenders fluid dynamics, cloth movement, organic body mechanics, and material reflections with natural real-world behavior.

  • Native synchronized multi-track audioGenerates matching dialogue, foley, ambience, and score cues with the picture, with independent on/off control via sound.

  • Extended 5–20 second generationProduces continuous video from 5 up to 20 seconds in a single run while maintaining camera momentum and narrative continuity.

  • Eight flexible aspect ratiosCovers cinematic 21:9 through vertical 9:16, plus intelligent auto composition matching for multi-platform delivery.

  • Deep camera and timeline adherenceFollows detailed directorial instructions such as dolly moves, pans, orbits, and sequential beat progressions.

  • Crisp 720p and 1080p tiersOffers high-definition output for both rapid creative iteration and final production delivery.

Parameters

ParameterRequirementDescription
promptRequired

String with minLength 1. Directs scene, subject motion, camera movement, and native audio.

durationOptional

Integer. Sets output length from 5 to 20 seconds; the Playground preselects 5 seconds.

Default5
resolutionOptional

String. Sets output resolution; the Playground preselects 720p.

Default720p1080p
aspect_ratioOptional

String. Controls framing; the Playground preselects auto.

Defaultauto21:92:116:94:31:13:49:16
soundOptional

Boolean. Controls whether native synchronized audio is generated; the Playground preselects true.

Defaulttruefalse

How to Use

  1. Write the scene promptDescribe subject action, lighting, camera movement, and the native audio you want in the shot.

  2. Choose output durationSet a length between 5 and 20 seconds based on pacing needs; default is 5 seconds.

  3. Select resolution tierChoose 720p or 1080p depending on turnaround versus visual fidelity; default is 720p.

  4. Select aspect ratioPick 16:9, 9:16, 21:9, auto, or another of the eight framing options to match your distribution layout.

  5. Configure audio generationLeave sound set to true for synchronized audio, or set it to false for silent footage.

  6. Review the cost and runReview the cost shown on the Run button, finish the prompt and settings, then click Run.

  7. Preview and download the videoWhen the task finishes, preview picture and sound in the output panel, then select Download video to save the result.

Pricing

Billed by output video seconds and resolution tier, with native synchronized audio included in the result. 1 credit = $0.005.

UsageRateDetails
720p34 credits/sec ($0.17/sec)Default 5s at 720p is 170 credits ($0.85).
1080p58 credits/sec ($0.29/sec)5s at 1080p is 290 credits ($1.45).

Best Use Cases

  • Cinematic visual previsualizationGenerate concept clips with detailed camera trajectories and lighting to validate shot pacing and blocking.

  • Commercial and promotional creativeProduce visually compelling brand concepts paired with synchronized audio environments.

  • Short-form and social storytellingCreate high-impact vertical content in 9:16 with native dialogue and sound effects.

  • Natural environments and physics dynamicsRender moving water, weather, and particle motion accompanied by matching ambient sound.

  • Audiovisual atmosphere shotsCombine scene and sound direction for rain-soaked streets or warm interiors to craft openers and transitions.

Pro Tips

  • For complex prompts, organize with the CASTLE schema: core summary, scene, subject, dynamic narrative, audio, and style & color.
  • Add explicit audio cues such as “Native audio: gentle footsteps and a soft violin melody” to guide the soundtrack.
  • Specify concrete camera behavior like “slow forward dolly at eye level” to stabilize spatial perspective.
  • Use 5–8 seconds for a single core action; expand to 10–20 seconds for multi-beat sequences.
  • Choose auto when you want framing to follow the composition described in the prompt.

Usage notes

  • FLUX 3 Text to Video is driven by a text prompt, with optional duration, resolution, aspect_ratio, and sound settings.
  • When sound is true, the result includes dialogue, effects, ambience, or score cues synchronized to on-screen events.
  • After an API submission, save the returned task_id to query progress and retrieve the final media URL.
  • Generated clips are delivered as standard video files ready for playback and editing software.

FLUX 3 Text to Video API frequently asked questions

What is the FLUX 3 Text to Video API?

FLUX 3 Text to Video is a Black Forest Labs model for generating video from text prompts. It creates 5–20 second high-fidelity clips with physically accurate motion and optional synchronized multi-track native audio at 720p or 1080p. Built on spatiotemporal generation with joint audiovisual modeling, it follows camera and sound design in the prompt while sustaining physical naturalness and shot continuity. You can call it programmatically or try it from the playground above.

How does FLUX 3 Text to Video control synchronized native audio?

With sound set to true by default, the model synthesizes dialogue, foley, ambience, and score cues in sync with the picture. Add audio directions in the prompt to guide the soundtrack; set sound to false when you want silent video only.

How should complex FLUX 3 Text to Video prompts follow the CASTLE principle?

Organize complex shots with the CASTLE six-part schema: core summary, scene, subject, dynamic narrative (camera and timing), audio cues, and style & color. Keep the subject description consistent, use concrete visible verbs for action and camera moves, and state the full arc first so multi-beat continuity stays stable.

How does FLUX 3 Text to Video keep physical stability across 20-second shots?

The model is trained for continuous dynamics including gravity, momentum, cloth, and rigid-body interactions. For near-20-second runs, keep subject trajectories and camera direction consistent in the prompt, and sequence beats in time order to maintain natural physics through the take.

Which aspect ratio should I choose for FLUX 3 Text to Video?

Choose among auto, 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16. Use a fixed ratio when the delivery channel is known; with auto, framing follows the composition described in the prompt. Full options are listed in the Parameters section on this page.

When should I use 720p versus 1080p with FLUX 3 Text to Video?

720p costs less per second and suits rapid concept iteration; 1080p delivers finer texture and edge detail for final edits. Both are billed by output seconds, so longer clips cost more—see the Pricing section for the exact credit rates.

How do I direct camera continuity in FLUX 3 Text to Video?

In the dynamic narrative, specify starting framing, move direction, and closing composition—for example, “medium tracking shot at eye level, slow push-in at the end.” A single 5–20 second run fits one continuous camera take or action arc; for longer chained storytelling, continue with the Extend Video or Keyframes endpoints listed under Related Models.