Flux Kontext Max Text to Image API

blackforestlabs/flux-kontext-max/text-to-image
size · output_format

Flux Kontext Max Text to Image transforms demanding prompts into studio-grade visuals with superior prompt adherence and typography. It resolves multi-subject spatial depth while maintaining cinematic balance.

Get API Key
Input
Required0/2000
Optional
Optional
OutputIdle

Generated images will appear here

Estimated cost: 16 credits · $0.080

16 credits ($0.080) per generation.

Continue using

Examples

max-text-to-image-03-output.png

A photograph taken from inside a small deep-sea research submersible, looking through a thick round acrylic porthole at a depth of 1,200 meters. Outside in the black water, a female deep-sea anglerfish hangs close to the glass, seen from the side: a round lumpy dark-brown body, a huge wide mouth full of long translucent needle teeth, and tiny eyes. Growing out of the top of the fish's own head is a thin, flexible fleshy fishing rod, part of its body, that arches forward over its mouth and ends in a small glowing blue-green bioluminescent bulb, the only strong light source outside. The glow lights the fish's teeth and skin from above. Drifting marine snow particles catch the light. On the inside surface of the porthole, faint reflections of the cabin appear: amber-lit analog gauges, a red emergency switch and the soft silhouette of a scientist's face lit by a tablet screen. Tiny condensation droplets bead on the lower edge of the acrylic. A thick steel ring with heavy hex bolts frames the porthole. Extremely low light, high dynamic range, realistic photographic noise. No text, no logos, no watermark.

max-text-to-image-01-output.png

A single hand-colored botanical plate from a nineteenth-century natural history book, photographed flat and straight on in soft daylight. On warm cream paper with faint foxing, a detailed watercolor and engraving of the carnivorous pitcher plant Nepenthes rajah: one large upright urn-shaped pitcher in deep wine red with yellow-green speckles, a ribbed glossy crimson peristome rim around its mouth, a round lid raised above the opening, and a curling tendril joining the base of the pitcher to the tip of a long leathery green leaf. Next to it, a small cutaway of the pitcher shows digestive fluid and a trapped beetle. Delicate engraved line hatching, soft watercolor washes and blooms, fine paper fibers and a slightly uneven plate-mark border. At the top center, one single line of elegant engraved serif capitals reads exactly "NEPENTHES RAJAH". That title is the only text on the page: there are no numbers, labels, captions, signatures, stamps or handwriting anywhere else.

max-text-to-image-02-output.jpg

An ultra-wide cinematic panorama of the Danakil Depression in Ethiopia at dawn. A long camel caravan of about twenty camels, each loaded with stacked rectangular slabs of white salt tied with rope, walks in single file from left to right across a vast, perfectly flat salt pan. Afar herders in white and ochre wraps walk beside the camels holding long sticks. The low sun sits just above the horizon on the right, casting very long blue shadows of the camels across the cracked hexagonal salt crust and turning the ground into a gradient of pink, apricot and pale lilac. A thin layer of standing water in the foreground mirrors the caravan and the sky. Distant dark volcanic hills line the horizon under a clear sky fading from peach to deep blue with a few high streaks of cloud. Shot on an anamorphic lens, low camera height, gentle heat haze, fine detail in the salt texture and camel fur, realistic scale. No text, no logos, no watermark.

Flux Kontext Max Text to Image

Flux Kontext Max Text to Image is the premier flagship model developed by Black Forest Labs, representing the pinnacle of fidelity in the Kontext lineup. Engineered on an expanded flow-matching backbone, it resolves complex multi-layered prompts, multi-subject staging, and intricate physical lighting, rendering dense multi-line typography for high-impact commercial campaigns.

Why Choose Flux Kontext Max Text to Image API?

  • Unrivaled Complex Prompt FollowingAccurately parses intricate, multi-sentence prompts without dropping subtle background elements, specific character costumes, or precise positional blocking.

  • Benchmark Dense In-Image TypographyExcels at rendering complete paragraphs, multi-tiered headlines, and stylized lettering on product packaging and architectural signboards.

  • Studio-Grade Physical Lighting and MaterialsSimulates complex optical phenomena including refractive glass caustics, metallic anisotropy, brushed textures, and subsurface skin scattering.

  • Uniform High Fidelity Across 7 RatiosNatively generates outputs across 7 standard proportions from ultrawide 21:9 to vertical 9:21, holding a constant ~1MP pixel density without distortion.

Parameters

ParameterRequirementDescription
promptRequired

Describe the image to generate using 1–2,000 characters, including at least one non-whitespace character.

sizeOptional

Output aspect ratio. Defaults to 1:1.

Default1:14:33:416:99:1621:99:21
output_formatOptional

Output image file format: png or jpg.

pngjpg

How to Use

  1. Compose High-Precision InstructionsDetail the composition, primary subjects, lighting setup, and materials; enclose headlines, slogans, or body copy in quotation marks to activate typographic rendering.

  2. Set Aspect Ratio and Image FormatChoose the optimal canvas ratio (such as 16:9 for widescreen visual assets or 9:16 for editorial vertical formats) and select png or jpg format.

  3. Submit Request and Access OutputSend the request to receive a unique task_id, check generation progress via polling or webhook callback, and fetch the high-resolution output URL.

Pricing

Each generation costs 16 credits ($0.080) and returns one flagship-grade image.

UsageRateDetails
Standard generation16 credits/generation · $0.080/generationApplies to every size aspect ratio and output_format option

Best Use Cases

  • Luxury Brand Stills and Product ShowcasesGenerate pristine, reflection-accurate studio scenes for high-end chronographs, perfume bottles, fine jewelry, and automotive concepts.

  • Outdoor Billboards and Key Visual PostersProduce visually striking event visuals featuring multi-line typographic lockups and sophisticated architectural perspectives.

  • Cinematic Pre-Production and Concept ArtBring atmospheric worldbuilding and intricate character ensembles to life with authentic optical lens traits and rich tonal palettes.

Pro Tips

  • Format multi-level typography: Clearly distinguish between main titles and secondary lines, e.g., A vintage poster titled "AURORA" with the subline "EXPLORE THE NORTH" underneath in light script.
  • Leverage precise tactile descriptors: Max excels when prompts specify nuanced material finishes like brushed titanium, fluted glass, matte carbon fiber, and direct ambient occlusion.
  • Direct photographic characteristics: Mention specific camera lenses and optics, such as shot on 65mm cine lens, anamorphic streak flares, and delicate f/1.2 bokeh, to unlock authentic cinematic aesthetic.

Notes

  • Text-only generation scope: This endpoint creates new imagery solely from descriptive prompts; to modify existing source images, use Flux Kontext Max Edit.
  • Character boundaries: Prompts must span between 1 and 2,000 characters after trimming leading and trailing whitespace.
  • Fixed pricing schedule: Every generation is billed at a flat 16 credits ($0.080); tasks interrupted by system failures are automatically reimbursed.

Related Models

Flux Kontext Max Text to Image API Frequently Asked Questions

What is the Flux Kontext Max Text to Image API?

Flux Kontext Max Text to Image is Black Forest Labs' flagship model for premium text-to-image generation. It synthesizes ~1MP high-fidelity imagery from complex descriptive prompts, offering industry-best in-image typography, unmatched semantic prompt adherence, and studio-grade physical lighting. Built on an expanded flow-matching transformer architecture, it resolves complex spatial arrangements and fine material reflections while preserving natural tonal balance. You can call it programmatically or try it from the playground above.

How does Flux Kontext Max Text to Image excel at complex prompts?

Flux Kontext Max possesses the most capable text-encoder coupling in the Kontext family. It effortlessly coordinates multi-sentence narratives involving several interacting subjects, layered depth occlusions, and diverse surface materials without omitting requested attributes.

Can Flux Kontext Max Text to Image render dense multi-line typography?

Yes. Max sets the benchmark for in-image lettering. It accurately produces not only isolated words but also structured multi-line text blocks across signage, packaging, and magazines, rendering clean, distortion-free font contours mapped to 3D surfaces.

What is the typical inference latency of Flux Kontext Max Text to Image?

Under standard production workloads, the median generation time is approximately 7 seconds. This modest increase over Pro yields significantly richer micro-textures and more rigorous compositional adherence.

Does Flux Kontext Max Text to Image support custom aspect ratios?

It supports 7 industry-standard aspect ratio presets: 1:1, 4:3, 3:4, 16:9, 9:16, 21:9, and 9:21. Each ratio maintains an optimal ~1 megapixel output resolution, ensuring sharpness across horizontal, square, and vertical orientations.

What is the pricing for Flux Kontext Max Text to Image?

Each completed image generation costs 16 credits ($0.080). This transparent flat rate applies across all aspect ratios and image formats, with balance verification performed upfront and automated refunds for incomplete tasks.

How should creative teams choose between Pro and Max for commercial posters?

Choose Flux Kontext Pro for rapid ideation, standard editorial compositions, and projects prioritizing 5–6s generation speeds. Choose Flux Kontext Max for final campaign key visuals requiring multi-line typographic headlines, complex lighting reflections, and zero-compromise prompt adherence.