Nano Banana 2 Lite Text to Image API

google/nano-banana-2-lite/text-to-image
10 ratios · 5 credits ($0.025) / generation

Nano Banana 2 Lite Text to Image transforms text prompts into native 1K images in approximately 4 seconds, featuring legible in-image text across 25+ languages, granular prompt adherence, and 10 versatile aspect ratios. It maintains balanced lighting and concrete subject fidelity during high-volume creative iteration while delivering cost-effective drafts at just 5 credits ($0.025) per generation.

Get API Key
Input
702
OutputReady
5 credits / generation · $0.025

Continue with

Examples

output.jpg

A cinematic editorial sports photograph of one adult female short-track speed skater carving a tight corner on an indoor ice rink. Extremely low camera close to the ice, full athlete and both skates visible, believable balanced deep lean and athletic anatomy, one gloved hand grazing the ice. Sharp ice crystals spray from the blades. Her plain midnight-blue racing suit has a single coral panel; clear goggles reflect long arena light strips. Keep the athlete crisp while distant spectator stands streak with panning motion blur. Cold silver ice, dramatic directional light, realistic ice texture and an arresting diagonal composition. No advertisements, brands, logos, signatures, text or watermarks.

output.jpg

An extraordinary natural macro photograph of a single pale pink orchid mantis perched on a closed purple artichoke flower bud. Show the entire small insect in a readable three-quarter view: triangular head, two fine antennae, folded raptorial forelegs and petal-like leg lobes with plausible insect anatomy. Tiny dew droplets sit on the tightly overlapping purple-green bracts. Shallow depth of field turns a distant ochre rock wall into a smooth warm background, with no other insects. Soft raking morning light reveals translucent pink edges and intricate organic textures; intimate, restrained, high-end wildlife macro photography. No advertisements, brands, logos, signatures, text or watermarks.

output.jpg

A handcrafted felt stop-motion film still viewed from an elevated three-quarter angle. A single tiny square-headed teal felt robot is repairing a bridge across a narrow deep crack in a wooden worktable. Two parallel full-length yellow pencils span the crack, clearly resting on both banks to form a small bridge. The robot stands safely on the near bank holding the end of one pencil. Oversized colorful buttons are stacked like canyon cliffs on the two banks, and one bright red thread spool waits behind the robot to cross. Coherent miniature scale, tactile wool fibers, stitched seams on the robot, real pencil wood and thread, warm practical workshop lighting and carefully composed depth. The crack and complete bridge remain clearly visible. No advertisements, brands, logos, signatures, text or watermarks.

Nano Banana 2 Lite Text to Image Overview

Nano Banana 2 Lite Text to Image is a speed-optimized, lightweight text-to-image synthesis model engineered by Google. Built as the dedicated high-throughput entry in the Nano Banana family, it renders native 1K high-fidelity visuals in roughly 4 seconds. Featuring reliable typography in over 25 languages and support for 10 native framing formats, it serves as the premier cost-efficient engine for high-volume prototyping, A/B creative testing, and scalable catalog asset production.

Why Choose Nano Banana 2 Lite Text to Image?

  • Ultra-Low Latency in ~4 SecondsDelivers ready-to-use visuals in approximately 4 seconds per image, dramatically accelerating feedback loops and interactive creative sessions.

  • Cost-Effective High-Throughput IterationFixed rate of just 5 credits ($0.025) per generation—roughly 26% lower than Google's base pricing—making large-scale exploration economical.

  • In-Image Typography in 25+ LanguagesRenders legible, well-aligned short text across more than 25 languages directly onto posters, product packaging, and storefront signage.

  • 10 Versatile Native Aspect RatiosGenerates natively in 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, and 21:9 framing to fit every social, mobile, and wide-display specification.

  • Concrete Subject Realism & Coherent LightingAccurately maps textural properties, spatial perspectives, and ambient illumination directly from natural language prompts.

Parameters

ParameterRequirementDescription
promptRequired

Non-blank string, at least 1 character after trimming. No upstream maximum length is specified.

sizeOptional

Output aspect ratio. Select auto to let the model choose the framing; the API request omits size.

Defaultauto1:12:33:23:44:34:55:49:1616:921:9

How to Use

  1. Draft a structured promptDetail your subject, materials, composition, and lighting. If your concept requires text rendered inside the image, enclose the exact wording in straight double quotes.

  2. Select your aspect ratioPick from 10 available aspect ratios (such as 1:1, 16:9, or 9:16) to fit your downstream distribution channel, or omit the parameter to apply standard framing.

  3. Submit asynchronous generation jobSend your API payload to instantly receive a unique task_id, kicking off background rendering across high-speed compute clusters.

  4. Poll for task completionQuery the task status endpoint using your task_id; the generated output is typically ready within 4 seconds.

  5. Download your 1K image assetRetrieve the public image URL from the completed response to download your native 1K asset for review or integration.

Pricing

Text to Image and Image Edit both cost 5 credits ($0.025) per generation. 1 credit = $0.005; no reference-image surcharge.

UsageRateDetails
Text to Image5 credits / generation$0.025 / generation; Google comparison: $0.034 / generation, about 26% less.

Best Use Cases

  • E-commerce catalog variants & product mockupsRapidly prototype dozens of colorways, material finishes, and background contexts before committing to final shoots.

  • Social ad creative A/B testingProduce diverse visual variants with distinct background themes and messaging treatments to evaluate performance benchmarks.

  • UI/UX screen mockups & contextual illustrationsGenerate conceptual spot graphics, hero backgrounds, and application interfaces for digital product roadmaps.

  • Character development & storyboard previsualizationRapidly iterate on character poses, costumes, and narrative scene beats to establish cohesive visual styles early.

Pro Tips

  • Follow the FRAME prompting checklist: Structure prompts across Focus (core subject), Rendering (visual medium), Angle (camera perspective), Mood (lighting tone), and Extras (text and props).
  • Enclose target lettering in quotes: For crisp in-image text on labels, signs, or shirts, wrap the exact phrase in straight double quotes, such as "SUMMER SALE".
  • Leverage high speed for variant exploration: Take advantage of the ~4-second latency and $0.025 price point to test several phrasing variations simultaneously rather than micro-tuning a single run.
  • Specify concrete lighting environments: Ground your visuals by requesting specific setups like "soft studio diffused lighting," "warm golden-hour backlight," or "overcast daylight."

Notes

  • The prompt parameter is required, non-empty, and must contain at least 1 character after trimming leading and trailing whitespace.
  • This endpoint only accepts text input; for image-guided generation or local edits, use the Nano Banana 2 Lite Edit endpoint.
  • Requests operate asynchronously and return a unique task_id for tracking generation progress and fetching the finished image asset.

Related Models

Nano Banana 2 Lite Text to Image API frequently asked questions

What is the Nano Banana 2 Lite Text to Image API?

Nano Banana 2 Lite Text to Image is a high-speed text-to-image synthesis model developed by Google. It produces native 1K high-fidelity visuals from natural language prompts in approximately 4 seconds, featuring legible in-image typography across 25+ languages and 10 native aspect ratios. Engineered with Google’s lightweight multimodal architecture, it faithfully preserves complex spatial and textural directions while enabling cost-efficient high-volume iteration at just $0.025 per generation. You can call it programmatically or try it from the playground above.

How fast is Nano Banana 2 Lite Text to Image generation?

According to official Google benchmarks and live production tests, the model averages approximately 4 seconds per generation (typically 4–6 seconds in real-world API requests). This low turnaround time makes it ideal for interactive creative tools and high-volume batch jobs.

Does Nano Banana 2 Lite Text to Image support in-image text rendering?

Yes. The model is specifically tuned for typographical legibility, producing sharp, well-formed short text in more than 25 languages on signs, merchandise, and labels. Putting the target text inside straight double quotes (e.g., "FRESH COFFEE") ensures optimal character precision.

What output resolution and aspect ratios does the model support?

It produces native 1K (1024px) resolution across 10 native aspect ratios: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, and 21:9. Selecting auto lets the model choose the framing.

How should I decide between Nano Banana 2 Lite and standard Nano Banana 2?

The two models serve complementary workflow stages: Nano Banana 2 Lite is engineered for rapid, low-cost exploration, drafts, and A/B variant testing, whereas standard Nano Banana 2 handles high-resolution final assets up to 4K and complex multi-object compositions.

How is Nano Banana 2 Lite Text to Image billed?

Generations are billed at a flat rate of 5 credits ($0.025) per image, saving roughly 26% compared to Google's standard $0.034 benchmark. Any task that encounters an unexpected processing error receives a full automatic refund.

What is the recommended prompt formula for Nano Banana 2 Lite Text to Image?

Adopt the FRAME methodology by clearly outlining the subject focus, rendering style, camera angle, lighting mood, and in-image text extras. Providing concrete tangible details over generic praise adjectives gives the diffusion engine the clearest guidance.