Nano Banana Text to Image API

google/nano-banana/text-to-image
10 ratios · 5 credits ($0.025) / generation

Nano Banana Text to Image transforms text prompts into high-fidelity visuals with legible typography and 10 versatile aspect ratios. It accurately maps complex spatial prompts while delivering naturally balanced lighting.

Get API Key
Input
485/5000
OutputReady
5 credits / generation · $0.025

Continue with

Examples

text-02-output.jpg

Extreme close-up photograph of a clockmaker's hands assembling a brass gear train under a warm bench lamp. Fine steel tweezers hold one small gear so its teeth mesh cleanly with the neighboring wheel; the rest of the train sits in an open wooden movement frame. Accurate mechanical contact, visible tool steel, brass wear, a drop of oil on a pivot, wood grain, shallow depth of field and workshop bokeh. Show only the hands and the mechanism, no face, no text, no logos, no watermarks.

text-01-output.jpg

Photorealistic documentary photograph of a train-station lost-and-found glass window in late afternoon. Two centered lines of white painted capital letters on the glass: the first line is exactly LOST and the second line is exactly FOUND. No ampersand, no extra letters, no other words on the glass. A small paper card taped inside the lower right corner shows only the time 8:00-18:00. Behind the glass, on one wooden shelf: a brass tuba lying on its side, a folded red umbrella, and a pair of white ice skates. Soft concourse reflections, warm side light, realistic glass and wood. No advertisements, brands, logos, signatures, or watermarks.

text-03-output.jpg

A behind-the-screen rehearsal photograph of a traditional shadow-puppet stage. In the foreground, two adult puppeteers hold a deer puppet and a boat puppet against a taut rice-paper screen; an oil lamp throws sharp silhouettes of the deer and the boat onto the paper. The carved leather puppets and the puppeteers' hands stay visible on the lamp side, while the silhouettes read clearly on the screen. Warm amber practical light, layered space, empty room, no audience. No text, no logos, no watermarks.

Nano Banana Text to Image

Nano Banana Text to Image is a high-speed text-to-image model developed by Google DeepMind, powered by the Gemini 2.5 Flash architecture. It excels at complex prompt interpretation, photorealistic lighting, crisp typography rendering, and real-world knowledge integration. Delivering 1K/2K visuals across 10 native ratios, it accelerates rapid creative prototyping.

Why Choose Nano Banana Text to Image?

  • Google DeepMind Native Multimodal ArchitectureTrained jointly on text and visual concepts within the Gemini 2.5 Flash Image framework, delivering robust world reasoning and faithful execution of intricate concepts.

  • Sharp In-Image Typography and SignageOvercomes conventional AI spelling artifacts by directly rendering legible, sharp, and well-aligned textual elements on product labels, posters, and urban signs.

  • Grounded Real-World Knowledge and LandmarksDraws upon deep factual intelligence to accurately depict authentic landmarks, geographical locations, and iconic industrial designs without anatomical distortion.

  • Native Support for 10 Aspect RatiosOffers 10 native framing formats including 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, and 21:9, ensuring composition precision across mobile and desktop displays.

  • Cost-Efficient High-Speed IterationStandard generations cost only 5 credits ($0.025), pairing ultra-low per-call expense with rapid throughput for expansive prototyping and A/B design testing.

Parameters

ParameterRequirementDescription
promptRequired

Describe the image to generate using 1–5,000 characters, including at least one non-whitespace character.

sizeOptional

Output image aspect ratio. Select auto to let the model choose the framing; the API request then omits size.

Defaultauto1:12:33:23:44:34:55:49:1616:921:9

How to Use

  1. Draft a structured text promptDescribe the central subject, visual texture, camera framing, and ambient lighting. If specific wording must appear in the visual, enclose the text in straight double quotes.

  2. Select the target aspect ratioChoose from 10 supported aspect ratios like 1:1, 16:9, or 9:16; choose auto to permit the model to adaptively determine framing based on your prompt.

  3. Initiate the asynchronous generation taskSubmit your POST payload to the API endpoint to receive an immediate task_id while backend GPU nodes perform rapid diffusion inference.

  4. Poll for task completionQuery the task status endpoint using your task_id, or specify a callback_url at the root level to receive an automated webhook when the task finishes.

  5. Download high-fidelity visualsRetrieve the rendered image URL from the completed response object to integrate directly into design layouts or web experiences.

Pricing

Each generation costs 5 credits ($0.025). 1 credit = $0.005.

UsageRateDetails
Text to Image5 credits / generation$0.025 / generation

Best Use Cases

  • E-Commerce Studio Shots and Packaging ConceptsRapidly generate crisp white-background studio images, situational product staging, and packaging mockups to accelerate catalog ideation.

  • Commercial Advertising and Poster LayoutsCreate cohesive promotional visuals featuring legible headline typography and striking lighting for social media banners, hero images, and cards.

  • Authentic Landmark and Travel ImageryLeverage built-in knowledge to render faithful architectural landscapes, historic locations, and travel scenes for editorial or travel publications.

  • Rapid Pre-Production Concept ExplorationTest multiple compositional angles, lighting conditions, and aesthetic moods with rapid iteration speeds and cost-effective pricing before final asset sign-off.

Pro Tips

  • Structure prompts with the FRAME methodology: Organize descriptions by Focus (subject), Rendering (medium/style), Angle (perspective), Mood (lighting/tone), and Extras (details) to ensure complete visual clarity.
  • Enclose in-image text in straight double quotes: When generating logos, shop signs, or product text, wrap target words in quotes (e.g., "SPECIAL COFFEE") to maximize character legibility.
  • Specify concrete illumination and surface shaders: Incorporate distinct phrases such as "soft studio softbox lighting" or "matte textured plastic" rather than vague buzzwords like "photorealistic".
  • Explore with adaptive ratios before finalizing: Use the default auto ratio during early creative exploration, then lock in exact dimensions like 16:9 or 9:16 once the concept is validated.

Notes

  • The prompt field is mandatory, requiring 1 to 5,000 characters with at least one non-whitespace character after trimming.
  • The Text to Image endpoint generates visual assets exclusively from text instructions; to supply visual reference photos, switch to the Nano Banana Edit endpoint.
  • Task execution is fully asynchronous, returning an immediate task_id; any request failure caused by system interruption or validation errors triggers an automated, full refund of deducted credits.

Related Models

Nano Banana Text to Image API frequently asked questions

What is the Nano Banana Text to Image API?

Nano Banana Text to Image is a Google DeepMind model for generating high-fidelity images from natural language text prompts. It produces high-resolution images with legible in-image typography rendering, deep real-world landmark understanding, and native support for 10 aspect ratios. Built on the Gemini 2.5 Flash Image multimodal architecture, it strictly follows complex spatial and lighting prompts while maintaining natural depth and balanced textures. You can call it programmatically or try it from the playground above.

Does Nano Banana Text to Image support in-image typography?

Yes. The model is specifically optimized for visual text rendering across multiple languages, accurately placing legible short phrases on posters, street signs, and packaging. Wrapping requested text in straight double quotes (e.g., "SUMMER SALE") ensures the highest spelling accuracy and crisp letterforms.

Can Nano Banana Text to Image accurately depict real-world landmarks?

Yes. Powered by Google's comprehensive world knowledge base, Nano Banana recognizes and accurately reconstructs recognizable landmarks such as the Eiffel Tower or Mount Fuji, ensuring structural authenticity without spatial distortions.

What aspect ratios does Nano Banana Text to Image support?

The model natively supports 10 aspect ratios: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, and 21:9. If you omit the size parameter or select auto, the model automatically selects an optimal framing format tailored to your prompt description.

How should I choose between Nano Banana and Nano Banana 2 Lite?

The models serve complementary roles: Nano Banana (Gemini 2.5 Flash) emphasizes deep prompt fidelity, factual landmark precision, and balanced photorealism, whereas Nano Banana 2 Lite is engineered for maximum speed (~4s generation) in latency-sensitive interactive prototyping.

How much does Nano Banana Text to Image cost?

Each generation costs a fixed 5 credits ($0.025, with 1 credit = $0.005) regardless of the chosen aspect ratio. In the event of a system failure or rejected task, all charged credits are automatically refunded.

What are best practices for crafting Nano Banana Text to Image prompts?

Follow the structured FRAME framework by detailing your main subject, rendering texture, camera angle, lighting setup, and typography. Using explicit descriptors such as "diffused studio rim lighting" yields superior clarity compared to generic aesthetic adjectives.