Gemini 3.7 Flash API

google/gemini-3.7-flash
1,048,576 tokens · 120 input / 12 cached input / 600 output credits / 1M tokens

Gemini 3.7 Flash turns text, images, audio, video, and documents into code, analysis, and useful answers. Its 1,048,576-token context supports extensive source material, while system instructions and generation settings guide the response.

Gemini 3.7 Flash

Google · chat-completions

Chat with Gemini 3.7 Flash

Each model has its own conversation. Switching never sends another model's history, and switching back resumes where you left off. Requests are billed from actual token usage.

Ctrl / ⌘ + Enter to send0 / 32,000

Continue with

Gemini 3.7 Flash

Gemini 3.7 Flash supports agentic coding, web development, multimodal analysis, and long-document work across a 1,048,576-token context window. Chat Completions accepts ordered messages, while Gemini Native uses contents, systemInstruction, and generationConfig. Both formats return generated text and token usage, with streaming available for interactive responses. The on-page Playground sends text conversations through either format; API requests can include supported multimodal inputs.

Why Choose Gemini 3.7 Flash?

  • Agentic software developmentConnect requirements, repository context, and tool results across coding and debugging tasks.

  • Web interface workUse design references and written requirements to plan interfaces, generate code, and review implementation details.

  • Long-context knowledge workExamine related documents and source files within a 1,048,576-token context window.

  • Multimodal understandingCombine text instructions with images, audio, video, and PDFs for focused analysis and extraction.

  • Two request formatsChoose ordered messages in Chat Completions or structured contents and generationConfig in Gemini Native.

Parameters

ParameterRequirementDescription
modelChat Completions: required

Set to google/gemini-3.7-flash.

messagesChat Completions: required

A non-empty ordered array of role and content messages.

temperatureOptional

Sampling temperature; the Playground accepts 0–2 in steps of 0.1.

Default1.0
max_tokensChat Completions: optional

Sets the output limit from 1 to 65,536 tokens in the Playground.

Default4096
top_pChat Completions: optional

Nucleus sampling probability; the Playground accepts 0–1 in steps of 0.1.

Default1.0
contentsGemini Native: required

A non-empty array of conversation turns with role and parts.

systemInstructionGemini Native: optional

System guidance supplied through parts.

generationConfig.temperatureGemini Native: optional

Sampling temperature; the Playground accepts 0–2 in steps of 0.1.

Default1.0
generationConfig.maxOutputTokensGemini Native: optional

Sets the output limit from 1 to 65,536 tokens in the Playground.

Default4096
generationConfig.topPGemini Native: optional

Nucleus sampling probability; the Playground accepts 0–1 in steps of 0.1.

Default1.0
streamChat Completions: optional

Set true for SSE output; Gemini Native uses the streamGenerateContent method.

Defaultfalsetrue

How to Use

  1. Describe the taskProvide the question, source material, and the output format you need.

  2. Choose an API formatSend messages to Chat Completions or contents to Gemini Native.

  3. Set generation controlsChoose an output limit and optional system instructions or sampling settings.

  4. Read the resultReview the generated content and token usage; add relevant context in the next turn.

Pricing

Input, cached input, and output are metered separately. Rates apply per 1M tokens and settle in credits.

UsageRateDetails
Input$0.600 · 120 credits / 1M tokensRequest tokens counted as standard input.
Cached input$0.060 · 12 credits / 1M tokensInput tokens read from cache when cache usage is reported.
Output (including thinking tokens)$3.00 · 600 credits / 1M tokensGenerated response and thinking tokens.

Best Use Cases

  • Repository developmentAnalyze requirements and related source files to draft code, debugging steps, or reviews.

  • Web interface implementationTurn a design reference and interaction requirements into frontend code and revision notes.

  • Long-document analysisCompare details across reports, PDFs, and research material in one task.

  • Multimodal reviewAsk focused questions about images, audio, video, and documents.

Pro Tips

  • Provide the task, relevant source material, and acceptance criteria together for coding work.
  • Use systemInstruction in Gemini Native to guide the answer format and task boundaries.
  • Set max_tokens or generationConfig.maxOutputTokens to the response length needed for the selected format.

Usage Notes

  • The input context window is 1,048,576 tokens, and maximum output is 65,536 tokens.
  • The text-chat Playground defaults to 4,096 output tokens and accepts values from 1 to 65,536.
  • Use google/gemini-3.7-flash as the Chat Completions model or the complete Gemini Native path ID.
  • The model offers low, medium, and high thinking levels through supported API requests.

Related Models

Gemini 3.7 Flash API — Frequently asked questions

What is the Gemini 3.7 Flash API?

Gemini 3.7 Flash is a Google model for coding, multimodal analysis, and long-context knowledge work. It turns text, images, audio, video, and documents into code, analysis, and specific answers. Its 1,048,576-token context window and generation settings help carry source material and instructions through complex tasks. Call it through Chat Completions or Gemini Native, or try a text conversation in the Playground above.

How large is the Gemini 3.7 Flash context window?

Gemini 3.7 Flash accepts up to 1,048,576 input tokens and generates up to 65,536 output tokens. Place related documents, code, and instructions together when the task depends on connections across sources.

Which inputs can Gemini 3.7 Flash analyze?

Gemini 3.7 Flash can analyze text, images, audio, video, and PDFs. Pair each source with a focused question, extraction goal, or development task in the API request.

How does Gemini 3.7 Flash support coding and web development?

Provide requirements, repository context, a design reference, and acceptance criteria. Gemini 3.7 Flash can plan changes, generate code, explain errors, and review interface implementations.

Which Gemini 3.7 Flash API formats can I use?

Send a model and messages to /v1/chat/completions, or send contents to /v1beta/models/google/gemini-3.7-flash:generateContent. Use the corresponding streaming method for SSE output.

How are cached input tokens priced for Gemini 3.7 Flash?

Cached input is billed at $0.060 or 12 credits per 1M tokens when the response reports cache reads. Standard input is $0.600 or 120 credits per 1M tokens; output, including thinking tokens, is $3.00 or 600 credits per 1M tokens.

Which thinking levels can Gemini 3.7 Flash use?

Gemini 3.7 Flash offers low, medium, and high thinking levels. Choose a level in an API request to balance response speed, token use, and reasoning depth for the task.

What output limit does the Gemini 3.7 Flash Playground offer?

Max Tokens starts at 4,096 and accepts any whole number from 1 to 65,536. Chat Completions sends max_tokens; Gemini Native sends generationConfig.maxOutputTokens.