Gemini 3.7 Flash API
google/gemini-3.7-flashGemini 3.7 Flash turns text, images, audio, video, and documents into code, analysis, and useful answers. Its 1,048,576-token context supports extensive source material, while system instructions and generation settings guide the response.
Gemini 3.7 Flash
Google · chat-completions
Quick start
Send a POST request to /v1/chat/completions with google/gemini-3.7-flash. Choose JSON or SSE with stream, and receive generated content and usage.
- Endpoint
- POST /v1/chat/completions
- Model ID
- google/gemini-3.7-flash
- Protocol
- Chat Completions
Request fields
| Field | Requirement | Description |
|---|---|---|
| model | Required | google/gemini-3.7-flash |
| messages | Required | A non-empty array of conversation messages. |
| max_tokens | Optional | Limits output tokens for this response. |
| temperature | Optional | Sampling temperature; the playground allows 0–2. |
| top_p | Optional | Nucleus sampling probability; the playground allows 0–1. |
| stream | Optional | Returns an SSE event stream when true. |
Response and usage
Successful responses include generated content and usage. Billing settles from actual input, output, and cache tokens.
Streaming
Set stream: true to read SSE. Finalize usage accounting from the terminal usage event.
curl --request POST \
--url https://api.vidgo.ai/v1/chat/completions \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"model": "google/gemini-3.7-flash",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Explain why deterministic retries matter in distributed systems."
}
],
"max_tokens": 4096,
"temperature": 1,
"top_p": 1,
"stream": false
}'Keep API keys on the server. Never expose them in browser code or public repositories.
Continue with
Related Models
Gemini 3.7 Flash API — Frequently asked questions
What is the Gemini 3.7 Flash API?
Gemini 3.7 Flash is a Google model for coding, multimodal analysis, and long-context knowledge work. It turns text, images, audio, video, and documents into code, analysis, and specific answers. Its 1,048,576-token context window and generation settings help carry source material and instructions through complex tasks. Call it through Chat Completions or Gemini Native, or try a text conversation in the Playground above.
How large is the Gemini 3.7 Flash context window?
Gemini 3.7 Flash accepts up to 1,048,576 input tokens and generates up to 65,536 output tokens. Place related documents, code, and instructions together when the task depends on connections across sources.
Which inputs can Gemini 3.7 Flash analyze?
Gemini 3.7 Flash can analyze text, images, audio, video, and PDFs. Pair each source with a focused question, extraction goal, or development task in the API request.
How does Gemini 3.7 Flash support coding and web development?
Provide requirements, repository context, a design reference, and acceptance criteria. Gemini 3.7 Flash can plan changes, generate code, explain errors, and review interface implementations.
Which Gemini 3.7 Flash API formats can I use?
Send a model and messages to /v1/chat/completions, or send contents to /v1beta/models/google/gemini-3.7-flash:generateContent. Use the corresponding streaming method for SSE output.
How are cached input tokens priced for Gemini 3.7 Flash?
Cached input is billed at $0.060 or 12 credits per 1M tokens when the response reports cache reads. Standard input is $0.600 or 120 credits per 1M tokens; output, including thinking tokens, is $3.00 or 600 credits per 1M tokens.
Which thinking levels can Gemini 3.7 Flash use?
Gemini 3.7 Flash offers low, medium, and high thinking levels. Choose a level in an API request to balance response speed, token use, and reasoning depth for the task.
What output limit does the Gemini 3.7 Flash Playground offer?
Max Tokens starts at 4,096 and accepts any whole number from 1 to 65,536. Chat Completions sends max_tokens; Gemini Native sends generationConfig.maxOutputTokens.