Gemini 3.5 Flash API
google/gemini-3.5-flashGemini 3.5 Flash turns text, images, audio, video, and documents into code, analysis, and grounded answers for demanding tasks. Its 1,048,576-token context window keeps extensive source material in view while system instructions and generation settings shape the response.
Gemini 3.5 Flash
Google · chat-completions
Quick start
POST to /v1beta/models/google/gemini-3.5-flash:generateContent. For SSE, use /v1beta/models/google/gemini-3.5-flash:streamGenerateContent?alt=sse; responses include generated content and usageMetadata.
- Endpoint
- POST /v1beta/models/google/gemini-3.5-flash:generateContent
- Model ID
- google/gemini-3.5-flash
- Protocol
- Gemini Native
Request fields
| Field | Requirement | Description |
|---|---|---|
| model | Required in path | google/gemini-3.5-flash |
| contents | Required | A non-empty conversation array with role and parts. |
| contents[].parts[].text | Text input | Text within a conversation turn. |
| contents[].parts[].inlineData | Multimodal input | Inline media with mimeType and Base64 data. |
| systemInstruction | Optional | System instructions for the conversation. |
| generationConfig.temperature | Optional | Sampling temperature. |
| generationConfig.maxOutputTokens | Optional | Output token control for this response. |
| generationConfig.topP | Optional | Nucleus sampling probability. |
| generationConfig.topK | Optional | Top K sampling setting. |
| generationConfig.stopSequences | Optional | Strings that stop generation. |
| safetySettings | Optional | Safety settings array. |
Response and usage
Generated content appears in candidates; usageMetadata reports input and output token counts.
Streaming
POST to /v1beta/models/google/gemini-3.5-flash:streamGenerateContent?alt=sse and read SSE; usage comes from usageMetadata.
curl --request POST \
--url https://api.vidgo.ai/v1beta/models/google/gemini-3.5-flash:generateContent \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"contents": [
{
"role": "user",
"parts": [
{
"text": "Summarize the key risks in this design."
}
]
}
],
"generationConfig": {
"temperature": 1,
"topP": 1,
"maxOutputTokens": 1024
}
}'Keep API keys on the server. Never expose them in browser code or public repositories.
Continue with
Related Models
Gemini 3.5 Flash API — Frequently asked questions
What is the Gemini 3.5 Flash API?
Gemini 3.5 Flash is a Google model that turns text and multimodal input into analysis, code, and answers. Its 1,048,576-token context window supports work across extensive source material. System instructions and generation settings help shape each response. You can call it through Chat Completions or Gemini Native, or try a text conversation in the Playground above.
How large is the Gemini 3.5 Flash context window?
Gemini 3.5 Flash accepts a context window of 1,048,576 tokens. Use it to connect large documents, code files, and conversation history in a single task.
How many output tokens can Gemini 3.5 Flash generate?
The model API supports up to 65,536 output tokens. Set max_tokens for Chat Completions or generationConfig.maxOutputTokens for Gemini Native. The on-page Playground control ranges from 256 to 8,192 tokens.
Which input types can Gemini 3.5 Flash analyze?
Gemini 3.5 Flash can analyze text, images, audio, video, and documents. Pair the material with a specific question or extraction goal in an API request.
How can Gemini 3.5 Flash help with coding?
Provide requirements, relevant source files, and test expectations. Gemini 3.5 Flash can trace relationships across the code, propose changes, explain errors, and review an implementation.
How do I call Gemini 3.5 Flash in Gemini Native format?
POST a contents array to /v1beta/models/google/gemini-3.5-flash:generateContent. Add systemInstruction or generationConfig when the task needs them. For SSE, call the streamGenerateContent method with alt=sse.
How is Gemini 3.5 Flash priced on Vidgo?
Input costs $0.900 or 180 credits per 1M tokens. Output costs $5.40 or 1080 credits per 1M tokens. The response usage reports the token counts for each request.