Gemini 3.5 Flash API

google/gemini-3.5-flash
1,048,576 tokens · 180 input credits / 1M tokens

Gemini 3.5 Flash turns text, images, audio, video, and documents into code, analysis, and grounded answers for demanding tasks. Its 1,048,576-token context window keeps extensive source material in view while system instructions and generation settings shape the response.

Gemini 3.5 Flash

Google · chat-completions

Chat with Gemini 3.5 Flash

Each model has its own conversation. Switching never sends another model's history, and switching back resumes where you left off. Requests are billed from actual token usage.

Ctrl / ⌘ + Enter to send0 / 32,000

Continue with

Gemini 3.5 Flash

Gemini 3.5 Flash supports coding, long-context analysis, and multimodal understanding across text, images, audio, video, and documents. Use Chat Completions for ordered conversation messages or Gemini Native for structured contents and generationConfig. Both interfaces offer a complete response or streaming output, with token usage in the result. The on-page Playground sends text conversations through either format, while API requests can carry supported multimodal input.

Why Choose Gemini 3.5 Flash?

  • Large context for connected evidenceAnalyze related source files, documents, and conversation history within a 1,048,576-token context window.

  • Coding and iterative workTurn requirements and repository context into code, debugging ideas, and review notes across multi-step tasks.

  • Multimodal understandingCombine written instructions with images, audio, video, or documents to examine material in its original form.

  • Two API formatsUse ordered messages with Chat Completions or native contents, systemInstruction, and generationConfig with Gemini Native.

Parameters

ParameterRequirementDescription
modelChat Completions: required

Set to google/gemini-3.5-flash.

messagesChat Completions: required

A non-empty ordered array of role and content messages.

max_tokensChat Completions: optional

Sets the response output limit, up to the model's 65,536-token API maximum.

contentsGemini Native: required

A non-empty array of conversation turns with role and parts.

systemInstructionGemini Native: optional

System guidance supplied through parts.

generationConfig.temperatureGemini Native: optional

Controls sampling temperature.

generationConfig.topPGemini Native: optional

Controls nucleus sampling.

generationConfig.topKGemini Native: optional

Sets the number of candidate tokens in top-k sampling.

generationConfig.stopSequencesGemini Native: optional

Stops generation when a listed sequence appears.

generationConfig.maxOutputTokensGemini Native: optional

Sets the response output limit, up to 65,536 tokens.

safetySettingsGemini Native: optional

Configures category-specific safety thresholds.

streamChat Completions: optional

Set true for SSE output; Gemini Native uses the streamGenerateContent method.

Defaultfalsetrue

How to Use

  1. Describe the taskProvide the question, source material, and the output format you need.

  2. Choose an API formatSend messages to Chat Completions or contents to Gemini Native.

  3. Set generation controlsChoose an output limit and optional system instructions or sampling settings.

  4. Read the resultReview the generated content and token usage; add relevant context in the next turn.

Pricing

Input and output token usage is measured separately. Rates below apply per 1M tokens and settle in credits.

UsageRateDetails
Input$0.900 · 180 credits / 1M tokensTokens counted as request input.
Output$5.40 · 1080 credits / 1M tokensTokens generated in the response, including thinking tokens.

Best Use Cases

  • Codebase analysisReview related files and requirements before implementing or debugging a change.

  • Long-document synthesisConnect details across reports, policies, and research materials in one request.

  • Multimodal reviewAsk focused questions about screenshots, recordings, video, or document content.

Pro Tips

  • Put the objective, relevant source material, and acceptance criteria in the same request.
  • Use systemInstruction in Gemini Native to keep output format and task boundaries consistent.
  • For long responses, set max_tokens or generationConfig.maxOutputTokens to match the selected API format.

Usage Notes

  • The context window is 1,048,576 tokens, and the model API supports up to 65,536 output tokens.
  • The text-chat Playground offers a 256–8,192-token output control and can switch between Chat Completions and Gemini Native.
  • Use google/gemini-3.5-flash as the Chat Completions model or as the complete Gemini Native path ID.

Related Models

Gemini 3.5 Flash API — Frequently asked questions

What is the Gemini 3.5 Flash API?

Gemini 3.5 Flash is a Google model that turns text and multimodal input into analysis, code, and answers. Its 1,048,576-token context window supports work across extensive source material. System instructions and generation settings help shape each response. You can call it through Chat Completions or Gemini Native, or try a text conversation in the Playground above.

How large is the Gemini 3.5 Flash context window?

Gemini 3.5 Flash accepts a context window of 1,048,576 tokens. Use it to connect large documents, code files, and conversation history in a single task.

How many output tokens can Gemini 3.5 Flash generate?

The model API supports up to 65,536 output tokens. Set max_tokens for Chat Completions or generationConfig.maxOutputTokens for Gemini Native. The on-page Playground control ranges from 256 to 8,192 tokens.

Which input types can Gemini 3.5 Flash analyze?

Gemini 3.5 Flash can analyze text, images, audio, video, and documents. Pair the material with a specific question or extraction goal in an API request.

How can Gemini 3.5 Flash help with coding?

Provide requirements, relevant source files, and test expectations. Gemini 3.5 Flash can trace relationships across the code, propose changes, explain errors, and review an implementation.

How do I call Gemini 3.5 Flash in Gemini Native format?

POST a contents array to /v1beta/models/google/gemini-3.5-flash:generateContent. Add systemInstruction or generationConfig when the task needs them. For SSE, call the streamGenerateContent method with alt=sse.

How is Gemini 3.5 Flash priced on Vidgo?

Input costs $0.900 or 180 credits per 1M tokens. Output costs $5.40 or 1080 credits per 1M tokens. The response usage reports the token counts for each request.