DeepSeek V4 Flash Chat Completions API
deepseek/deepseek-v4-flashDeepSeek V4 Flash Chat Completions turns conversation messages and long text context into answers, code explanations, and document analysis. Its 1M-token context helps connect details across large source collections while its efficient Flash profile suits frequent chat and coding assistance.
DeepSeek V4 Flash
DeepSeek · chat-completions
Quick start
Send one POST request to /v1/chat/completions. Non-streaming calls return the model result and usage directly, with no task creation or status polling.
- Endpoint
- POST /v1/chat/completions
- Model ID
- deepseek/deepseek-v4-flash
- Protocol
- Chat Completions
Request fields
| Field | Requirement | Description |
|---|---|---|
| model | Required | deepseek/deepseek-v4-flash |
| messages | Required | A non-empty array of conversation messages. |
| max_tokens | Optional | Limits output tokens for this response. |
| temperature | Optional | Sampling temperature; the playground allows 0–2. |
| top_p | Optional | Nucleus sampling probability; the playground allows 0–1. |
| stream | Optional | Returns an SSE event stream when true. |
Response and usage
Successful responses include generated content and usage. Billing settles from actual input, output, and cache tokens.
Streaming
Set stream: true to read SSE. Finalize usage accounting from the terminal usage event.
curl --request POST \
--url https://api.vidgo.ai/v1/chat/completions \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"model": "deepseek/deepseek-v4-flash",
"messages": [
{
"role": "user",
"content": "Explain why deterministic retries matter in distributed systems."
}
],
"max_tokens": 1024,
"temperature": 1,
"top_p": 1,
"stream": false
}'Keep API keys on the server. Never expose them in browser code or public repositories.
Continue with
Related Models
DeepSeek V4 Flash Chat Completions API — Frequently asked questions
What is the DeepSeek V4 Flash Chat Completions API?
DeepSeek V4 Flash Chat Completions is a DeepSeek model for turning conversation messages and long text into answers, code explanations, and document analysis. Its 1M-token context connects material across large inputs, while the 284B-total, 13B-active Flash profile is designed for efficient use. Clear source labels and instructions help direct each response. You can call it programmatically or try it from the playground above.
How much context can DeepSeek V4 Flash use?
DeepSeek V4 Flash has a 1M-token context window. Place related documents, code excerpts, and earlier messages together when the task requires connections across them.
How should I organize long documents for DeepSeek V4 Flash?
Give each document a title and stable section markers, then ask a focused question that names the evidence to compare. This makes the answer easier to check against the source material.
How can DeepSeek V4 Flash help with coding?
Provide the relevant code, the observed behavior, and the intended result. DeepSeek V4 Flash can explain functions, suggest debugging steps, and draft an initial change plan.
How does DeepSeek V4 Flash handle follow-up messages?
Send the earlier user and assistant turns in order in messages, followed by the new question. The model uses that conversation context to continue the task.
How do I control DeepSeek V4 Flash response length?
Set max_tokens to a positive integer for the requested response length. The on-page Playground offers values from 256 to 8,192 tokens.
How can DeepSeek V4 Flash help plan an agent task?
Describe the goal, available actions, constraints, and completion criteria in the conversation. The model can draft a first-pass sequence of steps for review.