DeepSeek V4 Flash API Pricing
As of October 10, 2026, DeepSeek V4 Flash costs $0.30 per 1M input tokens and $1.20 per 1M output tokens on DeepSeek, with cache reads at $0.006 and cache writes at $0 per 1M tokens. Across 22 providers on LiteLLM, the input price ranges from $0.08 to $0.44. It has a 1,000,000-token context window and returns up to 393,216 tokens per response.
- Input
- $0.30 / 1M tokens
- Output
- $1.20 / 1M tokens
- Cache read
- $0.006 / 1M tokens
- Context window
- 1,000,000 tokens
- Max output
- 393,216 tokens
- Providers
- 22
- Added to LiteLLM
- June 2026
DeepSeek V4 Flash pricing by provider
34 listings across 22 providers. Token prices are in US dollars per 1M tokens. The highlighted row is the price quoted above.
| Provider | Model name in LiteLLM | Input | Output | Cache read | Cache write | Context |
|---|---|---|---|---|---|---|
| DeepSeek | deepseek-v4-flash | $0.30 | $1.20 | $0.006 | $0 | 1M |
| DeepSeek | deepseek/deepseek-v4-flash | $0.30 | $1.20 | $0.006 | $0 | 1M |
| AIHubMix | aihubmix/deepseek-v4-flash | $0.142 | $0.284 | $0.0284 | – | 1M |
| Alibaba Cloud Model Studio | dashscope/deepseek-v4-flash | $0.20 | $0.40 | $0.04 | – | 1M |
| Alibaba Cloud Model Studiosnapshot 0731 | dashscope/deepseek-v4-flash-0731 | $0.20 | $0.40 | $0.04 | – | 1M |
| Azure AI Foundry | azure_ai/deepseek-v4-flash | $0.19 | $0.51 | $0.028 | – | 1M |
| Azure AI Foundrysnapshot 0731 | azure_ai/DeepSeek-V4-Flash-0731 | $0.44 | $1.32 | $0.014 | – | 1M |
| Basetensnapshot 0731 | baseten/deepseek-ai/DeepSeek-V4-Flash-0731 | $0.13 | $0.26 | $0.028 | – | 1M |
| Databrickssnapshot 0731 | databricks/databricks-deepseek-v4-flash-0731 | $0.14 | $0.28 | $0.028 | $0.14 | 1M |
| DeepInfra | deepinfra/deepseek-ai/DeepSeek-V4-Flash | $0.09 | $0.18 | $0.018 | – | 1M |
| DeepInfrasnapshot 0731 | deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731 | $0.08 | $0.18 | $0.016 | – | 1M |
| Fireworks AI | fireworks_ai/accounts/fireworks/models/deepseek-v4-flash | $0.14 | $0.28 | $0.028 | – | 1M |
| Fireworks AIsnapshot 0731 | fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731 | $0.22 | $0.66 | $0.007 | – | 1M |
| Fireworks AI | fireworks_ai/deepseek-v4-flash | $0.14 | $0.28 | $0.028 | – | 1M |
| Fireworks AIsnapshot 0731 | fireworks_ai/deepseek-v4-flash-0731 | $0.22 | $0.66 | $0.007 | – | 1M |
| LibertAI | libertai/deepseek-v4-flash | $0.25 | $1.75 | – | – | 200K |
| Nebius AI Studio | nebius/deepseek-ai/DeepSeek-V4-Flash | $0.14 | $0.28 | – | – | 1M |
| Nebius AI Studiosnapshot 0731 | nebius/deepseek-ai/DeepSeek-V4-Flash-0731 | $0.14 | $0.28 | – | – | 1M |
| Novita AI | novita/deepseek/deepseek-v4-flash | $0.14 | $0.28 | $0.028 | – | 1M |
| Novita AIsnapshot 0731 | novita/deepseek/deepseek-v4-flash-0731 | $0.44 | $1.32 | $0.028 | – | 1M |
| Perplexitysnapshot 0731 | perplexity/perplexity/deepseek-v4-flash-0731 | $0.13 | $0.26 | $0.028 | – | – |
| Pinstripes | pinstripes/ps/deepseek-v4-flash | $0.10 | $0.20 | – | – | 160K |
| Prism | prism/deepseek-v4-flash | $0.17 | $0.21 | $0.07 | – | 1M |
| Qianwen AI Platform | qwen_ai_platform/deepseek-v4-flash | $0.20 | $0.40 | $0.04 | – | 1M |
| Qianwen AI Platformsnapshot 0731 | qwen_ai_platform/deepseek-v4-flash-0731 | $0.20 | $0.40 | $0.04 | – | 1M |
Show 9 more listings
| Provider | Model name in LiteLLM | Input | Output | Cache read | Cache write | Context |
|---|---|---|---|---|---|---|
| QwenCloud | qwencloud/deepseek-v4-flash | $0.20 | $0.40 | $0.04 | – | 1M |
| QwenCloudsnapshot 0731 | qwencloud/deepseek-v4-flash-0731 | $0.20 | $0.40 | $0.04 | – | 1M |
| Sailsnapshot 0731 | sail/deepseek-ai/DeepSeek-V4-Flash-0731 | $0.09 | $0.18 | $0.02 | – | 1M |
| Scalewaysnapshot 0731 | scaleway/deepseek-v4-flash-0731 | $0.40 | $0.80 | $0.08 | – | 256K |
| Tencent TokenHub | tencent/deepseek-v4-flash | $0.14 | $0.28 | $0.0028 | $0 | 1M |
| Tensormesh | tensormesh/deepseek-ai/DeepSeek-V4-Flash | $0.14 | $0.28 | $0 | – | 32K |
| Together AIsnapshot 0731 | together_ai/deepseek-ai/DeepSeek-V4-Flash-0731 | $0.14 | $0.28 | $0.03 | – | 1M |
| Weights & Biases | wandb/deepseek-ai/DeepSeek-V4-Flash | $0.14 | $0.28 | $0.07 | – | 1M |
| Weights & Biasessnapshot 0731 | wandb/deepseek-ai/DeepSeek-V4-Flash-0731 | $0.13 | $0.28 | $0.07 | – | 262K |
What DeepSeek V4 Flash costs in practice
1,000 requests with 2,000 input and 500 output tokens each
| DeepSeek | $1.20 |
| DeepInfra | $0.25 |
| Sail | $0.27 |
| Pinstripes | $0.30 |
100 long-document requests with 100,000 input and 2,000 output tokens each
| DeepSeek | $3.24 |
| DeepInfra | $0.836 |
| Sail | $0.936 |
| Pinstripes | $1.04 |
DeepSeek V4 Flash features
- Function calling
- Parallel tool calls
- Structured output
- Image input
- Reasoning
- Prompt caching
- Web search
Use DeepSeek V4 Flash with LiteLLM
Call DeepSeek V4 Flash through the LiteLLM Python SDK, or put it behind the LiteLLM proxy and send OpenAI-format requests to /v1/chat/completions. LiteLLM tracks the cost of every request at these prices.
from litellm import completion
response = completion(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Hello!"}],
)model_list:
- model_name: deepseek-v4-flash
litellm_params:
model: deepseek-v4-flash
api_key: os.environ/DEEPSEEK_API_KEYcurl http://0.0.0.0:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-1234" \
-d '{
"model": "deepseek-v4-flash",
"messages": [{"role": "user", "content": "Hello!"}]
}'DeepSeek V4 Flash FAQ
How much does DeepSeek V4 Flash cost?
DeepSeek V4 Flash costs $0.30 per 1M input tokens and $1.20 per 1M output tokens on DeepSeek, with cache reads at $0.006 and cache writes at $0 per 1M tokens. Across 22 providers on LiteLLM, the input price ranges from $0.08 to $0.44.
What is the context window of DeepSeek V4 Flash?
DeepSeek V4 Flash has a 1,000,000-token context window and returns up to 393,216 tokens per response.
Which providers offer DeepSeek V4 Flash?
DeepSeek V4 Flash is available from 22 providers through LiteLLM: DeepSeek, AIHubMix, Alibaba Cloud Model Studio, Azure AI Foundry, Baseten, Databricks, DeepInfra, Fireworks AI, LibertAI, Nebius AI Studio, Novita AI and Perplexity and 10 more. The lowest price is on DeepInfra.
How do I call DeepSeek V4 Flash with an OpenAI-compatible API?
Use the LiteLLM Python SDK with model="deepseek-v4-flash", or add the model to the LiteLLM proxy's config.yaml and send requests to /v1/chat/completions from any OpenAI client. LiteLLM tracks the cost of every request at the prices on this page.
What features does DeepSeek V4 Flash support?
DeepSeek V4 Flash supports function calling, parallel tool calls, structured output, image input, reasoning, prompt caching and web search.
When did LiteLLM add DeepSeek V4 Flash?
DeepSeek V4 Flash was added to LiteLLM's model price file in June 2026.
Sources
Prices come from LiteLLM's open-source model_prices_and_context_window.json, which LiteLLM uses to calculate the cost of every request. Provider pricing pages:
- api-docs.deepseek.com/quick_start/pricing
- aihubmix.com/api/v1/models
- www.alibabacloud.com/help/en/model-studio/models
- prices.azure.com/api/retail/prices?$filter=serviceName%20eq%20'Foundry%20Models'%20and%20armRegionName%20eq%20'eastus'%20and%20priceType%20eq%20'Consumption'
LiteLLM docs: DeepSeek, AIHubMix, Alibaba Cloud Model Studio, Azure AI Foundry, Baseten, Databricks.
Provider pages: DeepSeek, AIHubMix, Alibaba Cloud Model Studio, Azure AI Foundry, Baseten, Databricks, DeepInfra, Fireworks AI, LibertAI, Nebius AI Studio, Novita AI, Perplexity.
