# DeepSeek V4 Flash API Pricing

As of October 10, 2026, DeepSeek V4 Flash costs $0.30 per 1M input tokens and $1.20 per 1M output tokens on DeepSeek, with cache reads at $0.006 and cache writes at $0 per 1M tokens. Across 22 providers on LiteLLM, the input price ranges from $0.08 to $0.44. It has a 1,000,000-token context window and returns up to 393,216 tokens per response.

- Page: https://models.litellm.ai/models/deepseek-v4-flash
- Lab: DeepSeek
- LiteLLM model name: `deepseek-v4-flash`
- Context window: 1,000,000 tokens
- Max output: 393,216 tokens
- Added to LiteLLM: June 2026
- Prices updated: October 10, 2026

## Pricing by provider

Token prices in USD per 1M tokens.

| Provider | Model name | Input | Output | Cache read | Per image | Per second | Context |
|---|---|---|---|---|---|---|---|
| DeepSeek | `deepseek-v4-flash` | $0.30 | $1.20 | $0.006 | – | – | 1M |
| DeepSeek | `deepseek/deepseek-v4-flash` | $0.30 | $1.20 | $0.006 | – | – | 1M |
| AIHubMix | `aihubmix/deepseek-v4-flash` | $0.142 | $0.284 | $0.0284 | – | – | 1M |
| Alibaba Cloud Model Studio | `dashscope/deepseek-v4-flash` | $0.20 | $0.40 | $0.04 | – | – | 1M |
| Alibaba Cloud Model Studio (snapshot 0731) | `dashscope/deepseek-v4-flash-0731` | $0.20 | $0.40 | $0.04 | – | – | 1M |
| Azure AI Foundry | `azure_ai/deepseek-v4-flash` | $0.19 | $0.51 | $0.028 | – | – | 1M |
| Azure AI Foundry (snapshot 0731) | `azure_ai/DeepSeek-V4-Flash-0731` | $0.44 | $1.32 | $0.014 | – | – | 1M |
| Baseten (snapshot 0731) | `baseten/deepseek-ai/DeepSeek-V4-Flash-0731` | $0.13 | $0.26 | $0.028 | – | – | 1M |
| Databricks (snapshot 0731) | `databricks/databricks-deepseek-v4-flash-0731` | $0.14 | $0.28 | $0.028 | – | – | 1M |
| DeepInfra | `deepinfra/deepseek-ai/DeepSeek-V4-Flash` | $0.09 | $0.18 | $0.018 | – | – | 1M |
| DeepInfra (snapshot 0731) | `deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731` | $0.08 | $0.18 | $0.016 | – | – | 1M |
| Fireworks AI | `fireworks_ai/accounts/fireworks/models/deepseek-v4-flash` | $0.14 | $0.28 | $0.028 | – | – | 1M |
| Fireworks AI (snapshot 0731) | `fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731` | $0.22 | $0.66 | $0.007 | – | – | 1M |
| Fireworks AI | `fireworks_ai/deepseek-v4-flash` | $0.14 | $0.28 | $0.028 | – | – | 1M |
| Fireworks AI (snapshot 0731) | `fireworks_ai/deepseek-v4-flash-0731` | $0.22 | $0.66 | $0.007 | – | – | 1M |
| LibertAI | `libertai/deepseek-v4-flash` | $0.25 | $1.75 | – | – | – | 200K |
| Nebius AI Studio | `nebius/deepseek-ai/DeepSeek-V4-Flash` | $0.14 | $0.28 | – | – | – | 1M |
| Nebius AI Studio (snapshot 0731) | `nebius/deepseek-ai/DeepSeek-V4-Flash-0731` | $0.14 | $0.28 | – | – | – | 1M |
| Novita AI | `novita/deepseek/deepseek-v4-flash` | $0.14 | $0.28 | $0.028 | – | – | 1M |
| Novita AI (snapshot 0731) | `novita/deepseek/deepseek-v4-flash-0731` | $0.44 | $1.32 | $0.028 | – | – | 1M |
| Perplexity (snapshot 0731) | `perplexity/perplexity/deepseek-v4-flash-0731` | $0.13 | $0.26 | $0.028 | – | – | – |
| Pinstripes | `pinstripes/ps/deepseek-v4-flash` | $0.10 | $0.20 | – | – | – | 160K |
| Prism | `prism/deepseek-v4-flash` | $0.17 | $0.21 | $0.07 | – | – | 1M |
| Qianwen AI Platform | `qwen_ai_platform/deepseek-v4-flash` | $0.20 | $0.40 | $0.04 | – | – | 1M |
| Qianwen AI Platform (snapshot 0731) | `qwen_ai_platform/deepseek-v4-flash-0731` | $0.20 | $0.40 | $0.04 | – | – | 1M |
| QwenCloud | `qwencloud/deepseek-v4-flash` | $0.20 | $0.40 | $0.04 | – | – | 1M |
| QwenCloud (snapshot 0731) | `qwencloud/deepseek-v4-flash-0731` | $0.20 | $0.40 | $0.04 | – | – | 1M |
| Sail (snapshot 0731) | `sail/deepseek-ai/DeepSeek-V4-Flash-0731` | $0.09 | $0.18 | $0.02 | – | – | 1M |
| Scaleway (snapshot 0731) | `scaleway/deepseek-v4-flash-0731` | $0.40 | $0.80 | $0.08 | – | – | 256K |
| Tencent TokenHub | `tencent/deepseek-v4-flash` | $0.14 | $0.28 | $0.0028 | – | – | 1M |
| Tensormesh | `tensormesh/deepseek-ai/DeepSeek-V4-Flash` | $0.14 | $0.28 | $0 | – | – | 32K |
| Together AI (snapshot 0731) | `together_ai/deepseek-ai/DeepSeek-V4-Flash-0731` | $0.14 | $0.28 | $0.03 | – | – | 1M |
| Weights & Biases | `wandb/deepseek-ai/DeepSeek-V4-Flash` | $0.14 | $0.28 | $0.07 | – | – | 1M |
| Weights & Biases (snapshot 0731) | `wandb/deepseek-ai/DeepSeek-V4-Flash-0731` | $0.13 | $0.28 | $0.07 | – | – | 262K |

## Cost examples

1,000 requests with 2,000 input and 500 output tokens each:

- DeepSeek: $1.20
- DeepInfra: $0.25
- Sail: $0.27
- Pinstripes: $0.30

100 long-document requests with 100,000 input and 2,000 output tokens each:

- DeepSeek: $3.24
- DeepInfra: $0.836
- Sail: $0.936
- Pinstripes: $1.04

## Features

Supported: Function calling, Parallel tool calls, Structured output, Image input, Reasoning, Prompt caching, Web search.

## Use it with LiteLLM

```python
from litellm import completion

response = completion(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
```

## FAQ

### How much does DeepSeek V4 Flash cost?

DeepSeek V4 Flash costs $0.30 per 1M input tokens and $1.20 per 1M output tokens on DeepSeek, with cache reads at $0.006 and cache writes at $0 per 1M tokens. Across 22 providers on LiteLLM, the input price ranges from $0.08 to $0.44.

### What is the context window of DeepSeek V4 Flash?

DeepSeek V4 Flash has a 1,000,000-token context window and returns up to 393,216 tokens per response.

### Which providers offer DeepSeek V4 Flash?

DeepSeek V4 Flash is available from 22 providers through LiteLLM: DeepSeek, AIHubMix, Alibaba Cloud Model Studio, Azure AI Foundry, Baseten, Databricks, DeepInfra, Fireworks AI, LibertAI, Nebius AI Studio, Novita AI and Perplexity and 10 more. The lowest price is on DeepInfra.

### How do I call DeepSeek V4 Flash with an OpenAI-compatible API?

Use the LiteLLM Python SDK with model="deepseek-v4-flash", or add the model to the LiteLLM proxy's config.yaml and send requests to /v1/chat/completions from any OpenAI client. LiteLLM tracks the cost of every request at the prices on this page.

### What features does DeepSeek V4 Flash support?

DeepSeek V4 Flash supports function calling, parallel tool calls, structured output, image input, reasoning, prompt caching and web search.

### When did LiteLLM add DeepSeek V4 Flash?

DeepSeek V4 Flash was added to LiteLLM's model price file in June 2026.

Source: LiteLLM model price file, https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json
