LiteLLMModels

GPT-4.1 Nano API Pricing

As of October 10, 2026, GPT-4.1 Nano costs $0.10 per 1M input tokens and $0.40 per 1M output tokens on OpenAI, with cache reads at $0.025 per 1M tokens. It has a 1,047,576-token context window and returns up to 32,768 tokens per response.

By OpenAIModel ID gpt-4.1-nanoPrices updated Checked
Scheduled for deprecation: OpenAI lists October 23, 2026 as the deprecation date for gpt-4.1-nano.
Input
$0.10 / 1M tokens
Output
$0.40 / 1M tokens
Cache read
$0.025 / 1M tokens
Context window
1,047,576 tokens
Max output
32,768 tokens
Providers
4
Added to LiteLLM
April 2025

GPT-4.1 Nano pricing by provider

10 listings across 4 providers. Token prices are in US dollars per 1M tokens. The highlighted row is the price quoted above.

ProviderModel name in LiteLLMInputOutputCache readCache writeBatch inBatch outContext
OpenAIgpt-4.1-nano$0.10$0.40$0.025–$0.05$0.201M
OpenAIsnapshot 2025-04-14gpt-4.1-nano-2025-04-14$0.10$0.40$0.025–$0.05$0.201M
OpenAIfine-tuned, snapshot 2025-04-14ft:gpt-4.1-nano-2025-04-14$0.20$0.80$0.05–$0.10$0.401M
Azure OpenAIazure/gpt-4.1-nano$0.10$0.40$0.025–$0.05$0.201M
Azure OpenAIsnapshot 2025-04-14azure/gpt-4.1-nano-2025-04-14$0.10$0.40$0.025–$0.05$0.201M
Replicatereplicate/openai/gpt-4.1-nano$0.10$0.40–––––
Vercel AI Gatewayvercel_ai_gateway/openai/gpt-4.1-nano$0.10$0.40$0.025$0––1M
Azure OpenAIeuazure/eu/gpt-4.1-nano$0.11$0.44$0.028–$0.055$0.22–
Azure OpenAIusazure/us/gpt-4.1-nano$0.11$0.44$0.028–$0.055$0.22–
Azure OpenAIus, snapshot 2025-04-14azure/us/gpt-4.1-nano-2025-04-14$0.11$0.44$0.028–$0.055$0.221M

Other GPT-4.1 Nano prices on OpenAI

Priority input$0.20 per 1M tokens
Priority output$0.80 per 1M tokens

What GPT-4.1 Nano costs in practice

1,000 requests with 2,000 input and 500 output tokens each

OpenAI$0.40
Azure OpenAI$0.40
Replicate$0.40
Vercel AI Gateway$0.40
OpenAI batch$0.20

100 long-document requests with 100,000 input and 2,000 output tokens each

OpenAI$1.08
Azure OpenAI$1.08
Replicate$1.08
Vercel AI Gateway$1.08

GPT-4.1 Nano features

Input: text, image. Output: text.

Use GPT-4.1 Nano with LiteLLM

Call GPT-4.1 Nano through the LiteLLM Python SDK, or put it behind the LiteLLM proxy and send OpenAI-format requests to /v1/chat/completions. LiteLLM tracks the cost of every request at these prices.

Python SDK
from litellm import completion

response = completion(
    model="gpt-4.1-nano",
    messages=[{"role": "user", "content": "Hello!"}],
)
Proxy config.yaml
model_list:
  - model_name: gpt-4.1-nano
    litellm_params:
      model: gpt-4.1-nano
      api_key: os.environ/OPENAI_API_KEY
Call the proxy
curl http://0.0.0.0:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-1234" \
  -d '{
    "model": "gpt-4.1-nano",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

GPT-4.1 Nano FAQ

How much does GPT-4.1 Nano cost?

GPT-4.1 Nano costs $0.10 per 1M input tokens and $0.40 per 1M output tokens on OpenAI, with cache reads at $0.025 per 1M tokens. Batch requests cost $0.05 per 1M input tokens and $0.20 per 1M output tokens.

What is the context window of GPT-4.1 Nano?

GPT-4.1 Nano has a 1,047,576-token context window and returns up to 32,768 tokens per response.

Which providers offer GPT-4.1 Nano?

GPT-4.1 Nano is available from 4 providers through LiteLLM: OpenAI, Azure OpenAI, Replicate and Vercel AI Gateway.

How do I call GPT-4.1 Nano with an OpenAI-compatible API?

Use the LiteLLM Python SDK with model="gpt-4.1-nano", or add the model to the LiteLLM proxy's config.yaml and send requests to /v1/chat/completions from any OpenAI client. LiteLLM tracks the cost of every request at the prices on this page.

What features does GPT-4.1 Nano support?

GPT-4.1 Nano supports function calling, parallel tool calls, structured output, image input, PDF input and prompt caching.

When did LiteLLM add GPT-4.1 Nano?

GPT-4.1 Nano was added to LiteLLM's model price file in April 2025.

Is GPT-4.1 Nano being deprecated?

OpenAI lists a deprecation date of October 23, 2026 for gpt-4.1-nano.

Sources

Prices come from LiteLLM's open-source model_prices_and_context_window.json, which LiteLLM uses to calculate the cost of every request. Provider pricing pages:

LiteLLM docs: OpenAI, Azure OpenAI, Replicate, Vercel AI Gateway.

Provider pages: OpenAI, Azure OpenAI, Replicate, Vercel AI Gateway.