LiteLLMModels

Claude 3.5 Haiku API Pricing

As of October 10, 2026, Claude 3.5 Haiku costs $0.80 per 1M input tokens and $4.00 per 1M output tokens from GradientAI, the lowest of 6 providers. It has a 200,000-token context window and returns up to 1,024 tokens per response.

By AnthropicModel ID gradient_ai/anthropic-claude-3.5-haikuPrices updated Checked
Input
$0.80 / 1M tokens
Output
$4.00 / 1M tokens
Context window
200,000 tokens
Max output
1,024 tokens
Providers
6
Added to LiteLLM
November 2024

Claude 3.5 Haiku pricing by provider

10 listings across 6 providers. Token prices are in US dollars per 1M tokens. The highlighted row is the price quoted above.

ProviderModel name in LiteLLMInputOutputCache readCache writeContext
GradientAIgradient_ai/anthropic-claude-3.5-haiku$0.80$4.00––200K
Amazon Bedrocksnapshot 20241022anthropic.claude-3-5-haiku-20241022-v1:0$0.80$4.00$0.08$1.00200K
Google Vertex AIvertex_ai/claude-3-5-haiku$1.00$5.00––200K
Google Vertex AIvertex_ai/claude-3-5-haiku@20241022$1.00$5.00––200K
Herokuheroku/claude-3-5-haiku––––200K
Replicatereplicate/anthropic/claude-3.5-haiku$1.00$5.00–––
Vercel AI Gatewayvercel_ai_gateway/anthropic/claude-3.5-haiku$0.80$4.00$0.08$1.00200K
Amazon BedrockUS cross-region, snapshot 20241022bedrock/us.anthropic.claude-3-5-haiku-20241022-v1:0$0.80$4.00$0.08$1.00200K
Amazon BedrockEU cross-region, snapshot 20241022eu.anthropic.claude-3-5-haiku-20241022-v1:0$0.80$4.00$0.08$1.00200K
Amazon BedrockUS cross-region, snapshot 20241022us.anthropic.claude-3-5-haiku-20241022-v1:0$0.80$4.00$0.08$1.00200K

What Claude 3.5 Haiku costs in practice

1,000 requests with 2,000 input and 500 output tokens each

GradientAI$3.60
Amazon Bedrock$3.60
Vercel AI Gateway$3.60
Google Vertex AI$4.50

100 long-document requests with 100,000 input and 2,000 output tokens each

GradientAI$8.80
Amazon Bedrock$8.80
Vercel AI Gateway$8.80
Google Vertex AI$11.00

Claude 3.5 Haiku features

Input: text.

Use Claude 3.5 Haiku with LiteLLM

Call Claude 3.5 Haiku through the LiteLLM Python SDK, or put it behind the LiteLLM proxy and send OpenAI-format requests to /v1/chat/completions. LiteLLM tracks the cost of every request at these prices.

Python SDK
from litellm import completion

response = completion(
    model="gradient_ai/anthropic-claude-3.5-haiku",
    messages=[{"role": "user", "content": "Hello!"}],
)
Proxy config.yaml
model_list:
  - model_name: claude-3-5-haiku
    litellm_params:
      model: gradient_ai/anthropic-claude-3.5-haiku
      # credentials: see the provider's LiteLLM docs page
Call the proxy
curl http://0.0.0.0:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-1234" \
  -d '{
    "model": "claude-3-5-haiku",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Claude 3.5 Haiku FAQ

How much does Claude 3.5 Haiku cost?

Claude 3.5 Haiku costs $0.80 per 1M input tokens and $4.00 per 1M output tokens from GradientAI, the lowest of 6 providers.

What is the context window of Claude 3.5 Haiku?

Claude 3.5 Haiku has a 200,000-token context window and returns up to 1,024 tokens per response.

Which providers offer Claude 3.5 Haiku?

Claude 3.5 Haiku is available from 6 providers through LiteLLM: GradientAI, Amazon Bedrock, Google Vertex AI, Heroku, Replicate and Vercel AI Gateway. The lowest price is on GradientAI.

How do I call Claude 3.5 Haiku with an OpenAI-compatible API?

Use the LiteLLM Python SDK with model="gradient_ai/anthropic-claude-3.5-haiku", or add the model to the LiteLLM proxy's config.yaml and send requests to /v1/chat/completions from any OpenAI client. LiteLLM tracks the cost of every request at the prices on this page.

What features does Claude 3.5 Haiku support?

Claude 3.5 Haiku supports function calling, parallel tool calls, structured output, image input, PDF input and prompt caching.

When did LiteLLM add Claude 3.5 Haiku?

Claude 3.5 Haiku was added to LiteLLM's model price file in November 2024.

Sources

Prices come from LiteLLM's open-source model_prices_and_context_window.json, which LiteLLM uses to calculate the cost of every request.

LiteLLM docs: GradientAI, Amazon Bedrock, Google Vertex AI, Heroku, Replicate, Vercel AI Gateway.

Provider pages: GradientAI, Amazon Bedrock, Google Vertex AI, Heroku, Replicate, Vercel AI Gateway.