LiteLLMModels

Claude Opus 4.8 Fast API Pricing

Claude Opus 4.8 Fast is available through LiteLLM. It has a 200,000-token context window and returns up to 64,000 tokens per response.

By AnthropicModel ID github_copilot/claude-opus-4.8-fastPrices updated Checked
Context window
200,000 tokens
Max output
64,000 tokens
Providers
1

Claude Opus 4.8 Fast pricing by provider

1 listing across 1 provider. Token prices are in US dollars per 1M tokens. The highlighted row is the price quoted above.

ProviderModel name in LiteLLMContext
GitHub Copilotgithub_copilot/claude-opus-4.8-fast200K

Claude Opus 4.8 Fast features

Use Claude Opus 4.8 Fast with LiteLLM

Call Claude Opus 4.8 Fast through the LiteLLM Python SDK, or put it behind the LiteLLM proxy and send OpenAI-format requests to /v1/chat/completions. LiteLLM tracks the cost of every request at these prices.

Python SDK
from litellm import completion

response = completion(
    model="github_copilot/claude-opus-4.8-fast",
    messages=[{"role": "user", "content": "Hello!"}],
)
Call the proxy
curl http://0.0.0.0:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-1234" \
  -d '{
    "model": "claude-opus-4.8-fast",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Claude Opus 4.8 Fast FAQ

What is the context window of Claude Opus 4.8 Fast?

Claude Opus 4.8 Fast has a 200,000-token context window and returns up to 64,000 tokens per response.

What features does Claude Opus 4.8 Fast support?

Claude Opus 4.8 Fast supports function calling, parallel tool calls, structured output, image input and reasoning.

Sources

Prices come from LiteLLM's open-source model_prices_and_context_window.json, which LiteLLM uses to calculate the cost of every request.

LiteLLM docs: GitHub Copilot.

Provider pages: GitHub Copilot.