LiteLLMModels

Llama 4 Scout 17B 16e Instruct API Pricing

As of October 10, 2026, Llama 4 Scout 17B 16e Instruct costs $0.05 per 1M input tokens and $0.10 per 1M output tokens from Lambda, the lowest of 10 providers. It has a 16,384-token context window and returns up to 8,192 tokens per response.

By MetaModel ID lambda_ai/llama-4-scout-17b-16e-instructPrices updated Checked
Input
$0.05 / 1M tokens
Output
$0.10 / 1M tokens
Context window
16,384 tokens
Max output
8,192 tokens
Providers
10
Added to LiteLLM
April 2025

Llama 4 Scout 17B 16e Instruct pricing by provider

10 listings across 10 providers. Token prices are in US dollars per 1M tokens. The highlighted row is the price quoted above.

ProviderModel name in LiteLLMInputOutputBatch inBatch outContext
Lambdalambda_ai/llama-4-scout-17b-16e-instruct$0.05$0.10––16K
Meta Llama APIfp8meta_llama/Llama-4-Scout-17B-16E-Instruct-FP8––––10M
Azure AI Foundryazure_ai/Llama-4-Scout-17B-16E-Instruct$0.20$0.78––10M
Cloudflare Workers AIcloudflare/@cf/meta/llama-4-scout-17b-16e-instruct$0.27$0.85––131K
DeepInfradeepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct$0.10$0.30––320K
Google Vertex AIvertex_ai/meta/llama-4-scout-17b-16e-instruct-maas$0.25$0.70$0.125$0.3510M
Novita AInovita/meta-llama/llama-4-scout-17b-16e-instruct$0.18$0.59––128K
Nscalenscale/meta-llama/Llama-4-Scout-17B-16E-Instruct$0.09$0.29–––
Oracle Cloud (OCI)oci/meta.llama-4-scout-17b-16e-instruct$0.72$0.72––10.5M
Together AItogether_ai/meta-llama/Llama-4-Scout-17B-16E-Instruct$0.18$0.59––1M

What Llama 4 Scout 17B 16e Instruct costs in practice

1,000 requests with 2,000 input and 500 output tokens each

Lambda$0.15
Nscale$0.325
DeepInfra$0.35
Novita AI$0.655

Llama 4 Scout 17B 16e Instruct features

Use Llama 4 Scout 17B 16e Instruct with LiteLLM

Call Llama 4 Scout 17B 16e Instruct through the LiteLLM Python SDK, or put it behind the LiteLLM proxy and send OpenAI-format requests to /v1/chat/completions. LiteLLM tracks the cost of every request at these prices.

Python SDK
from litellm import completion

response = completion(
    model="lambda_ai/llama-4-scout-17b-16e-instruct",
    messages=[{"role": "user", "content": "Hello!"}],
)
Proxy config.yaml
model_list:
  - model_name: llama-4-scout-17b-16e-instruct
    litellm_params:
      model: lambda_ai/llama-4-scout-17b-16e-instruct
      # credentials: see the provider's LiteLLM docs page
Call the proxy
curl http://0.0.0.0:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-1234" \
  -d '{
    "model": "llama-4-scout-17b-16e-instruct",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Llama 4 Scout 17B 16e Instruct FAQ

How much does Llama 4 Scout 17B 16e Instruct cost?

Llama 4 Scout 17B 16e Instruct costs $0.05 per 1M input tokens and $0.10 per 1M output tokens from Lambda, the lowest of 10 providers.

What is the context window of Llama 4 Scout 17B 16e Instruct?

Llama 4 Scout 17B 16e Instruct has a 16,384-token context window and returns up to 8,192 tokens per response.

Which providers offer Llama 4 Scout 17B 16e Instruct?

Llama 4 Scout 17B 16e Instruct is available from 10 providers through LiteLLM: Lambda, Meta Llama API, Azure AI Foundry, Cloudflare Workers AI, DeepInfra, Google Vertex AI, Novita AI, Nscale, Oracle Cloud (OCI) and Together AI. The lowest price is on Lambda.

How do I call Llama 4 Scout 17B 16e Instruct with an OpenAI-compatible API?

Use the LiteLLM Python SDK with model="lambda_ai/llama-4-scout-17b-16e-instruct", or add the model to the LiteLLM proxy's config.yaml and send requests to /v1/chat/completions from any OpenAI client. LiteLLM tracks the cost of every request at the prices on this page.

What features does Llama 4 Scout 17B 16e Instruct support?

Llama 4 Scout 17B 16e Instruct supports function calling, parallel tool calls and image input.

When did LiteLLM add Llama 4 Scout 17B 16e Instruct?

Llama 4 Scout 17B 16e Instruct was added to LiteLLM's model price file in April 2025.

Sources

Prices come from LiteLLM's open-source model_prices_and_context_window.json, which LiteLLM uses to calculate the cost of every request. Provider pricing pages:

LiteLLM docs: Lambda, Meta Llama API, Azure AI Foundry, Cloudflare Workers AI, DeepInfra, Google Vertex AI.

Provider pages: Lambda, Meta Llama API, Azure AI Foundry, Cloudflare Workers AI, DeepInfra, Google Vertex AI, Novita AI, Nscale, Oracle Cloud (OCI), Together AI.