LiteLLMModels

Llama3.2 11B Vision Instruct API Pricing

As of October 10, 2026, Llama3.2 11B Vision Instruct costs $0.015 per 1M input tokens and $0.025 per 1M output tokens from Lambda, the lowest of 5 providers. It has a 131,072-token context window and returns up to 131,072 tokens per response.

By MetaModel ID lambda_ai/llama3.2-11b-vision-instructPrices updated Checked
Input
$0.015 / 1M tokens
Output
$0.025 / 1M tokens
Context window
131,072 tokens
Max output
131,072 tokens
Providers
5
Added to LiteLLM
November 2024

Llama3.2 11B Vision Instruct pricing by provider

5 listings across 5 providers. Token prices are in US dollars per 1M tokens. The highlighted row is the price quoted above.

ProviderModel name in LiteLLMInputOutputContext
Lambdalambda_ai/llama3.2-11b-vision-instruct$0.015$0.025128K
Cloudflare Workers AIcloudflare/@cf/meta/llama-3.2-11b-vision-instruct$0.0485$0.676128K
DeepInfradeepinfra/meta-llama/Llama-3.2-11B-Vision-Instruct$0.049$0.049128K
IBM watsonx.aiwatsonx/meta-llama/llama-3-2-11b-vision-instruct$0.35$0.35128K
Oracle Cloud (OCI) †oci/meta.llama-3.2-11b-vision-instruct$2.00$2.00128K

† This price differs from the other providers by more than 8x and is left out of the headline price. Check the provider's pricing page.

What Llama3.2 11B Vision Instruct costs in practice

1,000 requests with 2,000 input and 500 output tokens each

Lambda$0.0425
Cloudflare Workers AI$0.435
DeepInfra$0.1225
IBM watsonx.ai$0.875

100 long-document requests with 100,000 input and 2,000 output tokens each

Lambda$0.155
Cloudflare Workers AI$0.6202
DeepInfra$0.4998
IBM watsonx.ai$3.57

Llama3.2 11B Vision Instruct features

Use Llama3.2 11B Vision Instruct with LiteLLM

Call Llama3.2 11B Vision Instruct through the LiteLLM Python SDK, or put it behind the LiteLLM proxy and send OpenAI-format requests to /v1/chat/completions. LiteLLM tracks the cost of every request at these prices.

Python SDK
from litellm import completion

response = completion(
    model="lambda_ai/llama3.2-11b-vision-instruct",
    messages=[{"role": "user", "content": "Hello!"}],
)
Proxy config.yaml
model_list:
  - model_name: llama-3.2-11b-vision-instruct
    litellm_params:
      model: lambda_ai/llama3.2-11b-vision-instruct
      # credentials: see the provider's LiteLLM docs page
Call the proxy
curl http://0.0.0.0:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-1234" \
  -d '{
    "model": "llama-3.2-11b-vision-instruct",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Llama3.2 11B Vision Instruct FAQ

How much does Llama3.2 11B Vision Instruct cost?

Llama3.2 11B Vision Instruct costs $0.015 per 1M input tokens and $0.025 per 1M output tokens from Lambda, the lowest of 5 providers.

What is the context window of Llama3.2 11B Vision Instruct?

Llama3.2 11B Vision Instruct has a 131,072-token context window and returns up to 131,072 tokens per response.

Which providers offer Llama3.2 11B Vision Instruct?

Llama3.2 11B Vision Instruct is available from 5 providers through LiteLLM: Lambda, Cloudflare Workers AI, DeepInfra, IBM watsonx.ai and Oracle Cloud (OCI). The lowest price is on Lambda.

How do I call Llama3.2 11B Vision Instruct with an OpenAI-compatible API?

Use the LiteLLM Python SDK with model="lambda_ai/llama3.2-11b-vision-instruct", or add the model to the LiteLLM proxy's config.yaml and send requests to /v1/chat/completions from any OpenAI client. LiteLLM tracks the cost of every request at the prices on this page.

What features does Llama3.2 11B Vision Instruct support?

Llama3.2 11B Vision Instruct supports function calling, parallel tool calls and image input.

When did LiteLLM add Llama3.2 11B Vision Instruct?

Llama3.2 11B Vision Instruct was added to LiteLLM's model price file in November 2024.

Sources

Prices come from LiteLLM's open-source model_prices_and_context_window.json, which LiteLLM uses to calculate the cost of every request. Provider pricing pages:

LiteLLM docs: Lambda, Cloudflare Workers AI, DeepInfra, IBM watsonx.ai, Oracle Cloud (OCI).

Provider pages: Lambda, Cloudflare Workers AI, DeepInfra, IBM watsonx.ai, Oracle Cloud (OCI).