GLM 5 API Pricing
As of October 10, 2026, GLM 5 costs $1.00 per 1M input tokens and $3.20 per 1M output tokens on Z.ai, with cache reads at $0.20 and cache writes at $0 per 1M tokens. Across 7 providers on LiteLLM, the input price ranges from $0.60 to $1.00. It has a 200,000-token context window and returns up to 128,000 tokens per response.
- Input
- $1.00 / 1M tokens
- Output
- $3.20 / 1M tokens
- Cache read
- $0.20 / 1M tokens
- Context window
- 200,000 tokens
- Max output
- 128,000 tokens
- Providers
- 7
- Added to LiteLLM
- February 2026
GLM 5 pricing by provider
9 listings across 7 providers. Token prices are in US dollars per 1M tokens. The highlighted row is the price quoted above.
| Provider | Model name in LiteLLM | Input | Output | Cache read | Cache write | Context |
|---|---|---|---|---|---|---|
| Z.ai | zai/glm-5 | $1.00 | $3.20 | $0.20 | $0 | 200K |
| Amazon Bedrock | zai.glm-5 | $1.00 | $3.20 | – | – | 200K |
| Baseten | baseten/zai-org/GLM-5 | $0.95 | $3.15 | – | – | – |
| DeepInfra | deepinfra/zai-org/GLM-5 | $0.60 | $2.08 | $0.12 | – | 198K |
| Google Vertex AI | vertex_ai/zai-org/glm-5-maas | $1.00 | $3.20 | $0.10 | – | 200K |
| Novita AI | novita/zai-org/glm-5 | $1.00 | $3.20 | $0.20 | – | 203K |
| Together AI | together_ai/zai-org/GLM-5 | $1.00 | $3.20 | – | – | 198K |
| Amazon Bedrockus-east-1 | bedrock/us-east-1/zai.glm-5 | $1.00 | $3.20 | – | – | 200K |
| Amazon Bedrockus-west-2 | bedrock/us-west-2/zai.glm-5 | $1.00 | $3.20 | – | – | 200K |
What GLM 5 costs in practice
1,000 requests with 2,000 input and 500 output tokens each
| Z.ai | $3.60 |
| DeepInfra | $2.24 |
| Baseten | $3.475 |
| Amazon Bedrock | $3.60 |
100 long-document requests with 100,000 input and 2,000 output tokens each
| Z.ai | $10.64 |
| DeepInfra | $6.416 |
| Baseten | $10.13 |
| Amazon Bedrock | $10.64 |
GLM 5 features
- Function calling
- Structured output
- Reasoning
- Prompt caching
Use GLM 5 with LiteLLM
Call GLM 5 through the LiteLLM Python SDK, or put it behind the LiteLLM proxy and send OpenAI-format requests to /v1/chat/completions. LiteLLM tracks the cost of every request at these prices.
from litellm import completion
response = completion(
model="zai/glm-5",
messages=[{"role": "user", "content": "Hello!"}],
)model_list:
- model_name: glm-5
litellm_params:
model: zai/glm-5
# credentials: see the provider's LiteLLM docs pagecurl http://0.0.0.0:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-1234" \
-d '{
"model": "glm-5",
"messages": [{"role": "user", "content": "Hello!"}]
}'GLM 5 FAQ
How much does GLM 5 cost?
GLM 5 costs $1.00 per 1M input tokens and $3.20 per 1M output tokens on Z.ai, with cache reads at $0.20 and cache writes at $0 per 1M tokens. Across 7 providers on LiteLLM, the input price ranges from $0.60 to $1.00.
What is the context window of GLM 5?
GLM 5 has a 200,000-token context window and returns up to 128,000 tokens per response.
Which providers offer GLM 5?
GLM 5 is available from 7 providers through LiteLLM: Z.ai, Amazon Bedrock, Baseten, DeepInfra, Google Vertex AI, Novita AI and Together AI. The lowest price is on DeepInfra.
How do I call GLM 5 with an OpenAI-compatible API?
Use the LiteLLM Python SDK with model="zai/glm-5", or add the model to the LiteLLM proxy's config.yaml and send requests to /v1/chat/completions from any OpenAI client. LiteLLM tracks the cost of every request at the prices on this page.
What features does GLM 5 support?
GLM 5 supports function calling, structured output, reasoning and prompt caching.
When did LiteLLM add GLM 5?
GLM 5 was added to LiteLLM's model price file in February 2026.
Sources
Prices come from LiteLLM's open-source model_prices_and_context_window.json, which LiteLLM uses to calculate the cost of every request. Provider pricing pages:
- docs.z.ai/guides/overview/pricing
- aws.amazon.com/bedrock/pricing/
- deepinfra.com/pricing
- cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing
LiteLLM docs: Z.ai, Amazon Bedrock, Baseten, DeepInfra, Google Vertex AI, Together AI.
Provider pages: Z.ai, Amazon Bedrock, Baseten, DeepInfra, Google Vertex AI, Novita AI, Together AI.
