# Gemma 4 31B IT API Pricing

As of October 10, 2026, Gemma 4 31B IT is free to use on Google AI Studio. Across 11 providers on LiteLLM, the input price ranges from $0 to $0.40. It has a 262,144-token context window and returns up to 32,768 tokens per response.

- Page: https://models.litellm.ai/models/gemma-4-31b-it
- Lab: Google
- LiteLLM model name: `gemini/gemma-4-31b-it`
- Context window: 262,144 tokens
- Max output: 32,768 tokens
- Added to LiteLLM: June 2026
- Prices updated: October 10, 2026

## Pricing by provider

Token prices in USD per 1M tokens.

| Provider | Model name | Input | Output | Cache read | Per image | Per second | Context |
|---|---|---|---|---|---|---|---|
| Google AI Studio | `gemini/gemma-4-31b-it` | $0 | $0 | – | – | – | 256K |
| AIHubMix | `aihubmix/gemma-4-31b-it` | $0.14 | $0.40 | – | – | – | 256K |
| DeepInfra | `deepinfra/google/gemma-4-31B-it` | $0.13 | $0.38 | – | – | – | 256K |
| FriendliAI | `friendliai/google/gemma-4-31B-it` | $0.14 | $0.40 | – | – | – | 256K |
| LibertAI | `libertai/gemma-4-31b-it` | $0.15 | $0.40 | – | – | – | 256K |
| Novita AI | `novita/google/gemma-4-31b-it` | $0.14 | $0.40 | – | – | – | 256K |
| Sail | `sail/google/gemma-4-31B-it` | $0.40 | $0.60 | $0.20 | – | – | 256K |
| SambaNova | `sambanova/gemma-4-31B-it` | $0.38 | $1.15 | – | – | – | 128K |
| Tensormesh | `tensormesh/google/gemma-4-31B-it` | $0.14 | $0.56 | $0 | – | – | 32K |
| Together AI | `together_ai/google/gemma-4-31B-it` | $0.39 | $0.97 | – | – | – | 256K |
| Weights & Biases | `wandb/google/gemma-4-31B-it` | $0.10 | $0.34 | – | – | – | 262K |
| Sail (nvfp4) | `sail/nvidia/Gemma-4-31B-IT-NVFP4` | $0.14 | $0.40 | $0.07 | – | – | 256K |

## Cost examples

1,000 requests with 2,000 input and 500 output tokens each:

- Google AI Studio: $0
- Weights & Biases: $0.37
- DeepInfra: $0.45
- AIHubMix: $0.48

100 long-document requests with 100,000 input and 2,000 output tokens each:

- Google AI Studio: $0
- Weights & Biases: $1.068
- DeepInfra: $1.376
- AIHubMix: $1.48

## Features

Supported: Function calling, Parallel tool calls, Structured output, Image input, Reasoning, Prompt caching.

## Use it with LiteLLM

```python
from litellm import completion

response = completion(
    model="gemini/gemma-4-31b-it",
    messages=[{"role": "user", "content": "Hello!"}],
)
```

## FAQ

### How much does Gemma 4 31B IT cost?

Gemma 4 31B IT is free to use on Google AI Studio. Across 11 providers on LiteLLM, the input price ranges from $0 to $0.40.

### What is the context window of Gemma 4 31B IT?

Gemma 4 31B IT has a 262,144-token context window and returns up to 32,768 tokens per response.

### Which providers offer Gemma 4 31B IT?

Gemma 4 31B IT is available from 11 providers through LiteLLM: Google AI Studio, AIHubMix, DeepInfra, FriendliAI, LibertAI, Novita AI, Sail, SambaNova, Tensormesh, Together AI and Weights & Biases. The lowest price is on Google AI Studio.

### How do I call Gemma 4 31B IT with an OpenAI-compatible API?

Use the LiteLLM Python SDK with model="gemini/gemma-4-31b-it", or add the model to the LiteLLM proxy's config.yaml and send requests to /v1/chat/completions from any OpenAI client. LiteLLM tracks the cost of every request at the prices on this page.

### What features does Gemma 4 31B IT support?

Gemma 4 31B IT supports function calling, parallel tool calls, structured output, image input, reasoning and prompt caching.

### When did LiteLLM add Gemma 4 31B IT?

Gemma 4 31B IT was added to LiteLLM's model price file in June 2026.

Source: LiteLLM model price file, https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json
