> ## Documentation Index
> Fetch the complete documentation index at: https://prices.voicegateway.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# LLM pricing

> 1095 LLM model(s) from 31 provider(s), largest catalogs first.

1095 LLM model(s) from 31 provider(s). Priced from `input_mtok` and `output_mtok`, US dollars per 1,000,000 tokens. Pick a vendor in the sidebar to see that vendor alone.

LLM rows carry no per-model verification status. Unlike voice rates, they are cross-checked against external aggregators rather than against one vendor page each. [How Fresh?](/how-fresh) explains the difference.

| Provider                                                    | Models | Cheapest input \$ / Mtok | Cheapest output \$ / Mtok |
| ----------------------------------------------------------- | -----: | -----------------------: | ------------------------: |
| [OpenRouter](/llm/openrouter)                               |    461 |                  \$0.005 |                    \$0.01 |
| [Together AI](/llm/together)                                |     72 |                    \$0.1 |                     \$0.1 |
| [AWS Bedrock](/llm/aws)                                     |     69 |                  \$0.035 |                    \$0.04 |
| [OpenAI](/llm/openai)                                       |     63 |                   \$0.02 |                     \$0.4 |
| [HuggingFace (novita)](/llm/huggingface-novita)             |     61 |                   \$0.02 |                    \$0.04 |
| [Novita](/llm/novita)                                       |     34 |                   \$0.02 |                    \$0.02 |
| [LiveKit Inference](/llm/livekit)                           |     33 |                   \$0.05 |                     \$0.4 |
| [Google](/llm/google)                                       |     31 |                 \$0.0375 |                    \$0.15 |
| [Groq](/llm/groq)                                           |     29 |                   \$0.04 |                    \$0.04 |
| [HuggingFace (nebius)](/llm/huggingface-nebius)             |     26 |                   \$0.02 |                    \$0.06 |
| [HuggingFace (together)](/llm/huggingface-together)         |     23 |                   \$0.02 |                    \$0.04 |
| [HuggingFace (nscale)](/llm/huggingface-nscale)             |     20 |                   \$0.01 |                    \$0.03 |
| [Anthropic](/llm/anthropic)                                 |     18 |                   \$0.25 |                    \$1.25 |
| [Microsoft Azure](/llm/azure)                               |     18 |                   \$0.02 |                     \$0.1 |
| [Mistral](/llm/mistral)                                     |     18 |                   \$0.04 |                    \$0.04 |
| [OVHcloud AI Endpoints](/llm/ovhcloud)                      |     15 |                   \$0.01 |                    \$0.11 |
| [Fireworks](/llm/fireworks)                                 |     13 |                   \$0.07 |                     \$0.1 |
| [HuggingFace (hyperbolic)](/llm/huggingface-hyperbolic)     |     12 |                    \$0.1 |                     \$0.1 |
| [X AI](/llm/x-ai)                                           |     12 |                    \$0.2 |                     \$0.5 |
| [MoonshotAi](/llm/moonshotai)                               |      9 |                    \$0.2 |                       \$2 |
| [HuggingFace (publicai)](/llm/huggingface-publicai)         |      8 |                    \$0.1 |                     \$0.2 |
| [HuggingFace (sambanova)](/llm/huggingface-sambanova)       |      8 |                    \$0.1 |                     \$0.2 |
| [Perplexity](/llm/perplexity)                               |      8 |                    \$0.2 |                     \$0.2 |
| [HuggingFace (ovhcloud)](/llm/huggingface-ovhcloud)         |      7 |                   \$0.05 |                    \$0.11 |
| [Cohere](/llm/cohere)                                       |      6 |                 \$0.0375 |                    \$0.15 |
| [HuggingFace (groq)](/llm/huggingface-groq)                 |      5 |                    \$0.1 |                    \$0.34 |
| [Avian](/llm/avian)                                         |      4 |                    \$0.1 |                     \$0.1 |
| [Cerebras](/llm/cerebras)                                   |      4 |                    \$0.1 |                     \$0.1 |
| [Deepseek](/llm/deepseek)                                   |      4 |                  \$0.135 |                    \$0.28 |
| [HuggingFace (fireworks-ai)](/llm/huggingface-fireworks-ai) |      3 |                   \$0.05 |                     \$0.2 |
| [HuggingFace (cerebras)](/llm/huggingface-cerebras)         |      1 |                    \$0.1 |                     \$0.1 |

## Via LiveKit Inference

What the same model costs going direct to the vendor versus through the gateway. Scale is the discounted tier.

| Model                             | Name                    | Direct (input \$ / Mtok) | LiveKit (input \$ / Mtok) | Scale (input \$ / Mtok) | vs direct |
| --------------------------------- | ----------------------- | -----------------------: | ------------------------: | ----------------------: | --------: |
| `google/gemini-2.5-pro`           | Gemini 2.5 Pro          |                   \$1.25 |                     \$2.5 |                       - |   +100.0% |
| `google/gemini-3.1-pro-preview`   | Gemini 3.1 Pro          |                      \$2 |                       \$4 |                       - |   +100.0% |
| `openai/gpt-5.4`                  | GPT-5.4                 |                    \$2.5 |                       \$5 |                       - |   +100.0% |
| `openai/gpt-5.5`                  | GPT-5.5                 |                      \$5 |                      \$10 |                       - |   +100.0% |
| `google/gemini-2.5-flash`         | Gemini 2.5 Flash        |                    \$0.3 |                     \$0.3 |                       - |      same |
| `google/gemini-2.5-flash-lite`    | Gemini 2.5 Flash-Lite   |                    \$0.1 |                     \$0.1 |                       - |      same |
| `google/gemini-3-flash-preview`   | Gemini 3 Flash          |                    \$0.5 |                     \$0.5 |                       - |      same |
| `google/gemini-3.1-flash-lite`    | Gemini 3.1 Flash Lite   |                   \$0.25 |                    \$0.25 |                       - |      same |
| `google/gemini-3.5-flash`         | Gemini 3.5 Flash        |                    \$1.5 |                     \$1.5 |                       - |      same |
| `moonshotai/kimi-k2.5`            | Kimi K2.5               |                    \$0.6 |                     \$0.6 |                       - |      same |
| `openai/gpt-4.1`                  | GPT-4.1                 |                      \$2 |                       \$2 |                       - |      same |
| `openai/gpt-4.1-mini`             | GPT-4.1 mini            |                    \$0.4 |                     \$0.4 |                       - |      same |
| `openai/gpt-4.1-nano`             | GPT-4.1 nano            |                    \$0.1 |                     \$0.1 |                       - |      same |
| `openai/gpt-4o`                   | GPT-4o                  |                    \$2.5 |                     \$2.5 |                       - |      same |
| `openai/gpt-4o-mini`              | GPT-4o mini             |                   \$0.15 |                    \$0.15 |                       - |      same |
| `openai/gpt-5`                    | GPT-5                   |                   \$1.25 |                    \$1.25 |                       - |      same |
| `openai/gpt-5-mini`               | GPT-5 mini              |                   \$0.25 |                    \$0.25 |                       - |      same |
| `openai/gpt-5-nano`               | GPT-5 nano              |                   \$0.05 |                    \$0.05 |                       - |      same |
| `openai/gpt-5.1`                  | GPT-5.1                 |                   \$1.25 |                    \$1.25 |                       - |      same |
| `openai/gpt-5.1-chat-latest`      | GPT-5.1 Chat            |                   \$1.25 |                    \$1.25 |                       - |      same |
| `openai/gpt-5.2`                  | GPT-5.2                 |                   \$1.75 |                    \$1.75 |                       - |      same |
| `openai/gpt-5.2-chat-latest`      | GPT-5.2 Chat            |                   \$1.75 |                    \$1.75 |                       - |      same |
| `openai/gpt-5.3-chat-latest`      | GPT-5.3 Chat            |                   \$1.75 |                    \$1.75 |                       - |      same |
| `openai/gpt-5.4-mini`             | GPT-5.4 mini            |                   \$0.75 |                    \$0.75 |                       - |      same |
| `openai/gpt-5.4-nano`             | GPT-5.4 nano            |                    \$0.2 |                     \$0.2 |                       - |      same |
| `xai/grok-4-1-fast-non-reasoning` | Grok 4.1 Fast           |                    \$0.2 |                     \$0.2 |                       - |      same |
| `xai/grok-4-1-fast-reasoning`     | Grok 4.1 Fast Reasoning |                    \$0.2 |                     \$0.2 |                       - |      same |

6 further LiveKit LLM model(s) have no direct-vendor entry to compare against.
