> ## Documentation Index
> Fetch the complete documentation index at: https://prices.voicegateway.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# HuggingFace (nscale) LLM pricing

> 20 LLM model(s) from HuggingFace (nscale), with input and output token rates.

Priced from `input_mtok` and `output_mtok`, US dollars per 1,000,000 tokens. Source of truth: [`prices/providers/huggingface_nscale.yml`](https://github.com/mahimailabs/voice-prices/blob/main/prices/providers/huggingface_nscale.yml).

| Model                                       | Name                           | Input \$ / Mtok | Output \$ / Mtok | Context |
| ------------------------------------------- | ------------------------------ | --------------: | ---------------: | ------: |
| `Qwen/QwQ-32B`                              | QwQ-32B                        |          \$0.18 |            \$0.2 | 131,072 |
| `Qwen/Qwen2.5-Coder-32B-Instruct`           | Qwen2.5-Coder-32B-Instruct     |          \$0.06 |            \$0.2 | 131,072 |
| `Qwen/Qwen2.5-Coder-3B-Instruct`            | Qwen2.5-Coder-3B-Instruct      |          \$0.01 |           \$0.03 |  32,768 |
| `Qwen/Qwen2.5-Coder-7B-Instruct`            | Qwen2.5-Coder-7B-Instruct      |          \$0.01 |           \$0.03 | 131,072 |
| `Qwen/Qwen3-14B`                            | Qwen3-14B                      |          \$0.07 |            \$0.2 |  40,960 |
| `Qwen/Qwen3-235B-A22B`                      | Qwen3-235B-A22B                |           \$0.2 |            \$0.6 |  32,000 |
| `Qwen/Qwen3-32B`                            | Qwen3-32B                      |          \$0.08 |           \$0.25 |  40,960 |
| `Qwen/Qwen3-4B-Instruct-2507`               | Qwen3-4B-Instruct-2507         |          \$0.01 |           \$0.03 | 262,144 |
| `Qwen/Qwen3-4B-Thinking-2507`               | Qwen3-4B-Thinking-2507         |          \$0.01 |           \$0.03 | 262,144 |
| `Qwen/Qwen3-8B`                             | Qwen3-8B                       |          \$0.07 |           \$0.18 |  40,960 |
| `deepseek-ai/DeepSeek-R1-Distill-Llama-70B` | DeepSeek-R1-Distill-Llama-70B  |          \$0.75 |           \$0.75 | 131,072 |
| `deepseek-ai/DeepSeek-R1-Distill-Llama-8B`  | DeepSeek-R1-Distill-Llama-8B   |          \$0.05 |           \$0.05 | 131,072 |
| `deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B` | DeepSeek-R1-Distill-Qwen-1.5B  |           \$0.1 |            \$0.1 | 131,072 |
| `deepseek-ai/DeepSeek-R1-Distill-Qwen-32B`  | DeepSeek-R1-Distill-Qwen-32B   |           \$0.3 |            \$0.3 | 131,072 |
| `deepseek-ai/DeepSeek-R1-Distill-Qwen-7B`   | DeepSeek-R1-Distill-Qwen-7B    |          \$0.15 |           \$0.15 | 131,072 |
| `meta-llama/Llama-3.1-8B-Instruct`          | Llama-3.1-8B-Instruct          |          \$0.06 |           \$0.06 | 131,072 |
| `meta-llama/Llama-3.3-70B-Instruct`         | Llama-3.3-70B-Instruct         |           \$0.4 |            \$0.4 | 131,072 |
| `meta-llama/Llama-4-Scout-17B-16E-Instruct` | Llama-4-Scout-17B-16E-Instruct |          \$0.09 |           \$0.29 | 890,000 |
| `openai/gpt-oss-120b`                       | gpt-oss-120b                   |           \$0.1 |            \$0.4 | 131,072 |
| `openai/gpt-oss-20b`                        | gpt-oss-20b                    |          \$0.05 |            \$0.2 | 131,072 |
