Models
Model specifications for all Tinker-supported models.
| Model | Org | Arch | Total | Active | Layers | d_model | Attention | Experts |
|---|---|---|---|---|---|---|---|---|
| Qwen3.5-397B-A17B | qwen | moe | 397B | 17B | — | — | GQA | — |
| Qwen3.5-35B-A3B | qwen | moe | 35B | 3B | — | — | GQA | — |
| Qwen3.5-27B | qwen | dense | 27B | 27B | — | — | GQA | — |
| Qwen3.5-4B | qwen | dense | 4B | 4B | — | — | GQA | — |
| Qwen3-VL-235B-A22B-Instruct | qwen | moe | 235B | 22B | 94 | 4096 | GQA | 128/8 |
| Qwen3-VL-30B-A3B-Instruct | qwen | moe | 30.5B | 3.3B | 48 | 2048 | GQA | 128/8 |
| Qwen3-235B-A22B-Instruct-2507 | qwen | moe | 235B | 22B | 94 | 4096 | GQA | 128/8 |
| Qwen3-30B-A3B-Instruct-2507 | qwen | moe | 30.5B | 3.3B | 48 | 2048 | GQA | 128/8 |
| Qwen3-30B-A3B | qwen | moe | 30.5B | 3.3B | 48 | 2048 | GQA | 128/8 |
| Qwen3-30B-A3B-Base | qwen | moe | 30.5B | 3.3B | 48 | 2048 | GQA | 128/8 |
| Qwen3-32B | qwen | dense | 32B | 32B | 64 | 5120 | GQA | — |
| Qwen3-8B | qwen | dense | 8B | 8B | 36 | 4096 | GQA | — |
| Qwen3-8B-Base | qwen | dense | 8B | 8B | 36 | 4096 | GQA | — |
| Qwen3-4B-Instruct-2507 | qwen | dense | 4B | 4B | 36 | 2560 | GQA | — |
| GPT-OSS-120B | openai | moe | 117B | 5.1B | — | — | GQA | — |
| GPT-OSS-20B | openai | moe | 21B | 3.6B | — | — | GQA | — |
| DeepSeek-V3.1 | deepseek | moe | 671B | 37B | 61 | 7168 | MLA | 256/8 |
| DeepSeek-V3.1-Base | deepseek | moe | 671B | 37B | 61 | 7168 | MLA | 256/8 |
| Llama-3.1-70B | meta | dense | 70B | 70B | 80 | 8192 | GQA | — |
| Llama-3.3-70B-Instruct | meta | dense | 70B | 70B | 80 | 8192 | GQA | — |
| Llama-3.1-8B | meta | dense | 8B | 8B | 32 | 4096 | GQA | — |
| Llama-3.1-8B-Instruct | meta | dense | 8B | 8B | 32 | 4096 | GQA | — |
| Llama-3.2-3B | meta | dense | 3B | 3B | 28 | 3072 | GQA | — |
| Llama-3.2-1B | meta | dense | 1B | 1B | 16 | 2048 | GQA | — |
| Kimi-K2-Thinking | moonshot | moe | 1000B | 32B | — | — | MLA | — |
| Kimi-K2.5 | moonshot | moe | 1000B | 32B | — | — | MLA | — |