Supported models and pricing
Compare model families, context windows, capabilities, traffic share, and input/output pricing in one live catalog.
- Pricing syncs automatically
- Seven model families in one catalog
- One API key across every provider
Transparent provider pricing.
Tokenly price = upstream list price × the provider service rate. Rates are shown before you make a request.
Latest models
New providers appear here first, then move into their family directory below.
How to read this catalog. Traffic share is usage analytics, not a whitelist or quota. Every Tokenly account can call every model listed here. Cached-input prices are shown only when available upstream.
Claude models
3 active models · all plans can call these models
| Model | Input / M | Output / M | Context | Capabilities | Traffic |
|---|---|---|---|---|---|
Claude Opus 5claude-opus-5 | $5.00cached · $5.00 | $25.00 | 1M / 128K | VisionToolsCache | 4.4% |
Claude Sonnet 5claude-sonnet-5 | $2.00cached · $2.00 | $10.00 | 1M / 128K | VisionToolsCache | 16.8% |
Claude Haiku 4.5claude-haiku-4-5 | $1.00cached · $1.00 | $5.00 | 200K / 64K | VisionToolsCache | 12.0% |
OpenAI / Codex models
3 active models · all plans can call these models
| Model | Input / M | Output / M | Context | Capabilities | Traffic |
|---|---|---|---|---|---|
GPT 5.6 solgpt-5.6-sol | $5.00cached · $5.00 | $30.00 | 922K / 128K | VisionToolsCache | 8.7% |
GPT 5.6 terragpt-5.6-terra | $2.00cached · $2.00 | $12.00 | 922K / 128K | VisionToolsCache | 6.4% |
GPT 6 lunagpt-6-luna | $0.10cached · $0.10 | $0.50 | 922K / 128K | VisionTools | <0.1% |
Google Gemini models
2 active models · all plans can call these models
| Model | Input / M | Output / M | Context | Capabilities | Traffic |
|---|---|---|---|---|---|
Gemini 3.5 Flashgemini-3.5-flash | $3.00cached · $3.00 | $18.00 | 1M / 65K | VisionToolsCache | — |
Gemini 3.5 Progemini-3.5-pro | $5.00cached · $5.00 | $20.00 | 1M / 65K | VisionToolsCache | <0.1% |
GLM models
2 active models · all plans can call these models
| Model | Input / M | Output / M | Context | Capabilities | Traffic |
|---|---|---|---|---|---|
GLM 5.2glm-5.2 | $2.38cached · $2.38 | $8.32 | 200K / 131K | VisionToolsCache | 0.7% |
GLM 5.3 Flashglm-5.3-flash | $0.14cached · $0.14 | $0.49 | 1M / 128K | VisionToolsCache | 2.1% |
Kimi models
1 active model · all plans can call these models
| Model | Input / M | Output / M | Context | Capabilities | Traffic |
|---|---|---|---|---|---|
Kimi K3kimi-k3 | $21.00cached · $21.00 | $105.00 | 1M / 131K | VisionToolsCache | 0.1% |
DeepSeek models
2 active models · all plans can call these models
| Model | Input / M | Output / M | Context | Capabilities | Traffic |
|---|---|---|---|---|---|
DeepSeek V4 Prodeepseek-v4-pro | $0.66cached · $0.66 | $1.98 | 1M / 384K | ReasoningTools | 1.7% |
DeepSeek V4.1 Flashdeepseek-v4.1-flash | $0.57cached · $0.57 | $1.70 | 1M / 384K | ReasoningFast | 11.3% |
Qwen models
2 active models · all plans can call these models
| Model | Input / M | Output / M | Context | Capabilities | Traffic |
|---|---|---|---|---|---|
Qwen 3.8 Flashqwen3.8-flash | $0.35cached · $0.35 | $1.05 | 1M / 128K | VisionToolsCache | 0.5% |
Qwen 3.8 Maxqwen3.8-max | $4.20cached · $4.20 | $12.59 | 1M / 128K | VisionToolsCache | 0.2% |
Use any model with one Tokenly key.
Create a token, choose a provider, and start with the same OpenAI-compatible endpoint.
Common questions.
How are model prices calculated?
We display the upstream list price and apply the provider service rate shown above. Your usage ledger records input, output, and cached input separately.
Can every plan call every model?
Yes. Plans control credits, quotas, and support. They do not hide models from the catalog.
How often does pricing update?
Pricing is synchronized regularly and the catalog shows the latest available rate for each provider.