Compare input, cached input and output rates per million tokens. Choose a model that fits your application and follow actual costs in your dashboard.

See what you use. Plan what’s next.
Models and token rates.
Availability and rates can change. Choose an available model and check its current prices before making a request.
Current and announced rates are shown separately.
| Model | Availability | Input / 1M | Cached / 1M | Output / 1M | Details |
|---|---|---|---|---|---|
| GLM 5.3 FlashGLM | Available | $0.125 | $0.050 | $0.50 | Rates |
| DeepSeek V4.1 FlashDeepSeek | Coming soon | Not published | Not published | Not published | Rates |
| Qwen 3.6 35B A3BQwen | Coming soon | Not published | Not published | Not published | Rates |
| Kimi K3Kimi | Roadmap | Not published | Not published | Not published | Rates |
| GLM 5.3GLM | Roadmap | Not published | Not published | Not published | Rates |
| Qwen 3.8 2.4T A95BQwen | Roadmap | Not published | Not published | Not published | Rates |
| DeepSeek V4 Pro 0813DeepSeek | Roadmap | Not published | Not published | Not published | Rates |
More room to build.
Check your BillingEVERY MONTH
$5 of free API usage
For every verified, active personal account, including existing users. Your allowance resets each month. Unused allowance does not roll over.
YOUR FIRST TOP-UP
1:1 match, up to $250
Add $50, receive $50 in bonus API credit. Add $250 or more, receive $250. Only your first successful top-up earns the match.
Bonus credit is non-refundable and expires one year after it is granted. Refunds or reversals remove the corresponding bonus. Purchased credit never expires.
Top-ups are available when checkout is enabled in Billing. One offer per person; creating additional accounts to repeat it is not permitted. Full offer terms.
Run a custom model of your choice with dedicated GPU inference. Configuration options and prices will be published before launch.
- Bring your model
- Custom models
- Inference
- Dedicated GPUs
- Metering
- Time-based
Join the update list for availability and pricing announcements.
Adapt a base model to your dataset and task. Training and custom-model inference are coming soon.
- Adapters
- LoRA
- Model weights
- Full fine-tuning (FFT)
- After training
- Custom-model inference
Join the training update list for release announcements.
Training rates and configuration options will be published before paid runs open.
Each model has separate input, cached input and output rates per million tokens. The table shows current rates and any announced standard rates. Cost depends on the model and the confirmed tokens used by your request.
Choose a model marked Available. Coming-soon and roadmap models cannot accept requests yet. Check the catalog for current availability.
Dedicated inference is coming soon, with time-based pricing. Configuration options and rates will be published before launch.
Requests with no confirmed usage are not charged. If an interrupted or failed response has confirmed token usage, those tokens may be charged at the applicable rate.
LoRA and full fine-tuning (FFT) are coming soon. Training prices and billing details will be published before paid runs open.
Payment methods, top-ups and custom spend-limit controls are not currently available in the dashboard. You can review recorded usage and cost there.
Cached tokens are a subset of input tokens. They use the cached input rate; the remaining input tokens use the input rate. Output tokens use the output rate.
Find the capabilities and token rates that fit your app, then get your first response.

