Pricing

Compare input, cached input and output rates per million tokens. Choose a model that fits your application and follow actual costs in your dashboard.

Sketched developers review a usage ledger between streams of input and output tokens.

See what you use. Plan what’s next.

Inference catalogGLM 5.3 Flash · Available

Models and token rates.

Availability and rates can change. Choose an available model and check its current prices before making a request.

USD per 1M tokens

Current and announced rates are shown separately.

Published per-million-token launch and standard rates with availability for every catalog model.
ModelAvailabilityInput / 1MCached / 1MOutput / 1MDetails
GLM 5.3 FlashGLMAvailable$0.125$0.050$0.50Rates
DeepSeek V4.1 FlashDeepSeekComing soonNot publishedNot publishedNot publishedRates
Qwen 3.6 35B A3BQwenComing soonNot publishedNot publishedNot publishedRates
Kimi K3KimiRoadmapNot publishedNot publishedNot publishedRates
GLM 5.3GLMRoadmapNot publishedNot publishedNot publishedRates
Qwen 3.8 2.4T A95BQwenRoadmapNot publishedNot publishedNot publishedRates
DeepSeek V4 Pro 0813DeepSeekRoadmapNot publishedNot publishedNot publishedRates
Roadmap models and release plans may change. Prices are shown when published.

More room to build.

Check your Billing

EVERY MONTH

$5 of free API usage

For every verified, active personal account, including existing users. Your allowance resets each month. Unused allowance does not roll over.

YOUR FIRST TOP-UP

1:1 match, up to $250

Add $50, receive $50 in bonus API credit. Add $250 or more, receive $250. Only your first successful top-up earns the match.

Bonus credit is non-refundable and expires one year after it is granted. Refunds or reversals remove the corresponding bonus. Purchased credit never expires.

Top-ups are available when checkout is enabled in Billing. One offer per person; creating additional accounts to repeat it is not permitted. Full offer terms.

Dedicated GPU

Run a custom model of your choice with dedicated GPU inference. Configuration options and prices will be published before launch.

Coming soonDedicated inference is coming soon.
Bring your model
Custom models
Inference
Dedicated GPUs
Metering
Time-based

Join the update list for availability and pricing announcements.

No account required.
Coming soon

Adapt a base model to your dataset and task. Training and custom-model inference are coming soon.

Coming soonNo active training price
Adapters
LoRA
Model weights
Full fine-tuning (FFT)
After training
Custom-model inference
Training stackUnsloth · Axolotl · Hugging Face TRLPlanned tooling includes Unsloth, Axolotl and Hugging Face TRL.
No account required.

Join the training update list for release announcements.

Training rates and configuration options will be published before paid runs open.

Questions

Each model has separate input, cached input and output rates per million tokens. The table shows current rates and any announced standard rates. Cost depends on the model and the confirmed tokens used by your request.

Choose a model marked Available. Coming-soon and roadmap models cannot accept requests yet. Check the catalog for current availability.

Dedicated inference is coming soon, with time-based pricing. Configuration options and rates will be published before launch.

Requests with no confirmed usage are not charged. If an interrupted or failed response has confirmed token usage, those tokens may be charged at the applicable rate.

LoRA and full fine-tuning (FFT) are coming soon. Training prices and billing details will be published before paid runs open.

Payment methods, top-ups and custom spend-limit controls are not currently available in the dashboard. You can review recorded usage and cost there.

Cached tokens are a subset of input tokens. They use the cached input rate; the remaining input tokens use the input rate. Output tokens use the output rate.

Start building

Find the capabilities and token rates that fit your app, then get your first response.

Sketched developers review a usage ledger between streams of input and output tokens.