Dedicated GPU inference
Run a custom model of your choice on dedicated GPU capacity.
Dedicated GPU inference
Coming soon
Dedicated inference is for running a custom model of your choice with capacity assigned to your workload. It is a separate option from calling a model in the shared catalog.
The intended workflow is to choose your model, configure its deployment and call it from your application. You can bring a model you already use or a checkpoint you have fine-tuned.
- Choose your own custom model.
- Use dedicated capacity for inference.
- Keep keys, usage and deployment access within your Arnict account.
- Follow dedicated inference updatesJoin the waitlist and see the current product plan.
Pricing and launch details
Dedicated inference is planned to use time-based pricing. Final configuration options, compatibility requirements and rates will be published before access opens. There is no active dedicated-inference charge today.
Prepare for a deployment
