Documentation
Coming soon

Dedicated GPU inference

Run a custom model of your choice on dedicated GPU capacity.

Dedicated GPU inference

Coming soon

Dedicated GPU inference is part of Arnict’s planned product offering. Access is not open yet.

Dedicated inference is for running a custom model of your choice with capacity assigned to your workload. It is a separate option from calling a model in the shared catalog.

The intended workflow is to choose your model, configure its deployment and call it from your application. You can bring a model you already use or a checkpoint you have fine-tuned.

  • Choose your own custom model.
  • Use dedicated capacity for inference.
  • Keep keys, usage and deployment access within your Arnict account.

Pricing and launch details

Dedicated inference is planned to use time-based pricing. Final configuration options, compatibility requirements and rates will be published before access opens. There is no active dedicated-inference charge today.

Prepare for a deployment

Know your model’s source, license, required input types and expected request volume. These help you select a suitable configuration when dedicated inference becomes available.