Key Benefits
Dedicated capacity
Dedicated capacity
Your deployments run on capacity reserved for your organization.
Predictable performance
Predictable performance
Reserved capacity provides predictable latency and throughput under load.
Custom weights
Custom weights
Deploy fine-tuned weights alongside standard model versions.
Configurable performance
Configurable performance
Tune capacity, draft models, model configuration, and quantization for your workload.
Advanced features
Advanced features
Use every Shared Inference capability, plus fine-tuning, weight management, and service tier controls.
Supported Models
Dedicated Inference supports many model families, parameter sizes, and variants, including-instruct and -thinking. You can also deploy custom weights and tune each deployment for your performance goals.
Alibaba Qwen: Qwen 3.8, Qwen3, and Qwen3-Coder
Alibaba Qwen: Qwen 3.8, Qwen3, and Qwen3-Coder
Qwen 3.8 27BQwen3-235B-A22BQwen3-32BQwen3-30B-A3BSmall and tiny variantsQwen3-Coder
OpenAI: GPT-OSS
OpenAI: GPT-OSS
MiniMax: M2.X
MiniMax: M2.X
Google: Gemma 4
Google: Gemma 4
Meta: Llama 3 and Llama 4
Meta: Llama 3 and Llama 4
Mistral: Mistral Small, Mistral Large 3, Devstral 2, and Mixtral
Mistral: Mistral Small, Mistral Large 3, Devstral 2, and Mixtral
Z.AI: GLM 4.X and GLM 5.X
Z.AI: GLM 4.X and GLM 5.X
Moonshot AI: Kimi K2.X
Moonshot AI: Kimi K2.X
StepFun: Step 3.X Flash
StepFun: Step 3.X Flash
ByteDance: Seed OSS
ByteDance: Seed OSS
ServiceNow: Apriel
ServiceNow: Apriel
Multimodal models, including Gemma 4, accept text and image inputs through Dedicated Inference. See Image Inputs.
Features
Dedicated Inference includes all Shared Inference capabilities, plus:- Fine-tuning: Deploy custom model weights behind your endpoint.
- Management API: Manage model versions, deployments, capacity, and endpoints through the API.
- Batch API: Run large asynchronous workloads on reserved capacity.
- Predicted Outputs: Reduce latency by supplying expected output content.
- Service tiers: Prioritize requests to meet your SLA requirements.
- Metrics: Monitor endpoint requests, tokens, latency, and health with Prometheus-compatible metrics.
Resource Model
Use these terms consistently when working with Dedicated Inference:
An inference request uses these resources in order: API route → endpoint → deployment → replica.
For compatibility with OpenAI clients, the inference API uses the
model field to select an inference target. For Shared Inference, use a model ID. For Dedicated Inference, use an endpoint ID.
