Deployments - Python SDK
client.deployments previews, creates, and manages model-serving workloads.
The SDK chooses the correct create/preview endpoint automatically: a body with
modelAssetId uses the asset route; otherwise it uses the Model Profile route.
Overview
Available Operations
| Method | Description |
|---|---|
capabilities() | Check whether Deployment mutations are currently allowed. |
features() | Read Deployment feature flags. |
cluster_options() | List clusters selectable for serving. |
model_inventory(cluster_id) | Read Model Profile placement inventory for a cluster. |
preview(body) | Run resource preflight using the same Profile/Asset routing as create. |
create(body) | Create a Profile, Upload, or Trained Model Deployment. |
list(...) | Search and page Deployments. |
get(id) | Read one Deployment. |
events(id) | Read lifecycle events for troubleshooting. |
yaml(id) | Read the rendered Kubernetes YAML. |
update(id, body) | Update mutable serving configuration. |
stop(id) | Stop the running workload while retaining its record. |
preview_restart(id) | Check restart blockers and current resource fit. |
restart(id) | Recreate a stopped or failed workload. |
delete(id) | Delete the Deployment and its managed workload resources. |
capabilities
Check whether Deployment mutations are currently allowed.
Request
This method has no parameters and sends no request body.
Response
InferenceServiceCapabilitiesVO with mutationsAllowed, unavailableReasonCode.
features
Read Deployment feature flags.
Request
This method has no parameters and sends no request body.
Response
DeploymentFeaturesVO with clusterSelectionEnabled.
cluster_options
List clusters selectable for serving.
Request
This method has no parameters and sends no request body.
Response
list[WorkloadClusterOptionVO].
model_inventory
Read Model Profile placement inventory for a cluster.
Request
cluster_id required.
Response
ModelInventoryVO.
preview
Run resource preflight using the same Profile/Asset routing as create.
Request
A mapping with the common create fields and exactly
one model reference: modelAssetId, or modelSource plus modelName.
Response
ModelServingResourcePreviewVO; see response fields.
create
Create a Profile, Upload, or Trained Model Deployment.
Request
A mapping with the common create fields and exactly
one model reference: modelAssetId, or modelSource plus modelName.
Response
InferenceServiceVO; see response fields.
list
Search and page Deployments.
Request
Optional page, page_size, keyword, status.
Response
PageResult[InferenceServiceVO].
get
Read one Deployment.
Request
id required.
Response
InferenceServiceVO.
events
Read lifecycle events for troubleshooting.
Request
id required.
Response
list[DeploymentEventVO].
yaml
Read the rendered Kubernetes YAML.
Request
id required.
Response
Raw str, not an API envelope.
update
Update mutable serving configuration.
Request
id and UpdateServiceRequest required.
Response
Updated InferenceServiceVO.
stop
Stop the running workload while retaining its record.
Request
id required.
Response
Updated InferenceServiceVO.
preview_restart
Check restart blockers and current resource fit.
Request
id required.
Response
DeploymentRestartPreviewVO.
restart
Recreate a stopped or failed workload.
Request
id required.
Response
Updated InferenceServiceVO.
delete
Delete the Deployment and its managed workload resources.
Request
id required.
Response
None.
Field Reference and Examples
Create Request Fields
| Field | Type | Required | Description |
|---|---|---|---|
name | str | Yes | Unique Deployment name. |
clusterId | int | Environment-dependent | Selected cluster. Use cluster_options(); required when cluster selection is enabled. |
backend | str | Yes | Serving backend, normally sglang or vllm. |
servingMode | str | No | Usually standard or disaggregated; supported values come from model capabilities. |
gpuType | str | Yes | Selected GPU type from capabilities. |
replicas | int | No | Standard-mode replica count. |
tensorParallel | int | No | Standard-mode tensor parallel size. |
prefillReplicas / decodeReplicas | int | No | Disaggregated replica counts. |
prefillTensorParallel / decodeTensorParallel | int | No | Disaggregated tensor parallel sizes. |
acceleratorSelection | dict | No | resourceName, productLabelKey, productLabelValues. |
deploymentProfileKey | str | No | Explicit deployment profile selected from model capabilities. |
kvCacheDistributed, mtpEnabled, kvFP8Enabled, operatorOptimizationEnabled | bool | No | Optional runtime features supported by the model profile. |
kvCacheMaxCapacityGB, rateLimit | int | No | Optional KV-cache capacity and QPS limit. |
Profile creation additionally requires modelSource and modelName; it may
include modelRevision. The machine-to-machine path may also supply
modelOssPath and ossCredential, but normal SDK users should not place cloud
credentials in Deployment requests. Asset creation requires modelAssetId
instead of modelSource/modelName.
Update behavior: Runtime V2 currently permits only description to change
in place. Changes to rateLimit, GPU, serving mode, replicas, parallelism,
KV-cache settings, or other runtime features return
DEPLOYMENT_V2_NEW_GENERATION_REQUIRED; create a new Deployment generation
for those changes.
Response Fields
| Type | Fields |
|---|---|
ModelServingResourcePreviewVO | clusterId, clusterName, creatable, blockerCodes, blockers, warnings, workloadPlan, preflightV2, acceleratorPool, runtimeBackend |
InferenceServiceVO | id, name, status, statusMessage, modelName, modelSource, backend, servingMode, clusterId, clusterName, endpoint, readyReplicas, totalReplicas, capabilities |
DeploymentRestartPreviewVO | deploymentId, restartable, blockerCode, blockerMessage, resourcePreview |
DeploymentEventVO | id, serviceId, eventType, message, details, createdAt |
profile_request = {
"name": "example-profile-deployment",
"clusterId": 1,
"modelSource": "gallery",
"modelName": "Qwen3-4B-Instruct-2507-FAST",
"backend": "sglang",
"servingMode": "standard",
"gpuType": "L20",
"replicas": 1,
}
asset_request = {
"name": "example-asset-deployment",
"clusterId": 1,
"modelAssetId": 42,
"backend": "sglang",
"servingMode": "standard",
"gpuType": "L20",
"replicas": 1,
}
features = client.deployments.features()
mutation_state = client.deployments.capabilities()
clusters = client.deployments.cluster_options()
inventory = client.deployments.model_inventory(cluster_id=1)
preview = client.deployments.preview(profile_request)
if preview["creatable"]:
created = client.deployments.create(profile_request)
asset_preview = client.deployments.preview(asset_request)
asset_deployment = client.deployments.create(asset_request)
page = client.deployments.list(page=1, page_size=20, status="RUNNING")
detail = client.deployments.get(asset_deployment["id"])
events = client.deployments.events(asset_deployment["id"])
rendered_yaml = client.deployments.yaml(asset_deployment["id"])
updated = client.deployments.update(
asset_deployment["id"],
{"description": "validated"},
)
stopped = client.deployments.stop(asset_deployment["id"])
restart_check = client.deployments.preview_restart(asset_deployment["id"])
if restart_check["restartable"]:
restarted = client.deployments.restart(asset_deployment["id"])
client.deployments.delete(asset_deployment["id"])
All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.