Skip to main content

Deployments - Python SDK

client.deployments previews, creates, and manages model-serving workloads. The SDK chooses the correct create/preview endpoint automatically: a body with modelAssetId uses the asset route; otherwise it uses the Model Profile route.

Overview​

Available Operations​

MethodDescription
capabilities()Check whether Deployment mutations are currently allowed.
features()Read Deployment feature flags.
cluster_options()List clusters selectable for serving.
model_inventory(cluster_id)Read Model Profile placement inventory for a cluster.
preview(body)Run resource preflight using the same Profile/Asset routing as create.
create(body)Create a Profile, Upload, or Trained Model Deployment.
list(...)Search and page Deployments.
get(id)Read one Deployment.
events(id)Read lifecycle events for troubleshooting.
yaml(id)Read the rendered Kubernetes YAML.
update(id, body)Update mutable serving configuration.
stop(id)Stop the running workload while retaining its record.
preview_restart(id)Check restart blockers and current resource fit.
restart(id)Recreate a stopped or failed workload.
delete(id)Delete the Deployment and its managed workload resources.

capabilities​

Check whether Deployment mutations are currently allowed.

Request​

This method has no parameters and sends no request body.

Response​

InferenceServiceCapabilitiesVO with mutationsAllowed, unavailableReasonCode.

features​

Read Deployment feature flags.

Request​

This method has no parameters and sends no request body.

Response​

DeploymentFeaturesVO with clusterSelectionEnabled.

cluster_options​

List clusters selectable for serving.

Request​

This method has no parameters and sends no request body.

Response​

list[WorkloadClusterOptionVO].

model_inventory​

Read Model Profile placement inventory for a cluster.

Request​

cluster_id required.

Response​

ModelInventoryVO.

preview​

Run resource preflight using the same Profile/Asset routing as create.

Request​

A mapping with the common create fields and exactly one model reference: modelAssetId, or modelSource plus modelName.

Response​

ModelServingResourcePreviewVO; see response fields.

create​

Create a Profile, Upload, or Trained Model Deployment.

Request​

A mapping with the common create fields and exactly one model reference: modelAssetId, or modelSource plus modelName.

Response​

InferenceServiceVO; see response fields.

list​

Search and page Deployments.

Request​

Optional page, page_size, keyword, status.

Response​

PageResult[InferenceServiceVO].

get​

Read one Deployment.

Request​

id required.

Response​

InferenceServiceVO.

events​

Read lifecycle events for troubleshooting.

Request​

id required.

Response​

list[DeploymentEventVO].

yaml​

Read the rendered Kubernetes YAML.

Request​

id required.

Response​

Raw str, not an API envelope.

update​

Update mutable serving configuration.

Request​

id and UpdateServiceRequest required.

Response​

Updated InferenceServiceVO.

stop​

Stop the running workload while retaining its record.

Request​

id required.

Response​

Updated InferenceServiceVO.

preview_restart​

Check restart blockers and current resource fit.

Request​

id required.

Response​

DeploymentRestartPreviewVO.

restart​

Recreate a stopped or failed workload.

Request​

id required.

Response​

Updated InferenceServiceVO.

delete​

Delete the Deployment and its managed workload resources.

Request​

id required.

Response​

None.

Field Reference and Examples​

Create Request Fields​

FieldTypeRequiredDescription
namestrYesUnique Deployment name.
clusterIdintEnvironment-dependentSelected cluster. Use cluster_options(); required when cluster selection is enabled.
backendstrYesServing backend, normally sglang or vllm.
servingModestrNoUsually standard or disaggregated; supported values come from model capabilities.
gpuTypestrYesSelected GPU type from capabilities.
replicasintNoStandard-mode replica count.
tensorParallelintNoStandard-mode tensor parallel size.
prefillReplicas / decodeReplicasintNoDisaggregated replica counts.
prefillTensorParallel / decodeTensorParallelintNoDisaggregated tensor parallel sizes.
acceleratorSelectiondictNoresourceName, productLabelKey, productLabelValues.
deploymentProfileKeystrNoExplicit deployment profile selected from model capabilities.
kvCacheDistributed, mtpEnabled, kvFP8Enabled, operatorOptimizationEnabledboolNoOptional runtime features supported by the model profile.
kvCacheMaxCapacityGB, rateLimitintNoOptional KV-cache capacity and QPS limit.

Profile creation additionally requires modelSource and modelName; it may include modelRevision. The machine-to-machine path may also supply modelOssPath and ossCredential, but normal SDK users should not place cloud credentials in Deployment requests. Asset creation requires modelAssetId instead of modelSource/modelName.

Update behavior: Runtime V2 currently permits only description to change in place. Changes to rateLimit, GPU, serving mode, replicas, parallelism, KV-cache settings, or other runtime features return DEPLOYMENT_V2_NEW_GENERATION_REQUIRED; create a new Deployment generation for those changes.

Response Fields​

TypeFields
ModelServingResourcePreviewVOclusterId, clusterName, creatable, blockerCodes, blockers, warnings, workloadPlan, preflightV2, acceleratorPool, runtimeBackend
InferenceServiceVOid, name, status, statusMessage, modelName, modelSource, backend, servingMode, clusterId, clusterName, endpoint, readyReplicas, totalReplicas, capabilities
DeploymentRestartPreviewVOdeploymentId, restartable, blockerCode, blockerMessage, resourcePreview
DeploymentEventVOid, serviceId, eventType, message, details, createdAt
profile_request = {
"name": "example-profile-deployment",
"clusterId": 1,
"modelSource": "gallery",
"modelName": "Qwen3-4B-Instruct-2507-FAST",
"backend": "sglang",
"servingMode": "standard",
"gpuType": "L20",
"replicas": 1,
}
asset_request = {
"name": "example-asset-deployment",
"clusterId": 1,
"modelAssetId": 42,
"backend": "sglang",
"servingMode": "standard",
"gpuType": "L20",
"replicas": 1,
}

features = client.deployments.features()
mutation_state = client.deployments.capabilities()
clusters = client.deployments.cluster_options()
inventory = client.deployments.model_inventory(cluster_id=1)

preview = client.deployments.preview(profile_request)
if preview["creatable"]:
created = client.deployments.create(profile_request)

asset_preview = client.deployments.preview(asset_request)
asset_deployment = client.deployments.create(asset_request)

page = client.deployments.list(page=1, page_size=20, status="RUNNING")
detail = client.deployments.get(asset_deployment["id"])
events = client.deployments.events(asset_deployment["id"])
rendered_yaml = client.deployments.yaml(asset_deployment["id"])
updated = client.deployments.update(
asset_deployment["id"],
{"description": "validated"},
)
stopped = client.deployments.stop(asset_deployment["id"])
restart_check = client.deployments.preview_restart(asset_deployment["id"])
if restart_check["restartable"]:
restarted = client.deployments.restart(asset_deployment["id"])
client.deployments.delete(asset_deployment["id"])

All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.