Deployments - TypeScript SDK
client.deployments previews, creates, and manages model-serving workloads.
The SDK chooses the create/preview endpoint automatically: a body with
modelAssetId uses the asset route; otherwise it uses the catalog route.
Overview
Available Operations
| Method | Description |
|---|---|
capabilities() | Check whether Deployment mutations are allowed. |
features() | Read Deployment feature flags. |
clusterOptions() | List clusters selectable for serving. |
modelInventory(clusterId) | Read catalog model placement inventory for a cluster. |
preview(body) | Run resource preflight through the same catalog/asset route as create. |
create(body) | Create a catalog, Upload, or Trained Model Deployment. |
list(options?) | Search and page Deployments. |
get(id) | Read one Deployment. |
events(id) | Read lifecycle events for troubleshooting. |
yaml(id) | Read the rendered Kubernetes YAML. |
update(id, body) | Update mutable serving configuration. |
stop(id) | Stop the running workload while retaining its record. |
previewRestart(id) | Check restart blockers and current resource fit. |
restart(id) | Recreate a stopped or failed workload. |
delete(id) | Delete the Deployment and managed workload resources. |
capabilities
Check whether Deployment mutations are allowed.
Request
This method has no parameters and sends no request body.
Response
InferenceServiceCapabilitiesVO with mutationsAllowed, unavailableReasonCode.
features
Read Deployment feature flags.
Request
This method has no parameters and sends no request body.
Response
DeploymentFeaturesVO with clusterSelectionEnabled.
clusterOptions
List clusters selectable for serving.
Request
This method has no parameters and sends no request body.
Response
WorkloadClusterOptionVO[].
modelInventory
Read Model Profile placement inventory for a cluster.
Request
clusterId required.
Response
ModelInventoryVO.
preview
Run resource preflight through the same Profile/Asset route as create.
Request
An object with the common create fields and exactly
one model reference: modelAssetId, or modelSource plus modelName.
Response
ModelServingResourcePreviewVO; see response fields.
create
Create a Model Profile, Upload, or Trained Model Deployment.
Request
An object with the common create fields and exactly
one model reference: modelAssetId, or modelSource plus modelName.
Response
InferenceServiceVO; see response fields.
list
Search and page Deployments.
Request
Optional page, pageSize, keyword, status.
Response
PageResult<InferenceServiceVO>.
get
Read one Deployment.
Request
id required.
Response
InferenceServiceVO.
events
Read lifecycle events for troubleshooting.
Request
id required.
Response
DeploymentEventVO[].
yaml
Read the rendered Kubernetes YAML.
Request
id required.
Response
Raw string, not an API envelope.
update
Update mutable serving configuration.
Request
id and update body required.
Response
Updated InferenceServiceVO.
stop
Stop the running workload while retaining its record.
Request
id required.
Response
Updated InferenceServiceVO.
previewRestart
Check restart blockers and current resource fit.
Request
id required.
Response
DeploymentRestartPreviewVO.
restart
Recreate a stopped or failed workload.
Request
id required.
Response
Updated InferenceServiceVO.
delete
Delete the Deployment and managed workload resources.
Request
id required.
Response
null.
Field Reference and Examples
Create Request Fields
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Yes | Unique Deployment name. |
clusterId | number | Environment-dependent | Required when cluster selection is enabled; use clusterOptions(). |
backend | string | Yes | Serving backend, normally sglang or vllm. |
servingMode | string | No | Usually standard or disaggregated; use model capabilities. |
gpuType | string | Yes | GPU type selected from capabilities. |
replicas | number | No | Standard-mode replica count. |
tensorParallel | number | No | Standard-mode tensor parallel size. |
prefillReplicas / decodeReplicas | number | No | Disaggregated replica counts. |
prefillTensorParallel / decodeTensorParallel | number | No | Disaggregated tensor parallel sizes. |
acceleratorSelection | object | No | resourceName, productLabelKey, productLabelValues. |
deploymentProfileKey | string | No | Explicit profile selected from model capabilities. |
kvCacheDistributed, mtpEnabled, kvFP8Enabled, operatorOptimizationEnabled | boolean | No | Optional features supported by the profile. |
kvCacheMaxCapacityGB, rateLimit | number | No | Optional KV-cache capacity and QPS limit. |
Catalog creation additionally requires modelSource and modelName, and may
include modelRevision. Asset creation requires modelAssetId instead of
modelSource and modelName. Machine-to-machine callers may have additional
source fields; normal SDK users should not put storage credentials in a
Deployment request.
Update behavior: Runtime V2 currently permits only description to change
in place. Changes to rateLimit, GPU, serving mode, replicas, parallelism,
KV-cache settings, or other runtime features return
DEPLOYMENT_V2_NEW_GENERATION_REQUIRED; create a new Deployment generation
for those changes.
Response Fields
| Type | Fields |
|---|---|
ModelServingResourcePreviewVO | clusterId, clusterName, creatable, blockerCodes, blockers, warnings, workloadPlan, preflightV2, acceleratorPool, runtimeBackend |
InferenceServiceVO | id, name, status, statusMessage, modelName, modelSource, backend, servingMode, clusterId, clusterName, endpoint, readyReplicas, totalReplicas, capabilities |
DeploymentRestartPreviewVO | deploymentId, restartable, blockerCode, blockerMessage, resourcePreview |
DeploymentEventVO | id, serviceId, eventType, message, details, createdAt |
const profileRequest = {
name: "example-profile-deployment",
clusterId: 1,
modelSource: "gallery",
modelName: "Qwen3-4B-Instruct-2507-FAST",
backend: "sglang",
servingMode: "standard",
gpuType: "L20",
replicas: 1,
};
const assetRequest = {
name: "example-asset-deployment",
clusterId: 1,
modelAssetId: 42,
backend: "sglang",
servingMode: "standard",
gpuType: "L20",
replicas: 1,
};
const features = await client.deployments.features();
const mutationState = await client.deployments.capabilities();
const clusters = await client.deployments.clusterOptions();
const inventory = await client.deployments.modelInventory(1);
const preview = await client.deployments.preview<{ creatable: boolean }>(profileRequest);
if (preview.creatable) await client.deployments.create(profileRequest);
const assetPreview = await client.deployments.preview(assetRequest);
const created = await client.deployments.create<{ id: number }>(assetRequest);
const page = await client.deployments.list({ page: 1, pageSize: 20, status: "RUNNING" });
const detail = await client.deployments.get(created.id);
const events = await client.deployments.events(created.id);
const renderedYaml = await client.deployments.yaml(created.id);
const updated = await client.deployments.update(created.id, {
description: "validated",
});
const stopped = await client.deployments.stop(created.id);
const restartCheck = await client.deployments.previewRestart<{ restartable: boolean }>(created.id);
if (restartCheck.restartable) await client.deployments.restart(created.id);
await client.deployments.delete(created.id);
All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.