Skip to main content

Deployments - TypeScript SDK

client.deployments previews, creates, and manages model-serving workloads. The SDK chooses the create/preview endpoint automatically: a body with modelAssetId uses the asset route; otherwise it uses the catalog route.

Overview​

Available Operations​

MethodDescription
capabilities()Check whether Deployment mutations are allowed.
features()Read Deployment feature flags.
clusterOptions()List clusters selectable for serving.
modelInventory(clusterId)Read catalog model placement inventory for a cluster.
preview(body)Run resource preflight through the same catalog/asset route as create.
create(body)Create a catalog, Upload, or Trained Model Deployment.
list(options?)Search and page Deployments.
get(id)Read one Deployment.
events(id)Read lifecycle events for troubleshooting.
yaml(id)Read the rendered Kubernetes YAML.
update(id, body)Update mutable serving configuration.
stop(id)Stop the running workload while retaining its record.
previewRestart(id)Check restart blockers and current resource fit.
restart(id)Recreate a stopped or failed workload.
delete(id)Delete the Deployment and managed workload resources.

capabilities​

Check whether Deployment mutations are allowed.

Request​

This method has no parameters and sends no request body.

Response​

InferenceServiceCapabilitiesVO with mutationsAllowed, unavailableReasonCode.

features​

Read Deployment feature flags.

Request​

This method has no parameters and sends no request body.

Response​

DeploymentFeaturesVO with clusterSelectionEnabled.

clusterOptions​

List clusters selectable for serving.

Request​

This method has no parameters and sends no request body.

Response​

WorkloadClusterOptionVO[].

modelInventory​

Read Model Profile placement inventory for a cluster.

Request​

clusterId required.

Response​

ModelInventoryVO.

preview​

Run resource preflight through the same Profile/Asset route as create.

Request​

An object with the common create fields and exactly one model reference: modelAssetId, or modelSource plus modelName.

Response​

ModelServingResourcePreviewVO; see response fields.

create​

Create a Model Profile, Upload, or Trained Model Deployment.

Request​

An object with the common create fields and exactly one model reference: modelAssetId, or modelSource plus modelName.

Response​

InferenceServiceVO; see response fields.

list​

Search and page Deployments.

Request​

Optional page, pageSize, keyword, status.

Response​

PageResult<InferenceServiceVO>.

get​

Read one Deployment.

Request​

id required.

Response​

InferenceServiceVO.

events​

Read lifecycle events for troubleshooting.

Request​

id required.

Response​

DeploymentEventVO[].

yaml​

Read the rendered Kubernetes YAML.

Request​

id required.

Response​

Raw string, not an API envelope.

update​

Update mutable serving configuration.

Request​

id and update body required.

Response​

Updated InferenceServiceVO.

stop​

Stop the running workload while retaining its record.

Request​

id required.

Response​

Updated InferenceServiceVO.

previewRestart​

Check restart blockers and current resource fit.

Request​

id required.

Response​

DeploymentRestartPreviewVO.

restart​

Recreate a stopped or failed workload.

Request​

id required.

Response​

Updated InferenceServiceVO.

delete​

Delete the Deployment and managed workload resources.

Request​

id required.

Response​

null.

Field Reference and Examples​

Create Request Fields​

FieldTypeRequiredDescription
namestringYesUnique Deployment name.
clusterIdnumberEnvironment-dependentRequired when cluster selection is enabled; use clusterOptions().
backendstringYesServing backend, normally sglang or vllm.
servingModestringNoUsually standard or disaggregated; use model capabilities.
gpuTypestringYesGPU type selected from capabilities.
replicasnumberNoStandard-mode replica count.
tensorParallelnumberNoStandard-mode tensor parallel size.
prefillReplicas / decodeReplicasnumberNoDisaggregated replica counts.
prefillTensorParallel / decodeTensorParallelnumberNoDisaggregated tensor parallel sizes.
acceleratorSelectionobjectNoresourceName, productLabelKey, productLabelValues.
deploymentProfileKeystringNoExplicit profile selected from model capabilities.
kvCacheDistributed, mtpEnabled, kvFP8Enabled, operatorOptimizationEnabledbooleanNoOptional features supported by the profile.
kvCacheMaxCapacityGB, rateLimitnumberNoOptional KV-cache capacity and QPS limit.

Catalog creation additionally requires modelSource and modelName, and may include modelRevision. Asset creation requires modelAssetId instead of modelSource and modelName. Machine-to-machine callers may have additional source fields; normal SDK users should not put storage credentials in a Deployment request.

Update behavior: Runtime V2 currently permits only description to change in place. Changes to rateLimit, GPU, serving mode, replicas, parallelism, KV-cache settings, or other runtime features return DEPLOYMENT_V2_NEW_GENERATION_REQUIRED; create a new Deployment generation for those changes.

Response Fields​

TypeFields
ModelServingResourcePreviewVOclusterId, clusterName, creatable, blockerCodes, blockers, warnings, workloadPlan, preflightV2, acceleratorPool, runtimeBackend
InferenceServiceVOid, name, status, statusMessage, modelName, modelSource, backend, servingMode, clusterId, clusterName, endpoint, readyReplicas, totalReplicas, capabilities
DeploymentRestartPreviewVOdeploymentId, restartable, blockerCode, blockerMessage, resourcePreview
DeploymentEventVOid, serviceId, eventType, message, details, createdAt
const profileRequest = {
name: "example-profile-deployment",
clusterId: 1,
modelSource: "gallery",
modelName: "Qwen3-4B-Instruct-2507-FAST",
backend: "sglang",
servingMode: "standard",
gpuType: "L20",
replicas: 1,
};
const assetRequest = {
name: "example-asset-deployment",
clusterId: 1,
modelAssetId: 42,
backend: "sglang",
servingMode: "standard",
gpuType: "L20",
replicas: 1,
};

const features = await client.deployments.features();
const mutationState = await client.deployments.capabilities();
const clusters = await client.deployments.clusterOptions();
const inventory = await client.deployments.modelInventory(1);

const preview = await client.deployments.preview<{ creatable: boolean }>(profileRequest);
if (preview.creatable) await client.deployments.create(profileRequest);

const assetPreview = await client.deployments.preview(assetRequest);
const created = await client.deployments.create<{ id: number }>(assetRequest);
const page = await client.deployments.list({ page: 1, pageSize: 20, status: "RUNNING" });
const detail = await client.deployments.get(created.id);
const events = await client.deployments.events(created.id);
const renderedYaml = await client.deployments.yaml(created.id);
const updated = await client.deployments.update(created.id, {
description: "validated",
});
const stopped = await client.deployments.stop(created.id);
const restartCheck = await client.deployments.previewRestart<{ restartable: boolean }>(created.id);
if (restartCheck.restartable) await client.deployments.restart(created.id);
await client.deployments.delete(created.id);

All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.