Models - Python SDK
client.models reads the deployable model catalog. It does not create My Model
assets; use client.my_models for caller-owned assets.
Pass the catalog's modelId to get() and deployment_capabilities(). Pass
its modelName unchanged when creating a Deployment; do not construct either
value from the provider label.
Overview
Available Operations
| Method | Description |
|---|---|
list(...) | Search and page the model catalog. |
get(model_id) | Read one catalog model. |
deployment_capabilities(model_id, cluster_id=None) | Resolve supported deployment choices, optionally for a cluster. |
list
Search and page the model catalog.
Request
Optional keyword, model_type, provider, source, page, page_size.
Response
PageResult[ModelInfoVO].
get
Read one catalog model.
Request
model_id: str (required).
Response
ModelInfoVO.
deployment_capabilities
Resolve supported deployment choices, optionally for a cluster.
Request
model_id: str required; cluster_id: str | None.
Response
ModelDeploymentCapabilitiesVO.
Field Reference and Examples
Key response fields
| Type | Fields |
|---|---|
ModelInfoVO | modelId, modelName, displayName, modelType, provider, source, modelSize, deployable, supportedGpuTypes, supportedServingModes, deploymentOptions |
ModelDeploymentCapabilitiesVO | modelId, modelName, clusterId, gpuTypeOptions, recommendedGpuType, acceleratorPools, clusterGpuTypes |
page = client.models.list(
keyword="Qwen",
model_type="LLM",
provider="Qwen",
source="gallery",
page=1,
page_size=20,
)
model = client.models.get("qwen3-4b")
choices = client.models.deployment_capabilities(
"qwen3-4b",
cluster_id="1",
)
All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.