Observability - Python SDK
client.observability reads caller-scoped cluster, Deployment, and generic
workload metrics. Snapshot methods return the latest state; timeseries methods
return bounded historical points.
Overview
Available Operations
| Method | Description |
|---|---|
cluster_overview(cluster_id) | Read aggregate cluster workload utilization. |
services(cluster_id=None) | List observable model-serving Deployments. |
service_snapshot(service_id) | Read the newest snapshot for one Deployment. |
service_timeseries(service_id, ...) | Read Deployment engine/resource points. |
workloads(cluster_id=None) | List observable workloads across supported domains. |
workload_snapshot(business_type, business_id) | Read one generic workload snapshot. |
workload_timeseries(business_type, business_id, ...) | Read generic CPU, memory, restart, and GPU points. |
history(...) | Cursor-page workload lifecycle history. |
cluster_overview
Read aggregate cluster workload utilization.
Request
cluster_id required.
Response
UserClusterObservabilityOverviewVO.
services
List observable model-serving Deployments.
Request
Optional keyword-only cluster_id.
Response
list[UserWorkloadObservabilityVO].
service_snapshot
Read the newest snapshot for one Deployment.
Request
service_id required.
Response
UserWorkloadObservabilityVO.
service_timeseries
Read Deployment engine/resource points.
Request
service_id required; optional range_hours, max_points.
Response
list[UserServiceMetricsPointVO].
workloads
List observable workloads across supported domains.
Request
Optional keyword-only cluster_id.
Response
list[UserWorkloadObservabilityVO].
workload_snapshot
Read one generic workload snapshot.
Request
business_type, business_id required.
Response
UserWorkloadObservabilityVO.
workload_timeseries
Read generic CPU, memory, restart, and GPU points.
Request
business_type, business_id required; optional range_hours, max_points.
Response
list[UserWorkloadMetricsPointVO].
history
Cursor-page workload lifecycle history.
Request
Optional cluster_id, cursor, limit.
Response
UserWorkloadHistoryPageVO.
Field Reference and Examples
Use the exact businessType and businessId returned by workloads() for
generic snapshot/timeseries calls rather than constructing identifiers.
Key response fields
| Type | Fields |
|---|---|
UserClusterObservabilityOverviewVO | clusterId, activeWorkloads, cpuUsageCores, memoryWorkingSetBytes, observedPlacements, plannedPlacements, scope |
UserWorkloadObservabilityVO | identity, lifecycle state, replicas/shards, CPU/memory/restarts, GPU allocation/utilization, serving QPS/latency/throughput, completeness, evidence |
UserServiceMetricsPointVO | snapshotTime, replicas/shards, latency percentiles, QPS/request count, TTFT/ITL, error rates, throughput, KV-cache hit rate, input/output lengths |
UserWorkloadMetricsPointVO | snapshotTime, CPU, working set/RSS memory, restart count, GPU activity/utilization/memory, completeness states |
UserWorkloadHistoryPageVO | items, nextCursor, contractVersion, scope |
overview = client.observability.cluster_overview(cluster_id=1)
services = client.observability.services(cluster_id=1)
snapshot = client.observability.service_snapshot(service_id=42)
service_points = client.observability.service_timeseries(
service_id=42,
range_hours=1,
max_points=120,
)
workloads = client.observability.workloads(cluster_id=1)
if workloads:
workload = workloads[0]
latest = client.observability.workload_snapshot(
workload["businessType"],
workload["businessId"],
)
points = client.observability.workload_timeseries(
workload["businessType"],
workload["businessId"],
range_hours=1,
max_points=120,
)
history = client.observability.history(cluster_id=1, limit=100)
All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.