Evaluations - Python SDK
client.evaluations creates benchmark, automatic, and comparison evaluations
against external models, API keys, deployed services, or model assets.
Overview
Available Operations
| Method | Description |
|---|---|
available_models() | List model references currently available to Evaluation. |
cluster_options() | List clusters selectable for Evaluation. |
preview(body) | Preflight the selected cluster. |
create(body) | Create an Evaluation job. |
get(id) | Read Evaluation status, scores, metrics, and report state. |
report(id) | Create a temporary URL for the completed report. |
artifact_download_url(id, artifact_ref) | Create a temporary URL for one VLM/media artifact. |
available_models
List model references currently available to Evaluation.
Request
This method has no parameters and sends no request body.
Response
AvailableModelsVO.
cluster_options
List clusters selectable for Evaluation.
Request
This method has no parameters and sends no request body.
Response
list[WorkloadClusterOptionVO].
preview
Preflight the selected cluster.
Request
Body requires clusterId.
Response
WorkloadAdmissionPreviewVO.
create
Create an Evaluation job.
Request
CreateEvalJobRequest fields below.
Response
CreateEvalJobVO with jobId, resourcePreview.
get
Read Evaluation status, scores, metrics, and report state.
Request
id: str required.
Response
EvalJobDetailVO.
report
Create a temporary URL for the completed report.
Request
id: str required.
Response
DownloadUrl.
artifact_download_url
Create a temporary URL for one VLM/media artifact.
Request
Job id and a server-provided artifact_ref required.
Response
DownloadUrl.
Field Reference and Examples
Ordinary LLM evaluations expose their completed report through report(id).
artifact_download_url(...) is only usable when an Evaluation media result
provides an artifactRef; callers should not manufacture this identifier.
Create body fields
| Field | Type | Required | Constraints |
|---|---|---|---|
kind | str | Yes | benchmark, auto, or compare. |
modelType | str | Yes | LLM or VLM. |
models | list[ModelRef] | Yes | Up to two models. |
judge | ModelRef | No | Optional judge model. |
dataset | str | Yes | Dataset name/reference, maximum 128 characters. |
metricConfig | dict | No | Metric-specific options. |
maxSamples | int | No | Maximum 1,000,000. |
clusterId | int | Yes | Selected Evaluation cluster. |
ModelRef uses type plus snake_case references. An external model uses
provider_key_id and model_id. A deployed service uses both
msp_api_key_id and service_id; neither field is optional for that form.
Key EvalJobDetailVO fields: jobId, status, progress, kind,
evaluationMethod, evaluationType, modelType, datasetName, clusterId,
scores, metrics, reportAvailable, error, and timestamps.
available = client.evaluations.available_models()
clusters = client.evaluations.cluster_options()
resource = client.evaluations.preview({"clusterId": 1})
created = client.evaluations.create({
"kind": "benchmark",
"modelType": "LLM",
"models": [{
"type": "service",
"msp_api_key_id": "7",
"service_id": "42",
}],
"dataset": "evaluation-dataset",
"maxSamples": 100,
"clusterId": 1,
})
job_id = created["jobId"]
detail = client.evaluations.get(job_id)
report = client.evaluations.report(job_id)
All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.