Evaluations - TypeScript SDK
client.evaluations creates benchmark, automatic, and comparison evaluations
against external models, API keys, deployed services, or model assets.
Overview
Available Operations
| Method | Description |
|---|---|
availableModels() | List model references available to Evaluation. |
clusterOptions() | List clusters selectable for Evaluation. |
preview(body) | Preflight the selected cluster. |
create(body) | Create an Evaluation job. |
get(id) | Read status, scores, metrics, and report state. |
report(id) | Create a temporary URL for the completed report. |
artifactDownloadUrl(id, artifactRef) | Create a temporary URL for one VLM/media artifact. |
availableModels
List model references available to Evaluation.
Request
This method has no parameters and sends no request body.
Response
AvailableModelsVO.
clusterOptions
List clusters selectable for Evaluation.
Request
This method has no parameters and sends no request body.
Response
WorkloadClusterOptionVO[].
preview
Preflight the selected cluster.
Request
Body requires clusterId.
Response
WorkloadAdmissionPreviewVO.
create
Create an Evaluation job.
Request
Fields below.
Response
CreateEvalJobVO with jobId, resourcePreview.
get
Read status, scores, metrics, and report state.
Request
id: string required.
Response
EvalJobDetailVO.
report
Create a temporary URL for the completed report.
Request
id: string required.
Response
DownloadUrl.
artifactDownloadUrl
Create a temporary URL for one VLM/media artifact.
Request
Job id and a server-provided artifactRef required.
Response
DownloadUrl.
Field Reference and Examples
Ordinary LLM evaluations expose their completed report through report(id).
artifactDownloadUrl(...) is only usable when an Evaluation media result
provides an artifactRef; callers should not manufacture this identifier.
Create body fields
| Field | Type | Required | Constraints |
|---|---|---|---|
kind | string | Yes | benchmark, auto, or compare. |
modelType | string | Yes | LLM or VLM. |
models | ModelRef[] | Yes | Up to two models. |
judge | ModelRef | No | Optional judge model. |
dataset | string | Yes | Dataset name/reference, maximum 128 characters. |
metricConfig | object | No | Metric-specific options. |
maxSamples | number | No | Maximum 1,000,000. |
clusterId | number | Yes | Selected Evaluation cluster. |
ModelRef uses type plus snake_case references. An external model uses
provider_key_id and model_id. A deployed service uses both
msp_api_key_id and service_id; neither field is optional for that form.
Key EvalJobDetailVO fields: jobId, status, progress, kind,
evaluationMethod, evaluationType, modelType, datasetName, clusterId,
scores, metrics, reportAvailable, error, and timestamps.
const available = await client.evaluations.availableModels();
const clusters = await client.evaluations.clusterOptions();
const resource = await client.evaluations.preview({ clusterId: 1 });
const created = await client.evaluations.create<{ jobId: string }>({
kind: "benchmark",
modelType: "LLM",
models: [{
type: "service",
msp_api_key_id: "7",
service_id: "42",
}],
dataset: "evaluation-dataset",
maxSamples: 100,
clusterId: 1,
});
const detail = await client.evaluations.get(created.jobId);
const report = await client.evaluations.report(created.jobId);
All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.