Skip to main content

Evaluations - TypeScript SDK

client.evaluations creates benchmark, automatic, and comparison evaluations against external models, API keys, deployed services, or model assets.

Overview​

Available Operations​

MethodDescription
availableModels()List model references available to Evaluation.
clusterOptions()List clusters selectable for Evaluation.
preview(body)Preflight the selected cluster.
create(body)Create an Evaluation job.
get(id)Read status, scores, metrics, and report state.
report(id)Create a temporary URL for the completed report.
artifactDownloadUrl(id, artifactRef)Create a temporary URL for one VLM/media artifact.

availableModels​

List model references available to Evaluation.

Request​

This method has no parameters and sends no request body.

Response​

AvailableModelsVO.

clusterOptions​

List clusters selectable for Evaluation.

Request​

This method has no parameters and sends no request body.

Response​

WorkloadClusterOptionVO[].

preview​

Preflight the selected cluster.

Request​

Body requires clusterId.

Response​

WorkloadAdmissionPreviewVO.

create​

Create an Evaluation job.

Request​

Fields below.

Response​

CreateEvalJobVO with jobId, resourcePreview.

get​

Read status, scores, metrics, and report state.

Request​

id: string required.

Response​

EvalJobDetailVO.

report​

Create a temporary URL for the completed report.

Request​

id: string required.

Response​

DownloadUrl.

artifactDownloadUrl​

Create a temporary URL for one VLM/media artifact.

Request​

Job id and a server-provided artifactRef required.

Response​

DownloadUrl.

Field Reference and Examples​

Ordinary LLM evaluations expose their completed report through report(id). artifactDownloadUrl(...) is only usable when an Evaluation media result provides an artifactRef; callers should not manufacture this identifier.

Create body fields

FieldTypeRequiredConstraints
kindstringYesbenchmark, auto, or compare.
modelTypestringYesLLM or VLM.
modelsModelRef[]YesUp to two models.
judgeModelRefNoOptional judge model.
datasetstringYesDataset name/reference, maximum 128 characters.
metricConfigobjectNoMetric-specific options.
maxSamplesnumberNoMaximum 1,000,000.
clusterIdnumberYesSelected Evaluation cluster.

ModelRef uses type plus snake_case references. An external model uses provider_key_id and model_id. A deployed service uses both msp_api_key_id and service_id; neither field is optional for that form.

Key EvalJobDetailVO fields: jobId, status, progress, kind, evaluationMethod, evaluationType, modelType, datasetName, clusterId, scores, metrics, reportAvailable, error, and timestamps.

const available = await client.evaluations.availableModels();
const clusters = await client.evaluations.clusterOptions();
const resource = await client.evaluations.preview({ clusterId: 1 });

const created = await client.evaluations.create<{ jobId: string }>({
kind: "benchmark",
modelType: "LLM",
models: [{
type: "service",
msp_api_key_id: "7",
service_id: "42",
}],
dataset: "evaluation-dataset",
maxSamples: 100,
clusterId: 1,
});
const detail = await client.evaluations.get(created.jobId);
const report = await client.evaluations.report(created.jobId);

All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.