Skip to main content

Training - Python SDK

client.training manages fine-tuning jobs. Inputs are logical model, Dataset, recipe, and placement references; raw paths, images, commands, environments, and credentials are not accepted by the create API.

Overview​

Available Operations​

MethodDescription
capabilities()Read the active recipe/resource catalog.
cluster_options()List clusters selectable for Training.
preview(body)Preflight a resource specification in a cluster.
knowledge_teacher_models(...)Find compatible knowledge-distillation teachers.
create(body)Create a fine-tuning job.
list(body=None)Page Training jobs.
get(id)Read job progress, losses, actions, and artifacts.
artifact_download_url(artifact_id)Create temporary download URLs for a completed artifact.
cancel(id)Request cancellation of a non-terminal job.

capabilities​

Read the active recipe/resource catalog.

Request​

This method has no parameters and sends no request body.

Response​

TrainCapabilitiesVO.

cluster_options​

List clusters selectable for Training.

Request​

This method has no parameters and sends no request body.

Response​

list[WorkloadClusterOptionVO].

preview​

Preflight a resource specification in a cluster.

Request​

clusterId, resourceSpecId required.

Response​

WorkloadAdmissionPreviewVO.

knowledge_teacher_models​

Find compatible knowledge-distillation teachers.

Request​

Keyword-only student_model_id, recipe_id required.

Response​

list[TrainKnowledgeTeacherVO].

create​

Create a fine-tuning job.

Request​

CreateTrainJobRequest fields below.

Response​

{jobId}.

list​

Page Training jobs.

Request​

Optional body: pageNum (default 1), pageSize (default 20), status.

Response​

PageResult[TrainJobDetailVO].

get​

Read job progress, losses, actions, and artifacts.

Request​

id: str required.

Response​

TrainJobDetailVO.

artifact_download_url​

Create temporary download URLs for a completed artifact.

Request​

artifact_id: str required.

Response​

{"urls": [str], "files"?: [{"path", "sizeBytes", "url"}]}.

cancel​

Request cancellation of a non-terminal job.

Request​

id: str required.

Response​

None.

Field Reference and Examples​

Create body fields

FieldTypeDescription
clientTokenstrIdempotency token for safe create retries.
displayNamestrHuman-facing job name.
outputModelNamestrOutput model suffix. Its maximum length depends on baseModelRef.id because the final deployable name is <base>-FT-<suffix>; use the capability limit (the tested 4B Profile permits 14 characters).
recipeId, recipeVersionstrRecipe identity selected from capabilities().
baseModelRefdict{type, id} logical base-model reference.
teacherRefdictOptional teacher reference; fields described below.
datasetRefslist[dict]Each item uses {datasetId, role}.
placementdict{clusterId, nodeId?, resourceSpecId}.
paramsdictRecipe hyperparameters.

teacherRef is the exception to the API's normal camelCase convention: its wire keys are provider_key_id, model_id, service_id, and model_asset_id, plus type. Do not send camelCase variants.

Key response fields: TrainCapabilitiesVO contains schemaVersion, catalogVersion, catalogDigest, and entries. TrainJobDetailVO contains id, status, stage, progress, taskDisplayName, baseModel, trainingMethod, artifacts, actions, loss series, deployment references, timestamps, and error/status details.

catalog = client.training.capabilities()
clusters = client.training.cluster_options()
resource = client.training.preview({
"clusterId": 1,
"resourceSpecId": "resource-spec-id",
})
teachers = client.training.knowledge_teacher_models(
student_model_id="model-id",
recipe_id="recipe-id",
)

request = {
"clientToken": "training-request-001",
"displayName": "example-training",
"outputModelName": "example-output",
"recipeId": "recipe-id",
"recipeVersion": "recipe-version",
"baseModelRef": {"type": "recipe_model", "id": "model-id"},
"datasetRefs": [{"datasetId": "dataset-id", "role": "train"}],
"placement": {"clusterId": "1", "resourceSpecId": "resource-spec-id"},
"params": {},
}
created = client.training.create(request)
job_id = created["jobId"]
page = client.training.list({"pageNum": 1, "pageSize": 20, "status": "RUNNING"})
detail = client.training.get(job_id)
if detail.get("artifacts"):
download = client.training.artifact_download_url(
detail["artifacts"][0]["artifactId"]
)
first_url = (download.get("urls") or [download["url"]])[0]
client.training.cancel(job_id)

All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.