Training - Python SDK
client.training manages fine-tuning jobs. Inputs are logical model, Dataset,
recipe, and placement references; raw paths, images, commands, environments,
and credentials are not accepted by the create API.
Overview
Available Operations
| Method | Description |
|---|---|
capabilities() | Read the active recipe/resource catalog. |
cluster_options() | List clusters selectable for Training. |
preview(body) | Preflight a resource specification in a cluster. |
knowledge_teacher_models(...) | Find compatible knowledge-distillation teachers. |
create(body) | Create a fine-tuning job. |
list(body=None) | Page Training jobs. |
get(id) | Read job progress, losses, actions, and artifacts. |
artifact_download_url(artifact_id) | Create temporary download URLs for a completed artifact. |
cancel(id) | Request cancellation of a non-terminal job. |
capabilities
Read the active recipe/resource catalog.
Request
This method has no parameters and sends no request body.
Response
TrainCapabilitiesVO.
cluster_options
List clusters selectable for Training.
Request
This method has no parameters and sends no request body.
Response
list[WorkloadClusterOptionVO].
preview
Preflight a resource specification in a cluster.
Request
clusterId, resourceSpecId required.
Response
WorkloadAdmissionPreviewVO.
knowledge_teacher_models
Find compatible knowledge-distillation teachers.
Request
Keyword-only student_model_id, recipe_id required.
Response
list[TrainKnowledgeTeacherVO].
create
Create a fine-tuning job.
Request
CreateTrainJobRequest fields below.
Response
{jobId}.
list
Page Training jobs.
Request
Optional body: pageNum (default 1), pageSize (default 20), status.
Response
PageResult[TrainJobDetailVO].
get
Read job progress, losses, actions, and artifacts.
Request
id: str required.
Response
TrainJobDetailVO.
artifact_download_url
Create temporary download URLs for a completed artifact.
Request
artifact_id: str required.
Response
{"urls": [str], "files"?: [{"path", "sizeBytes", "url"}]}.
cancel
Request cancellation of a non-terminal job.
Request
id: str required.
Response
None.
Field Reference and Examples
Create body fields
| Field | Type | Description |
|---|---|---|
clientToken | str | Idempotency token for safe create retries. |
displayName | str | Human-facing job name. |
outputModelName | str | Output model suffix. Its maximum length depends on baseModelRef.id because the final deployable name is <base>-FT-<suffix>; use the capability limit (the tested 4B Profile permits 14 characters). |
recipeId, recipeVersion | str | Recipe identity selected from capabilities(). |
baseModelRef | dict | {type, id} logical base-model reference. |
teacherRef | dict | Optional teacher reference; fields described below. |
datasetRefs | list[dict] | Each item uses {datasetId, role}. |
placement | dict | {clusterId, nodeId?, resourceSpecId}. |
params | dict | Recipe hyperparameters. |
teacherRef is the exception to the API's normal camelCase convention: its
wire keys are provider_key_id, model_id, service_id, and
model_asset_id, plus type. Do not send camelCase variants.
Key response fields: TrainCapabilitiesVO contains schemaVersion,
catalogVersion, catalogDigest, and entries. TrainJobDetailVO contains
id, status, stage, progress, taskDisplayName, baseModel,
trainingMethod, artifacts, actions, loss series, deployment references,
timestamps, and error/status details.
catalog = client.training.capabilities()
clusters = client.training.cluster_options()
resource = client.training.preview({
"clusterId": 1,
"resourceSpecId": "resource-spec-id",
})
teachers = client.training.knowledge_teacher_models(
student_model_id="model-id",
recipe_id="recipe-id",
)
request = {
"clientToken": "training-request-001",
"displayName": "example-training",
"outputModelName": "example-output",
"recipeId": "recipe-id",
"recipeVersion": "recipe-version",
"baseModelRef": {"type": "recipe_model", "id": "model-id"},
"datasetRefs": [{"datasetId": "dataset-id", "role": "train"}],
"placement": {"clusterId": "1", "resourceSpecId": "resource-spec-id"},
"params": {},
}
created = client.training.create(request)
job_id = created["jobId"]
page = client.training.list({"pageNum": 1, "pageSize": 20, "status": "RUNNING"})
detail = client.training.get(job_id)
if detail.get("artifacts"):
download = client.training.artifact_download_url(
detail["artifacts"][0]["artifactId"]
)
first_url = (download.get("urls") or [download["url"]])[0]
client.training.cancel(job_id)
All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.