Skip to main content

AI Dataset Preparations - Python SDK

client.datasets.preparations manages AI-assisted dataset preparation and labeling tasks. Request bodies use server wire names exactly as shown below.

Overview​

Available Operations​

MethodDescription
cluster_options()List clusters selectable for preparation workloads.
available_provider_keys()List caller-owned provider keys suitable for labeling.
preview(body)Preflight the selected cluster.
create(body)Create a preparation task.
list(body)Search and page preparation tasks.
get(id=...)Read task configuration, progress, and artifacts.
update(id, body)Update an editable preparation task.
generate_rules(id=...)Start labeling-rule generation.
start_labeling(body)Start labeling with optional rule text.
action(body)Cancel, retry, or delete a task.
download_url(body)Create a temporary URL for a task artifact.

cluster_options​

List clusters selectable for preparation workloads.

Request​

This method has no parameters and sends no request body.

Response​

list[WorkloadClusterOptionVO].

available_provider_keys​

List caller-owned provider keys suitable for labeling.

Request​

This method has no parameters and sends no request body.

Response​

PrepProviderKeysVO.

preview​

Preflight the selected cluster.

Request​

Body: clusterId (required).

Response​

WorkloadAdmissionPreviewVO.

create​

Create a preparation task.

Request​

PrepTaskCreateRequest; required fields below.

Response​

PrepTaskVO.

list​

Search and page preparation tasks.

Request​

PrepTaskQueryRequest.

Response​

PageResult[PrepTaskVO].

get​

Read task configuration, progress, and artifacts.

Request​

Keyword-only id required.

Response​

PrepTaskDetailVO.

update​

Update an editable preparation task.

Request​

id and PrepTaskUpdateRequest required.

Response​

Updated PrepTaskVO.

generate_rules​

Start labeling-rule generation.

Request​

Keyword-only id required.

Response​

None.

start_labeling​

Start labeling with optional rule text.

Request​

preparationId required; optional labelRules.

Response​

None.

action​

Cancel, retry, or delete a task.

Request​

preparationId and action (CANCEL, RETRY, DELETE) required.

Response​

None.

download_url​

Create a temporary URL for a task artifact.

Request​

preparationId plus artifactPath or storageRef.

Response​

DownloadUrl.

Field Reference and Examples​

Use client.wait_for_dataset_preparation(id, until="rules_ready") after rule generation, and call it without until after labeling to wait for completed. The rules_ready target also succeeds when the task has already advanced to a later successful processing phase or completed between polls.

Create body fields

FieldTypeRequiredConstraints
taskNamestrYesMaximum 255 characters; no control characters.
clusterIdintYesSelected cluster ID.
modelTypestrYesllm or vlm.
postTrainingMethodstrYessft or ref_distill.
preparationModestrYesbase or sss-bench.
providerKeyIdslist[int]YesUp to 50 provider-key IDs; may be empty when the mode does not require one.
scenariostrNoScenario text, maximum 4000 characters.
autoSplitPercentintNo1..100.
generatedDatasetFileLinesintNoRequested generated record count.
processingModestrNoauto or manual.
unlabeledDataFiles / evaluationDataFileslist[DatasetFileItem]NoUp to 100 items each.
tBenchConfigdictNoBenchmark-mode options.

DatasetFileItem accepts fileName, filePath, fileSize, storageRef, sourceType, and preparationGroup. Update accepts the same configuration fields as create, but all are optional.

Key response fields: PrepTaskVO includes id, taskName, clusterId, clusterName, status, progress, progressInfo, savedDatasetId, savedDatasetName, error, resourcePreview, createdAt, and updatedAt. PrepTaskDetailVO additionally contains configuration, selected keys, input files, artifacts, rules, execution summary, and result summary.

clusters = client.datasets.preparations.cluster_options()
provider_keys = client.datasets.preparations.available_provider_keys()
resource = client.datasets.preparations.preview({"clusterId": 1})

task = client.datasets.preparations.create({
"taskName": "prepare-example-dataset",
"clusterId": 1,
"modelType": "llm",
"postTrainingMethod": "sft",
"preparationMode": "base",
"providerKeyIds": [],
})
task_id = task["id"]

page = client.datasets.preparations.list({
"taskName": "prepare-example",
"status": ["pending", "completed"],
"pageNum": 1,
"pageSize": 20,
})
detail = client.datasets.preparations.get(id=task_id)
updated = client.datasets.preparations.update(task_id, {"taskName": "new-name"})
client.datasets.preparations.generate_rules(id=task_id)
client.datasets.preparations.start_labeling({
"preparationId": task_id,
"labelRules": "Return one concise label.",
})
artifact = client.datasets.preparations.download_url({
"preparationId": task_id,
"artifactPath": "outputs/result.jsonl",
})
client.datasets.preparations.action({"preparationId": task_id, "action": "CANCEL"})

All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.