AI Dataset Preparations - Python SDK
client.datasets.preparations manages AI-assisted dataset preparation and
labeling tasks. Request bodies use server wire names exactly as shown below.
Overview
Available Operations
| Method | Description |
|---|---|
cluster_options() | List clusters selectable for preparation workloads. |
available_provider_keys() | List caller-owned provider keys suitable for labeling. |
preview(body) | Preflight the selected cluster. |
create(body) | Create a preparation task. |
list(body) | Search and page preparation tasks. |
get(id=...) | Read task configuration, progress, and artifacts. |
update(id, body) | Update an editable preparation task. |
generate_rules(id=...) | Start labeling-rule generation. |
start_labeling(body) | Start labeling with optional rule text. |
action(body) | Cancel, retry, or delete a task. |
download_url(body) | Create a temporary URL for a task artifact. |
cluster_options
List clusters selectable for preparation workloads.
Request
This method has no parameters and sends no request body.
Response
list[WorkloadClusterOptionVO].
available_provider_keys
List caller-owned provider keys suitable for labeling.
Request
This method has no parameters and sends no request body.
Response
PrepProviderKeysVO.
preview
Preflight the selected cluster.
Request
Body: clusterId (required).
Response
WorkloadAdmissionPreviewVO.
create
Create a preparation task.
Request
PrepTaskCreateRequest; required fields below.
Response
PrepTaskVO.
list
Search and page preparation tasks.
Request
PrepTaskQueryRequest.
Response
PageResult[PrepTaskVO].
get
Read task configuration, progress, and artifacts.
Request
Keyword-only id required.
Response
PrepTaskDetailVO.
update
Update an editable preparation task.
Request
id and PrepTaskUpdateRequest required.
Response
Updated PrepTaskVO.
generate_rules
Start labeling-rule generation.
Request
Keyword-only id required.
Response
None.
start_labeling
Start labeling with optional rule text.
Request
preparationId required; optional labelRules.
Response
None.
action
Cancel, retry, or delete a task.
Request
preparationId and action (CANCEL, RETRY, DELETE) required.
Response
None.
download_url
Create a temporary URL for a task artifact.
Request
preparationId plus artifactPath or storageRef.
Response
DownloadUrl.
Field Reference and Examples
Use client.wait_for_dataset_preparation(id, until="rules_ready") after rule
generation, and call it without until after labeling to wait for completed.
The rules_ready target also succeeds when the task has already advanced to a
later successful processing phase or completed between polls.
Create body fields
| Field | Type | Required | Constraints |
|---|---|---|---|
taskName | str | Yes | Maximum 255 characters; no control characters. |
clusterId | int | Yes | Selected cluster ID. |
modelType | str | Yes | llm or vlm. |
postTrainingMethod | str | Yes | sft or ref_distill. |
preparationMode | str | Yes | base or sss-bench. |
providerKeyIds | list[int] | Yes | Up to 50 provider-key IDs; may be empty when the mode does not require one. |
scenario | str | No | Scenario text, maximum 4000 characters. |
autoSplitPercent | int | No | 1..100. |
generatedDatasetFileLines | int | No | Requested generated record count. |
processingMode | str | No | auto or manual. |
unlabeledDataFiles / evaluationDataFiles | list[DatasetFileItem] | No | Up to 100 items each. |
tBenchConfig | dict | No | Benchmark-mode options. |
DatasetFileItem accepts fileName, filePath, fileSize, storageRef,
sourceType, and preparationGroup. Update accepts the same configuration
fields as create, but all are optional.
Key response fields: PrepTaskVO includes id, taskName, clusterId,
clusterName, status, progress, progressInfo, savedDatasetId,
savedDatasetName, error, resourcePreview, createdAt, and updatedAt.
PrepTaskDetailVO additionally contains configuration, selected keys, input
files, artifacts, rules, execution summary, and result summary.
clusters = client.datasets.preparations.cluster_options()
provider_keys = client.datasets.preparations.available_provider_keys()
resource = client.datasets.preparations.preview({"clusterId": 1})
task = client.datasets.preparations.create({
"taskName": "prepare-example-dataset",
"clusterId": 1,
"modelType": "llm",
"postTrainingMethod": "sft",
"preparationMode": "base",
"providerKeyIds": [],
})
task_id = task["id"]
page = client.datasets.preparations.list({
"taskName": "prepare-example",
"status": ["pending", "completed"],
"pageNum": 1,
"pageSize": 20,
})
detail = client.datasets.preparations.get(id=task_id)
updated = client.datasets.preparations.update(task_id, {"taskName": "new-name"})
client.datasets.preparations.generate_rules(id=task_id)
client.datasets.preparations.start_labeling({
"preparationId": task_id,
"labelRules": "Return one concise label.",
})
artifact = client.datasets.preparations.download_url({
"preparationId": task_id,
"artifactPath": "outputs/result.jsonl",
})
client.datasets.preparations.action({"preparationId": task_id, "action": "CANCEL"})
All methods use the shared authentication and typed error behavior described in Response Conventions and Retry and Security.