Skip to main content

Common Workflows

These are Human integration examples for the complete SDK surface. The msp-operations Skill must check its current capability manifest before using any command or method. In the first Skill release, workload creation, upload, update, cancellation, restart, stop, and deletion remain disabled pending contract tests, sandbox end-to-end evidence, and Agent evaluation.

Upload and Deploy a Model​

The upload helper validates files, calculates content metadata, transfers the directory, and completes the model asset before returning it. Deployment creation then uses the asset id.

model = client.my_models.upload(
"./model",
name="example-model",
model_type="LLM",
)

request = {
"name": "example-deployment",
"modelAssetId": model["id"],
"backend": "sglang",
"servingMode": "standard",
"gpuType": "your-gpu-type",
"replicas": 1,
"clusterId": 1,
}

preview = client.deployments.preview(request)
if preview.get("creatable"):
deployment = client.deployments.create(request)
running = client.wait_for_deployment(deployment["id"])

Choose the cluster and accelerator values returned by the discovery and preview methods. The placeholder values above are not deployment recommendations.

Upload a Dataset and Start Training​

Build the request from the active Training capabilities, a selected cluster, the uploaded Dataset, and a deployable base model. Preview the exact resource specification before creating the job.

dataset = client.datasets.upload(
"./train.jsonl",
name="example-dataset",
dataset_type="training",
training_category="sft-llm",
)

preview = client.training.preview({
"clusterId": 1,
"resourceSpecId": "resource-spec-id",
})

request = {
"clientToken": "training-request-001",
"displayName": "example-training",
"outputModelName": "example-output",
"recipeId": "recipe-id",
"recipeVersion": "recipe-version",
"baseModelRef": {"type": "recipe_model", "id": "model-id"},
"datasetRefs": [{"datasetId": str(dataset["id"]), "role": "train"}],
"placement": {"clusterId": "1", "resourceSpecId": "resource-spec-id"},
"params": {},
}

if preview.get("decision") == "FIT":
job = client.training.create(request)
completed = client.wait_for_training_job(job["jobId"])

Replace the placeholder recipe, model, cluster, and resource identifiers with values returned by the current environment. Do not copy resource selections between environments. outputModelName is only the output suffix. Keep it within the model-specific limit returned by Training capabilities so the final <base>-FT-<suffix> name remains deployable.

Create an Evaluation​

request = {
"kind": "benchmark",
"modelType": "LLM",
"models": [{
"type": "external",
"provider_key_id": "provider-key-id",
"model_id": "model-id",
}],
"dataset": "evaluation-dataset",
"maxSamples": 100,
"clusterId": 1,
}

preview = client.evaluations.preview({"clusterId": request["clusterId"]})
if preview.get("decision") == "FIT":
job = client.evaluations.create(request)
client.wait_for_evaluation_job(job["jobId"])
report = client.evaluations.report(job["jobId"])

Use the report method for ordinary evaluations. Media-oriented evaluations may return a separate artifact reference for download.

Safety Rules​

  • Preview compute requirements before creating workloads.
  • Treat create, update, stop, restart, cancel, and delete as state-changing operations.
  • Use idempotency fields supplied by the API when retrying creation after an uncertain result.
  • Fetch pre-signed object URLs without adding the platform bearer token.
  • Verify final state with get, list, or a waiter instead of relying only on the initial response.