Model Lab — Python SDK
Interactive model playground: list models, manage sessions, and stream inference.
Overview
The Model Lab resource powers the in-Console playground: a filterable model list, multi-turn session creation, and streaming (SSE) inference that yields OpenAI-compatible chunks.
Accessed via client.model_lab.
Available Operations
| Method | Description |
|---|---|
list() | Get model list for the Model Lab |
createSession() | Create a new chat session for multi-turn conversations |
stream() | Call model stream inference (SSE) |
list
Get model list for the Model Lab.
page_num and page_size are typed as optional but the backend requires them (returns code=1000 if missing). Recommended defaults: page_num=1, page_size=20.
Example Usage
from wlt import WltClient
client = WltClient(api_key="your-api-key", base_url="https://console.example.com")
# List inference models
models = client.model_lab.list(usage_type="inference")
for m in models.data:
print(m.modelName, m.provider)
# Filter by provider and capability
models = client.model_lab.list(
provider="dashscope",
capability="text_to_text",
inputs=["text"],
outputs=["text"],
)
Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| model_name | str | No | None | Model name (fuzzy search) |
| provider | str | No | None | Provider identifier |
| series_provider | str | No | None | Series provider identifier |
| model_type | str | No | None | Model type |
| view_all_flag | int | No | None | View all flag (1=view all) |
| page_size | int | No | None | Items per page. Required by backend (code=1000 if missing). Recommended default: 20 |
| page_num | int | No | None | Page number. Required by backend (code=1000 if missing). Recommended default: 1 |
| usage_type | str | No | None | Usage type: "inference" or "train" |
| training_method | str | No | None | Training method: "sft" or "dpo" |
| tags | list[str] | No | None | Tag filter list |
| capability | str | No | None | Capability filter, e.g. "text_to_text" |
| model_type_list | list[str] | No | None | Model type list (multi-select) |
| origin_providers | list[str] | No | None | Origin provider list (multi-select) |
| inputs | list[str] | No | None | Input type list, e.g. ["text"] |
| outputs | list[str] | No | None | Output type list (max 1 item), e.g. ["image"] |
Response
Returns BaseResponse[list[ModelInfoVO]].
ModelInfoVO fields: see
client.secrets.model_list()above.
Errors
| Code | Exception | When |
|---|---|---|
2000 / 2002 | AuthenticationError | API Key invalid |
2007 | PermissionError | Permission denied |
3001 | AuthenticationError | Token expired and refresh failed |
createSession
Create a new chat session for multi-turn conversations.
Example Usage
from wlt import WltClient
client = WltClient(api_key="your-api-key", base_url="https://console.example.com")
session = client.model_lab.create_session()
session_id = session.data # UUID string
print(f"Session ID: {session_id}")
Parameters
None.
Response
Returns BaseResponse[str]. The data field contains a UUID string (session ID).
Errors
| Code | Exception | When |
|---|---|---|
2000 / 2002 | AuthenticationError | API Key invalid |
2007 | PermissionError | Permission denied |
3001 | AuthenticationError | Token expired and refresh failed |
stream
Call model stream inference (SSE). Yields each SSE data chunk as a raw JSON string.
Example Usage
import json
from wlt import WltClient
from wlt.models.model_lab import OptionsDTO
client = WltClient(api_key="your-api-key", base_url="https://console.example.com")
# Create a session for multi-turn conversation
session = client.model_lab.create_session()
# Stream inference (use a provider + model that is actually served by the
# gateway -- discover via `client.model_lab.list()` / `client.usage.gateway_providers()`.
# On the daily environment, `openrouter` + `openai/gpt-5.4-nano` is a known-good pair).
for chunk in client.model_lab.stream(
provider="openrouter",
model_name="openai/gpt-5.4-nano",
user_prompt="Explain quantum computing in one paragraph.",
session_id=session.data,
options=OptionsDTO(
temperature=0.7,
maxToken=2048,
topP=0.9,
systemPrompt="You are a helpful assistant.",
),
):
data = json.loads(chunk)
# Each chunk contains delta content, finish_reason, usage, etc.
if "content" in data:
print(data["content"], end="", flush=True)
print() # newline after streaming completes
Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| provider | str | Yes | -- | Provider identifier |
| model_name | str | Yes | -- | Model name |
| user_prompt | str | Yes (unless retry) | None | User prompt text |
| session_id | str | No | None | Session ID (for multi-turn conversations, use UUID) |
| task_id | str | No | None | Task ID (for retry, references the failed task) |
| task_type | str | No | "TEXT" | Task type |
| is_retry | bool | No | False | Whether this is a retry request (requires task_id) |
| options | OptionsDTO | No | None | Model parameter configuration (see below) |
OptionsDTO fields:
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| temperature | float | No | None | Temperature, controls output randomness |
| maxToken | int | No | None | Maximum generated token count |
| topP | float | No | None | Top P sampling parameter |
| topK | int | No | None | Top K sampling parameter |
| presencePenalty | float | No | None | Presence penalty |
| frequencyPenalty | float | No | None | Frequency penalty |
| stopSequences | str | No | None | Stop sequences |
| systemPrompt | str | No | None | System prompt |
| enableThinking | bool | No | None | Enable deep thinking |
| enableProgress | bool | No | None | Enable progress push |
| timeoutSeconds | int | No | None | Task timeout (seconds) |
| enableCache | bool | No | None | Enable result caching |
| enableDocumentInlining | bool | No | None | Enable document inlining |
| budgetTokens | int | No | None | Budget token count |
Response
Returns a generator yielding SSE data chunks as raw JSON strings. Each chunk follows the OpenAI-compatible streaming format. The stream ends with a [DONE] sentinel.
Errors
| Code | Exception | When |
|---|---|---|
1001 | ValidationError | Invalid request parameters |
2000 / 2002 | AuthenticationError | API Key invalid |
2007 | PermissionError | Permission denied |
3001 | AuthenticationError | Token expired and refresh failed |
3002 | RateLimitError | Rate limit exceeded |