Frequently Asked Questions
This page answers common questions about Smart Studio. Expand a question to read the answer. For step-by-step procedures and complete requirements, follow the related documentation links.
Deployment and Infrastructure
What are the minimum hardware requirements for Smart Studio and hosted models?
Smart Studio supports customer-provided infrastructure. The required GPU, CPU, memory, and storage configuration depends on the models and workloads you plan to host. See BYO-GPU Requirements for the detailed requirements.
If you plan to run an API resale business, see API Router Requirements.
Does Smart Studio support both on-premise and cloud deployment?
Yes. Smart Studio supports on-premise deployment as well as deployment on Alibaba Cloud or other clouds using the customer's own resources and account.
Can Smart Studio run on GPU bare metal?
Yes. GPU bare metal can be used for a basic deployment. High-availability deployments require both CPU and GPU resources.
Can Smart Studio use GPUs from multiple regions or cities?
Yes, with limitations:
- Smart Studio supports one cluster per region, with clusters managed centrally.
- A single cluster cannot span multiple regions.
- GPUs in the same region can be connected through one cluster.
- Cloud-based models can be used while model services are deployed to different clusters.
- Proprietary models need to be deployed separately in each region where the compute resources are located.
How do I add new GPU or CPU nodes?
Adding new GPU or CPU nodes to the Kubernetes cluster is supported. Because this is a potentially disruptive operation, contact the Smart Studio technical team before making the change.
High Availability
What is the minimum number of nodes required for an HA cluster?
An HA cluster requires a minimum of 3 nodes. Fewer than 3 nodes cannot provide standard Kubernetes control-plane high availability.
See High Availability Deployment for the complete architecture and planning guidance.
How are node roles divided in an HA cluster?
Three nodes provide Kubernetes control-plane high availability. The remaining nodes provide workload and user-console high availability.
How are 3-node HA clusters deployed?
In a 3-node deployment, all three nodes act as Masters and run both control- plane components and workload or user-console services.
How many node failures can a 3-node HA cluster tolerate?
A standard 3-node HA cluster can tolerate the loss of at most 1 node. When a node becomes unavailable, workloads and the user console are rescheduled to healthy nodes according to health checks.
How does an HA cluster scale beyond 3 nodes?
Adding nodes increases available workload capacity. Each additional node provides approximately 0.8x linear throughput scaling, with the actual value depending on the workload and the limiting CPU or memory ratio.
How should I plan an HA cluster for production?
Use at least 3 nodes with equivalent specifications, keep at least 2 nodes healthy to sustain service, and add nodes to increase throughput rather than over-provisioning a single node. See the HA deployment guide for resource planning details.
Network and API
Can I use SSH keys instead of a password?
Yes. When SSH keys are used for VM-to-VM interconnection, the password can be left empty. The password is used for internal VM-to-VM interconnection, not for external VM access.
How do I configure HTTPS for Smart Studio?
Use the following architecture:
User Browser --HTTPS--> Customer LB/WAF --HTTP--> Server:80 (Kubernetes Ingress)
Configure the customer-side load balancer to:
- Listen on port 443 with a TLS certificate.
- Forward traffic to port 80 on the server.
- Set the
X-Forwarded-Forheader. - Set
X-Forwarded-Proto: https. - Set WebSocket and SSE timeouts to 600 seconds or longer.
What network bandwidth is recommended for multi-data-center deployment?
For deployments across different IDCs or between a customer data center and the cloud, bandwidth of 1000 Mbps or higher is recommended. This is mainly used to download the sglang image and model weight files during deployment.
How are TPM and RPM limits handled?
For Router Monetization, TPM limits follow Model Studio's rules and higher quotas can be requested for enterprise customers. For GPU Monetization, customers can set TPM limits for end users in the Admin Console.
How can I estimate TPM capacity for a 3-node API Router cluster?
For the reference configuration of 3 CPU nodes, with 64 vCPU and 128 GB memory per node, reserve resources for the highly available platform first:
- Console resources:
12C × 3 = 36Cand24G × 3 = 72G. - Total resources:
64C × 3 = 192Cand128G × 3 = 384G. - Gateway resources:
192C − 36C = 156Cand384G − 72G = 312G.
If one Gateway unit of 4C + 8G serves approximately 20M TPM, the estimated capacity is:
min(156C / 4C, 312G / 8G) × 20M TPM
= min(39, 39) × 20M TPM
= 780M TPM
This is a capacity estimate for the reference configuration. Actual capacity depends on the model, request length, concurrency, and workload pattern.
Which TPM setting takes effect when both an admin and an end user set a limit?
The effective TPM value is the lower of the administrator's setting and the end user's setting.
Models
How are Smart Studio platform models updated?
In a networked environment, new models can be updated from the same day to T+3 after release. Commercial Edition deployment and updates are handled by the Smart Studio technical team. Self-service Edition requires network connectivity throughout the update process.
Can customers deploy their own models?
Yes. Customers can deploy their own models on Smart Studio. Customer-owned models do not receive the platform's standard inference acceleration. A value-added acceleration service may be available for fine-tuned models based on supported open-source models.
Does auto-routing support self-deployed models?
No. Auto-routing currently supports Router Monetization and API-based model calling. Self-deployed models are not currently supported.
What is the cache service used in Smart Studio?
Smart Studio uses Advanced KV-Cache, a proprietary implementation developed by Smart Studio, different from the standard KV Cache. The optimized model deployed on Smart Studio has higher TPS, lower TTFT, and lower latency. KV-Cache is an internal optimization for the model itself, not an external service, and only applies to self-deployed open-source models.
Billing and Usage
How does Smart Studio pricing work?
Smart Studio uses a token revenue-sharing model. See the Billing documentation for the current pricing rules.
How are token revenue-share charges calculated?
The calculation is:
Consumed Tokens × Standard Model Price × Sharing Rate
The sharing rate depends on the model and monetization scenario.
Are the Minimum Guarantee and Token Revenue Share charged twice?
No. The customer pays the higher of the Minimum Guarantee and the calculated Token Revenue Share.
Are prepaid and pay-as-you-go payment models supported?
Yes. Smart Studio supports both prepaid and pay-as-you-go models.
Can administrators set credits, spending limits, or token quotas for end users?
Yes. Administrators can provide free credits and set rate limits for end users.
Can API access be suspended automatically after a quota or spending threshold is reached?
Yes. API access can be suspended automatically when a predefined quota or spending threshold is reached.