BYO-GPU Requirements
This guide covers the requirements and infrastructure needed to deploy Smart Studio on your own local IDC (Internet Data Center) with your GPU cluster.
Standard Requirements for IDC Deployment
The following CPU and storage servers are required. These servers must be on the same local virtual network as the compute servers that host the Compute Cards.
| No. | Basic Server Configuration | Quantity | Purpose |
|---|---|---|---|
| 1 | CPU: 64 cores, Memory: 128GB, Storage: 1T+ | 3 | Deploy the Kubernetes control plane and Smart Studio platform services (required) |
| 2 | GPU: depends on the local GPU configuration of the customer and the size of the LLM | N | See the Minimum Configuration for Different Model Deployments |
Demo of Network Topology
Minimum Requirements for IDC Deployment (POC)
For a minimal proof-of-concept setup, a single GPU server can be used.
| No. | Basic Server Configuration | Quantity | Purpose |
|---|---|---|---|
| 1 | GPU: H20/H100/H200/RTX PRO 6000, CPU: 128+ cores, Memory: 1TB, Storage: 1T+ | 1 | All services — Kubernetes control plane, CNI, Smart Studio platform, and Model Serving deployed on one GPU server (without HA guarantee) |
Demo of Network Topology
Minimum Configuration for Different Model Deployments
You can provide your required model name to us, and we will respond with the GPU requirements.
| Model | Minimum GPU Specifications (Estimate) | Base Memory Pool Size |
|---|---|---|
| Qwen3.8-Flash-Next (125B-A6B, FP8) | 4 Cards: H20-96G, H20-141G, A800-80G, RTX PRO 6000-96G, L20-48G, L40S-48G, or equivalent with ≥48 GB GPU memory per card | >1 TB |
| Qwen3.8-27B (FP8) | 1 Card: H20-96G, H20-141G, A800-80G, or equivalent with ≥80 GB GPU memory per card 2 Cards: L20-48G, L40S-48G, or equivalent with ≥48 GB GPU memory per card | >1 TB |
| Qwen3.7-Max | 8 Cards: ≥288 GB GPU memory per card | >15 TB |
| Qwen3.7-Plus | 8 Cards: H20-141G, H20-96G, or equivalent with ≥80 GB GPU memory per card (8-card deployment recommended) | >4 TB |
| Qwen3.6-Flash (35B-A3B, FP8) | 1 Card: H20-96G, H20-141G, A800-80G, or equivalent with ≥80 GB GPU memory per card 2 Cards: L20-48G, L40S-48G, or equivalent with ≥48 GB GPU memory per card | >1 TB |
| Qwen3.6-27B (FP8) / Qwen3.6-35B-A3B (FP8) | 1 Card: H20-96G, H20-141G, A800-80G, or equivalent with ≥80 GB GPU memory per card 2 Cards: L20-48G, L40S-48G, or equivalent with ≥48 GB GPU memory per card | >1.5 TB |
| Qwen3.5-397B-A22B (FP8) | 4 Cards: ≥188 GB GPU memory per card 8 Cards: H20-96G, H20-141G, or equivalent with ≥80 GB GPU memory per card | >4 TB |
| DeepSeek-V4-Flash-0731 (284B, FP4/FP8 mixed precision) | 4 Cards: H20-96G, H20-141G, A800-80G, RTX PRO 6000-96G, or equivalent with ≥80 GB GPU memory per card | >1 TB |
| Deepseek V4-Flash | 4 Cards: H20-96G, H20-141G, RTX PRO 6000-96G, or equivalent with ≥80 GB GPU memory per card | >1 TB |
| GLM-5.3-Flash (320B-A18B, FP8) | 8 Cards: H20-96G, H20-141G, A800-80G, RTX PRO 6000-96G, or equivalent with ≥80 GB GPU memory per card | >1 TB |
| GLM-5.2 (FP8) | 4 Cards: ≥288 GB GPU memory per card 8 Cards: H20-141G, or equivalent with ≥141 GB GPU memory per card 16 Cards: H20-96G, or equivalent with ≥80 GB GPU memory per card | >8 TB |
Note
- Estimations for Qwen3.7-Plus/Max are based on FP8 dtype and may not be fully accurate.
- You are responsible for providing the required GPU resources (purchased or self-owned), as the platform does not include them.
- For more technical details about deployment or other model deployment configurations, please contact us.
Infrastructure for Deploying Smart Studio on a Local IDC
GPU/CPU infrastructure → Scheduler and Model Serving → Smart Studio control plane
Next Step
Once your infrastructure meets the requirements above, proceed to Activate Cluster to connect your servers and deploy the Smart Studio platform.