High Availability Deployment
This reference describes the high-availability architecture for running Smart Studio on customer-managed local IDC Kubernetes clusters, including storage, databases, GPU management, networking, and cross-region control.
Smart Studio High Availability (HA) Deployment Plan
This document describes the high-availability architecture of Smart Studio running on customer-managed local IDC infrastructure. It covers the Kubernetes control plane, CPU-based platform services, storage, databases, GPU management, model serving, networking, and cross-region orchestration.
Contents
1Overview & Design Principles
Smart Studio runs as a self-contained platform in customer-owned local IDC data centers. Models, images, data, and logs stay on local HA infrastructure, while training, inference, and the AI Gateway run on local GPU and CPU clusters.
Core Design Principles
- Local-first persistence. Models, images, datasets, and logs reside on HA local infrastructure.
- Control-plane autonomy. The Kubernetes cluster keeps scheduling and serving workloads when the cloud Ops connection is unavailable.
- Non-intrusive compute integration. GPU Operator, DCGM, NFD, and Prometheus adapt to the customer's GPU environment.
- Pluggable scheduling. The design supports Volcano, Leader Worker Set (LWS), and Ray scheduling patterns.
2IDC Environment Constraints
HA decisions are driven by the constraints commonly found in customer IDC environments.
Constraint #1 — No Public IP
Use a secure Cloud Assistant channel for remote operations and monitoring without exposing a public IP.
Constraint #2 — No OSS / NAS
Self-build the persistence layer for databases, the image registry, and model files.
Constraint #3 — Bare K8s Only
Middleware, monitoring, storage, and registry services must be deployed in the customer cluster.
3Reference Topology
The foundation is a three-node stacked control plane behind a Sealos lvscare VIP, with GPU and CPU worker nodes carrying AI workloads.
┌──────────────────────────────────────────────┐ │ Kubernetes Node Layer · Cilium · Containerd │ │ ┌──────────── Master HA (3 nodes) ─────────┐ │ │ │ Master-1 · Master-2 · Master-3 │ │ │ │ etcd · API Server · Scheduler · CNI │ │ │ └───────────────┬──────────────────────────┘ │ │ lvscare VIP → kube-apiserver │ │ ┌────────────── Worker Nodes ──────────────┐ │ │ │ GPU workers · device plugin · DCGM │ │ │ └──────────────────────────────────────────┘ │ └──────────────────────────────────────────────┘
4Storage High Availability
A self-built NFS backend provides shared persistence for databases, Harbor, and model files, so NFS availability is a first-class HA concern.
| Workload | Recommended pattern | Purpose |
|---|---|---|
| MySQL / Harbor PVCs | nfs-client StorageClass | Dynamic isolated subdirectories per PVC |
| Model files | Static PV/PVC with RWX | Concurrent read-only access by inference pods |
| NFS servers | Keepalived VIP + Active-Backup | Simple failover over customer-provided backend storage |
5Database High Availability
Databases are self-deployed in the bare Kubernetes cluster and use replication to limit the impact of a node or pod failure.
MySQL
Use Bitnami MySQL with primary/secondary replication for the current deployment. Percona Operator is the recommended evolution for automatic failover, hot backup, and read/write splitting.
Redis
Use Bitnami Redis in replication and Sentinel mode. Sentinel elects a new master automatically, with a bounded failover window.
6GPU Management High Availability
GPU Operator manages the driver and runtime lifecycle, while Volcano and LWS provide workload-aware scheduling for training and inference.
| Tool | Scheduling level | Use case |
|---|---|---|
| Volcano | Cluster / job | Distributed training, gang scheduling, queues, and preemption |
| Ray | Application task / actor | Fine-grained distributed training and online workloads |
| LWS | Leader-worker pod group | Topology-aware LLM inference |
7Network & Proxy High Availability
Cilium provides the cluster network and observability. The Cloud Assistant xt-assist-server proxy carries the Smart Studio interaction channel over the existing outbound agent connection.
Why use the proxy?
- Removes the standalone frp/FPC tunnel and public FPS dependency.
- Reuses the existing
xt-agentlong-lived connection. - Keeps the IDC free of inbound ports and external exposure.
Accessible services
8Cross-Region & IDC Control High Availability
ACK One registers local IDC and regional ACK Pro clusters as a fleet. Workloads can be scheduled across healthy GPU and CPU pools while the local IDC continues operating autonomously during an Ops-plane outage.
AliCloud fleet plane (ACK One)
│ job scheduling
┌───────┴────────┐
Local IDC cluster Regional ACK Pro clusters
local K8s + HA GPU / CPU pools
services + GPUs independent scaling9Failure Domain & Recovery Summary
| Layer | HA mechanism | Failover behavior |
|---|---|---|
| Control plane | 3-node etcd + lvscare VIP | Tolerates one master loss |
| Storage | Keepalived VIP + Active-Backup | 3–10 second VIP failover |
| MySQL | Primary / secondary | Manual promotion today; automation planned |
| Redis | Replication + Sentinel | Automatic master election |
| GPU scheduling | GPU Operator + Volcano / LWS | Gang scheduling and pod restart |
| Remote access | xt-assist-server proxy | Outbound connection without inbound ports |
| Cross-region | ACK One fleet + local autonomy | Independent regional scheduling |