Skip to main content

High Availability Deployment

This reference describes the high-availability architecture for running Smart Studio on customer-managed local IDC Kubernetes clusters, including storage, databases, GPU management, networking, and cross-region control.

Architecture & Reliability Design

Smart Studio High Availability (HA) Deployment Plan

This document describes the high-availability architecture of Smart Studio running on customer-managed local IDC infrastructure. It covers the Kubernetes control plane, CPU-based platform services, storage, databases, GPU management, model serving, networking, and cross-region orchestration.

Scope Third-party local IDC take-overPlatform K8s · Sealos · CiliumOrchestration ACK One + local K8s + ACK ProAudience Solution architects / SRE

1Overview & Design Principles

Smart Studio runs as a self-contained platform in customer-owned local IDC data centers. Models, images, data, and logs stay on local HA infrastructure, while training, inference, and the AI Gateway run on local GPU and CPU clusters.

3-NodeStacked etcd and control-plane quorum
Ops-IndependentLocal scheduling and serving continue during uplink outages
Zero-TouchNo changes to the IDC AI-compute network

Core Design Principles

  • Local-first persistence. Models, images, datasets, and logs reside on HA local infrastructure.
  • Control-plane autonomy. The Kubernetes cluster keeps scheduling and serving workloads when the cloud Ops connection is unavailable.
  • Non-intrusive compute integration. GPU Operator, DCGM, NFD, and Prometheus adapt to the customer's GPU environment.
  • Pluggable scheduling. The design supports Volcano, Leader Worker Set (LWS), and Ray scheduling patterns.
HA Objective Every layer is designed so a single node, pod, or backend failure degrades service without halting the platform.

2IDC Environment Constraints

HA decisions are driven by the constraints commonly found in customer IDC environments.

Constraint #1 — No Public IP

Use a secure Cloud Assistant channel for remote operations and monitoring without exposing a public IP.

Constraint #2 — No OSS / NAS

Self-build the persistence layer for databases, the image registry, and model files.

Constraint #3 — Bare K8s Only

Middleware, monitoring, storage, and registry services must be deployed in the customer cluster.

3Reference Topology

The foundation is a three-node stacked control plane behind a Sealos lvscare VIP, with GPU and CPU worker nodes carrying AI workloads.

Kubernetes Node Layer — High Availability
┌──────────────────────────────────────────────┐
│ Kubernetes Node Layer · Cilium · Containerd   │
│  ┌──────────── Master HA (3 nodes) ─────────┐  │
│  │ Master-1 · Master-2 · Master-3           │  │
│  │ etcd · API Server · Scheduler · CNI      │  │
│  └───────────────┬──────────────────────────┘  │
│       lvscare VIP → kube-apiserver             │
│  ┌────────────── Worker Nodes ──────────────┐  │
│  │ GPU workers · device plugin · DCGM        │  │
│  └──────────────────────────────────────────┘  │
└──────────────────────────────────────────────┘
HA MechanismThe three-node etcd cluster tolerates the loss of one master while retaining write quorum.

4Storage High Availability

A self-built NFS backend provides shared persistence for databases, Harbor, and model files, so NFS availability is a first-class HA concern.

WorkloadRecommended patternPurpose
MySQL / Harbor PVCsnfs-client StorageClassDynamic isolated subdirectories per PVC
Model filesStatic PV/PVC with RWXConcurrent read-only access by inference pods
NFS serversKeepalived VIP + Active-BackupSimple failover over customer-provided backend storage
Known LimitNFS failover is not zero-interruption. Transient I/O errors may occur during the 3–10 second VIP failover window, so applications must implement retries.

5Database High Availability

Databases are self-deployed in the bare Kubernetes cluster and use replication to limit the impact of a node or pod failure.

MySQL

Use Bitnami MySQL with primary/secondary replication for the current deployment. Percona Operator is the recommended evolution for automatic failover, hot backup, and read/write splitting.

Redis

Use Bitnami Redis in replication and Sentinel mode. Sentinel elects a new master automatically, with a bounded failover window.

Known LimitMySQL secondary promotion is manual in the current design and should be treated as a near-term hardening item.

6GPU Management High Availability

GPU Operator manages the driver and runtime lifecycle, while Volcano and LWS provide workload-aware scheduling for training and inference.

ToolScheduling levelUse case
VolcanoCluster / jobDistributed training, gang scheduling, queues, and preemption
RayApplication task / actorFine-grained distributed training and online workloads
LWSLeader-worker pod groupTopology-aware LLM inference
GPU OperatorNFDnvidia-container-toolkitGPU Feature Discoverydevice pluginDCGM exporter
Known LimitsThe current Smart Studio integration deploys models on a single node. Volcano and GPU Operator versions must be revalidated after Kubernetes or driver upgrades.

7Network & Proxy High Availability

Cilium provides the cluster network and observability. The Cloud Assistant xt-assist-server proxy carries the Smart Studio interaction channel over the existing outbound agent connection.

Why use the proxy?

  • Removes the standalone frp/FPC tunnel and public FPS dependency.
  • Reuses the existing xt-agent long-lived connection.
  • Keeps the IDC free of inbound ports and external exposure.

Accessible services

HarborGrafanapipe-mgtsss-deploy
Known LimitThe proxy is the single access channel to the cluster. Large model-file transfers should use in-IDC NFS or offline media instead.

8Cross-Region & IDC Control High Availability

ACK One registers local IDC and regional ACK Pro clusters as a fleet. Workloads can be scheduled across healthy GPU and CPU pools while the local IDC continues operating autonomously during an Ops-plane outage.

Cross-Region + Local IDC Control
AliCloud fleet plane (ACK One)
             │ job scheduling
     ┌───────┴────────┐
 Local IDC cluster   Regional ACK Pro clusters
  local K8s + HA       GPU / CPU pools
  services + GPUs      independent scaling
HA OutcomeControl is distributed across the cloud fleet plane, the local IDC control plane, and regional clusters, so one domain can fail without taking down the entire platform.

9Failure Domain & Recovery Summary

LayerHA mechanismFailover behavior
Control plane3-node etcd + lvscare VIPTolerates one master loss
StorageKeepalived VIP + Active-Backup3–10 second VIP failover
MySQLPrimary / secondaryManual promotion today; automation planned
RedisReplication + SentinelAutomatic master election
GPU schedulingGPU Operator + Volcano / LWSGang scheduling and pod restart
Remote accessxt-assist-server proxyOutbound connection without inbound ports
Cross-regionACK One fleet + local autonomyIndependent regional scheduling
Overall PostureThe design provides layered HA with bounded failover windows. The key hardening items are automated MySQL failover and protection of the xt-assist-server interaction channel.