Skip to main content

Model Serving Platform

The Model Serving Platform (MSP) is the centralized workspace for deploying, training, and managing AI models on your GPU infrastructure. It provides end-to-end capabilities for cluster management, model deployment, fine-tuning, and observability.

Key capabilities​

CapabilityDescription
Cluster ManagementRegister and monitor Kubernetes GPU clusters, configure resource groups, and track node health.
Model DeploymentDeploy pre-trained models from Hugging Face or custom sources to GPU clusters.
Model TrainingFine-tune models with SFT, DPO, or CPT methods using your datasets.
ObservabilityMonitor deployment performance, resource utilization, and serving metrics in real-time.

Access the Model Serving Platform​

Navigate to GPU Monetization > Model Serving Platform or visit the Model Serving Platform directly.

Prerequisites​

Before activating the Model Serving Platform, verify that your GPU infrastructure meets the following requirements:

RequirementDescription
Kubernetes ClusterA running Kubernetes cluster with GPU nodes.
GPU ResourcesAvailable GPU resources based on your model's requirements.
kubeconfigA valid kubeconfig file with sufficient permissions for the platform to manage workloads.
StorageAdequate storage capacity for model weights and artifacts.

Get started​

To start using the Model Serving Platform:

  1. Visit the Model Serving Platform activation page to activate the service.
  2. Register your GPU cluster, deploy models, and configure observability.

Next Step​

After deploying your model service in MSP, return to the Admin Console to set pricing and list the model for your customers.