Model Serving Platform
The Model Serving Platform (MSP) is the centralized workspace for deploying, training, and managing AI models on your GPU infrastructure. It provides end-to-end capabilities for cluster management, model deployment, fine-tuning, and observability.
Key capabilities
| Capability | Description |
|---|---|
| Cluster Management | Register and monitor Kubernetes GPU clusters, configure resource groups, and track node health. |
| Model Deployment | Deploy pre-trained models from Hugging Face or custom sources to GPU clusters. |
| Model Training | Fine-tune models with SFT, DPO, or CPT methods using your datasets. |
| Observability | Monitor deployment performance, resource utilization, and serving metrics in real-time. |
Access the Model Serving Platform
Navigate to GPU Monetization > Model Serving Platform or visit the Model Serving Platform directly.
Prerequisites
Before activating the Model Serving Platform, verify that your GPU infrastructure meets the following requirements:
| Requirement | Description |
|---|---|
| Kubernetes Cluster | A running Kubernetes cluster with GPU nodes. |
| GPU Resources | Available GPU resources based on your model's requirements. |
| kubeconfig | A valid kubeconfig file with sufficient permissions for the platform to manage workloads. |
| Storage | Adequate storage capacity for model weights and artifacts. |
Get started
To start using the Model Serving Platform:
- Visit the Model Serving Platform activation page to activate the service.
- Register your GPU cluster, deploy models, and configure observability.
Next Step
After deploying your model service in MSP, return to the Admin Console to set pricing and list the model for your customers.