Overview
Explore training methods and supported models for fine-tuning AI models on MSP.
Fine-tuning adapts pre-trained models to specific tasks and datasets to improve domain-specific performance. MSP supports three training methods (SFT, DPO, and CPT) and multiple fine-tuning methods (LoRA, Full-Parameter Fine-Tuning) to make fine-tuning efficient and effective.
Training Methods
MSP supports the following training methods for different use cases and data types.
SFT: Supervised Fine-Tuning
Supervised Fine-Tuning trains models on labeled input-output pairs to learn specific tasks or adapt to new domains.
Optimize language models for text-based tasks such as:
- Conversational AI and chatbots
- Text classification and sentiment analysis
- Content generation and summarization
- Domain-specific language understanding
Adapt vision-language models for multimodal tasks including:
- Image captioning and description
- Visual question answering
- Document understanding and OCR
- Medical image analysis
DPO: Direct Preference Optimization
Advanced training method that optimizes models based on human preferences and comparative feedback.
Ideal for scenarios requiring human-aligned outputs:
- Improving response quality and safety
- Aligning model behavior with human values
- Reducing harmful or biased outputs
- Fine-tuning based on preference rankings
Align vision-language model outputs with human preferences for multimodal scenarios such as:
- Image description quality improvement
- Visual reasoning accuracy alignment
- Multimodal response safety enhancement
CPT: Continual Pre-Training
Continual Pre-Training adapts a pre-trained model to new domains by continuing the pre-training process on domain-specific data. Unlike SFT which teaches the model specific tasks through question-answer pairs, CPT injects new domain knowledge and terminology into the model's weights.
Domain Adaptation - Expand a model's knowledge base for specialized domains such as medical, legal, or finance. CPT is commonly used to adapt models to new industries or support new languages.
Terminology Learning - Inject domain-specific terminology and concepts into the model's weights. CPT uses full-parameter training to ensure thorough domain adaptation.
- Use SFT-LLM for language tasks with clear input-output examples.
- Use SFT-VLM for multimodal applications involving images and text.
- Use DPO when you have preference data or need to align model behavior with human values.
- Use CPT when you need to inject domain knowledge or adapt a model to a specialized industry.
Fine-Tuning Methods
MSP supports the following fine-tuning methods to control how model parameters are updated during training. LoRA is used by default for all training methods.
Parameter-efficient fine-tuning. Trains small adapter modules while keeping original weights frozen.
Best for:
- Most use cases
- Cost-efficient training
- Quick experiments
- Limited GPU resources
Availability: All training methods
Updates all model parameters during training. Provides maximum model capacity and flexibility.
Best for:
- Maximum customization needs
- When LoRA performance is insufficient
- Large high-quality datasets
Availability: SFT-LLM, SFT-VLM, CPT
Supported Models
MSP supports fine-tuning for a wide range of state-of-the-art models across different providers and model types.
| Model Name | Model Type | Training Method | Model Size |
|---|---|---|---|
Qwen/Qwen3-4B-Instruct-2507 | LLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 4B |
Qwen/Qwen3.5-0.8B | VLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 0.8B |
Qwen/Qwen3.5-2B | VLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 2B |
Qwen/Qwen3.5-4B | VLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 4B |
Qwen/Qwen3.5-9B | VLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 9B |
Qwen/Qwen3.5-27B | VLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 27B |
Qwen/Qwen3.5-35B-A3B | VLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 35B |
Qwen/Qwen3.5-122B-A10B | VLM | SFT-LoRA only | 122B |
Qwen/Qwen3.6-27B | VLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 27B |
Qwen/Qwen3.6-35B-A3B | VLM | SFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full | 35B |
We continuously add support for new models. If you need a specific model that's not listed, please contact our support team or check our Model Gallery for the latest additions.
Next Steps
Ready to start fine-tuning? Choose your approach based on your specific requirements: