Skip to main content

Overview

Explore training methods and supported models for fine-tuning AI models on MSP.

Fine-tuning adapts pre-trained models to specific tasks and datasets to improve domain-specific performance. MSP supports three training methods (SFT, DPO, and CPT) and multiple fine-tuning methods (LoRA, Full-Parameter Fine-Tuning) to make fine-tuning efficient and effective.

Training Methods​

MSP supports the following training methods for different use cases and data types.

SFT: Supervised Fine-Tuning​

Supervised Fine-Tuning trains models on labeled input-output pairs to learn specific tasks or adapt to new domains.

Supervised Fine-Tuning - LLM

Optimize language models for text-based tasks such as:

  • Conversational AI and chatbots
  • Text classification and sentiment analysis
  • Content generation and summarization
  • Domain-specific language understanding
Supervised Fine-Tuning - VLM

Adapt vision-language models for multimodal tasks including:

  • Image captioning and description
  • Visual question answering
  • Document understanding and OCR
  • Medical image analysis

DPO: Direct Preference Optimization​

Advanced training method that optimizes models based on human preferences and comparative feedback.

Direct Preference Optimization - LLM

Ideal for scenarios requiring human-aligned outputs:

  • Improving response quality and safety
  • Aligning model behavior with human values
  • Reducing harmful or biased outputs
  • Fine-tuning based on preference rankings
Direct Preference Optimization - VLM

Align vision-language model outputs with human preferences for multimodal scenarios such as:

  • Image description quality improvement
  • Visual reasoning accuracy alignment
  • Multimodal response safety enhancement

CPT: Continual Pre-Training​

Continual Pre-Training adapts a pre-trained model to new domains by continuing the pre-training process on domain-specific data. Unlike SFT which teaches the model specific tasks through question-answer pairs, CPT injects new domain knowledge and terminology into the model's weights.

Continual Pre-Training

Domain Adaptation - Expand a model's knowledge base for specialized domains such as medical, legal, or finance. CPT is commonly used to adapt models to new industries or support new languages.

Knowledge Injection

Terminology Learning - Inject domain-specific terminology and concepts into the model's weights. CPT uses full-parameter training to ensure thorough domain adaptation.

Choosing the Right Method
  • Use SFT-LLM for language tasks with clear input-output examples.
  • Use SFT-VLM for multimodal applications involving images and text.
  • Use DPO when you have preference data or need to align model behavior with human values.
  • Use CPT when you need to inject domain knowledge or adapt a model to a specialized industry.

Fine-Tuning Methods​

MSP supports the following fine-tuning methods to control how model parameters are updated during training. LoRA is used by default for all training methods.

LoRA (Default)

Parameter-efficient fine-tuning. Trains small adapter modules while keeping original weights frozen.

Best for:

  • Most use cases
  • Cost-efficient training
  • Quick experiments
  • Limited GPU resources

Availability: All training methods

Full-Parameter Fine-Tuning

Updates all model parameters during training. Provides maximum model capacity and flexibility.

Best for:

  • Maximum customization needs
  • When LoRA performance is insufficient
  • Large high-quality datasets

Availability: SFT-LLM, SFT-VLM, CPT

Supported Models​

MSP supports fine-tuning for a wide range of state-of-the-art models across different providers and model types.

Model NameModel TypeTraining MethodModel Size
Qwen/Qwen3-4B-Instruct-2507LLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full4B
Qwen/Qwen3.5-0.8BVLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full0.8B
Qwen/Qwen3.5-2BVLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full2B
Qwen/Qwen3.5-4BVLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full4B
Qwen/Qwen3.5-9BVLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full9B
Qwen/Qwen3.5-27BVLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full27B
Qwen/Qwen3.5-35B-A3BVLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full35B
Qwen/Qwen3.5-122B-A10BVLMSFT-LoRA only122B
Qwen/Qwen3.6-27BVLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full27B
Qwen/Qwen3.6-35B-A3BVLMSFT-LoRA, SFT-Full, DPO-LoRA, CPT-Full35B
Expanding Model Support

We continuously add support for new models. If you need a specific model that's not listed, please contact our support team or check our Model Gallery for the latest additions.

Next Steps​

Ready to start fine-tuning? Choose your approach based on your specific requirements:

Text Fine-Tuning

Learn how to fine-tune language models for text-based applications.

Vision Fine-Tuning

Explore multimodal fine-tuning for vision-language tasks.

Preference Optimization

Implement DPO for human-aligned model behavior.