Skip to main content

AI Auto Evaluation

AI Auto Evaluation tests your model with custom datasets, LLM-as-Judge scoring, and customizable metrics. This method is ideal for real-world or domain-specific testing.

Step 1: Configure basic settings​

Select a target cluster, then configure the evaluation parameters.

  • Cluster: Choose an available cluster for model evaluation. Evaluation only consumes CPU resources.
  • Model Type: Select LLM (Large Language Models for text tasks) or VLM (Vision-Language Models for multimodal tasks). Note: VLM currently supports AI Auto Evaluation only.
  • Evaluation Method: Choose how to evaluate your model.
  • Evaluation Type: Select Single or Comparative.
  • Model: Select the model to evaluate. Both platform-deployed models and Router models are supported.
  • Judge Model (LLM-as-Judge): Select a judge model to compare the model’s predicted answers against the ground truth answers.

Tip: For Comparative Evaluation, select two models to evaluate side by side.


Basic Configuration


Token Consumption

Both the Router model and Judge Model consume large token volumes. Confirm your Provider Key is configured. Check that your account has sufficient balance before proceeding.

Step 2: Configure datasets & metrics​

Upload an evaluation dataset and define the metrics for evaluating model performance.

1. Dataset setup​

Upload your evaluation dataset:

  • Use OSS: Provide the OSS path to your JSONL dataset.
  • Select Existing Dataset: Choose an existing dataset from the console.
  • Create Dataset: Create a new evaluation dataset. For best results, use the AI Data Prep Tool to prepare your evaluation data.
Dataset Format

All datasets must be in JSONL format. For detailed format requirements and examples, refer to the Dataset Format Requirements guide.


Dataset Setup


  • MAX-Sample Count (Optional): Set the maximum number of samples to evaluate. Leave empty to evaluate all samples.

2. Metrics configuration​

Define the evaluation scene and select the metrics for evaluating model performance.

  1. Select a Scene for your evaluation:
SceneDescription
Definitive QuestionsEvaluates responses to questions with standard, verifiable answers.
Open-ended QAEvaluates responses to open-ended questions without a single correct answer.

Scene Description: Review and edit the scene description to provide context for the evaluation.

  1. Configure metrics for your selected scene:
  • Predefined Metrics: Select predefined metrics to evaluate the model.
  • Custom Metrics: Create custom metrics to meet specific needs (e.g, Instruction Following, Safety).
  • Weight Distribution: Assign a weight to each metric.
Evaluation Criteria Preview

Preview the system prompt and user prompt template before execution.


Metrics Configuration


Verify all configurations, then click Start Evaluation. The evaluation progress will display in the list.


Eval-list


Step 3: View evaluation report​

After the evaluation completes, review the detailed report to analyze model performance. The layout adapts for Single or Comparative evaluations.

Score overview​

Score Overview provides a summary of the evaluation results.

  • Quality Score: Represents the model's overall weighted score. In a Comparative Evaluation, scores for all models appear side-by-side.
  • Evaluated Items: Shows the number of evaluated samples out of the total.
  • Metadata: Lists key information about the job.
  • Download Results: Click to download the complete evaluation results.

Quality-score


Performance metrics​

Performance Metrics provides visual insights into your evaluation results. Use these insights to identify where each model excels.

  • Score Comparison: Displays model accuracy across different questions and benchmark categories.
  • Performance Comparison: The radar chart compares model capabilities across all configured metrics.
  • Detailed Metrics Comparison: Lists the exact numerical scores for each evaluation metric across all models.

Performance Metrics


Case analysis​

Review individual evaluation cases for error analysis. This helps you understand specific model behaviors.

Single Evaluation: The report defaults to Failure Case Analysis. Low-performance samples are highlighted.

Comparative Evaluation: The default Side-by-Side View shows how different models responded to the same prompt. Review the judge's score and reasoning for each response to understand performance differences.


side-by-side comparison


Analysis & recommendations (comparative evaluation only)​

Analysis & Recommendations provides a model recommendation based on the comparison.


analysis


Dataset format​

Select the dataset format that matches your evaluation type.

AI-Auto evaluation datasets must match the format generated by the AI-Data-Prep tool.

{
"init": { // Initial version
"messages": [ // Conversation list
{ "role": "user", "content": "..." }, // User question
{ "role": "assistant", "content": "..." } // Model answer
],
"question_type": "fact_retrieval", // Question type
"config_post_training_method": "sft", // Training method
"metadata": { // Metadata
"entity": "...", // Core entity
"entities_used": ["..."] // List of entities used
}
},
"confirm": "confirmed", // Review status
"refined": { // Refined version
"messages": [ // Same structure as init.messages
{ "role": "user", "content": "..." },
{ "role": "assistant", "content": "..." }
]
}
}

Next steps​

Fine-Tune Based on Results

Adapt a base model to your domain using SFT, DPO, or CPT.

Deploy the Best Model

Deploy a model from the gallery or upload your own model weights.