Skip to main content

Benchmark Evaluation

Benchmark Evaluation measures model performance against official datasets and standardized metrics. This provides precise, quantifiable insights into model capabilities.

Currently, only LLM Benchmark Single evaluation is supported.

Step 1: Configure basic settings​

Select a target cluster, then configure the evaluation parameters.

  • Cluster: Choose an available cluster from the dropdown. The system checks whether the cluster has sufficient capacity. If not, choose a different cluster.
  • Display Name: Enter a name for this evaluation task.
  • Model Type: Select LLM (Large Language Models for text tasks). VLM support coming soon.
  • Evaluation Method: Choose how to evaluate your model.
  • Evaluation Type: Select Single. (Comparative eval coming soon)
  • Model: Select the model to evaluate. Both platform-deployed models and Router models are supported.

Benchmark Evaluation


Token Consumption

The Router model consumes large token volumes. Confirm your Provider Key is configured. Check that your account has sufficient balance before proceeding.

Step 2: Select benchmark dataset & metrics​

Select a public benchmark dataset to measure model performance against industry standards.

  1. Choose a Benchmark Dataset Choose from standard benchmarks:

General knowledge & reasoning​

  • MMLU: 57 subjects across STEM, humanities, and social sciences
  • HellaSwag: Commonsense reasoning
  • DROP: Discrete reasoning and numerical inference
  • ARC: Science and world‑knowledge reasoning
  • BBH: Hard BIG‑bench subset requiring multi‑step reasoning

Math reasoning​

  • GSM8K: Multi‑step elementary math
  • MATH: Competition‑level math and formal reasoning

Code generation​

  • HumanEval: Functional correctness via unit tests

Multilingual evaluation​

  • MGSM: Multilingual GSM8K across 10 languages
  • MMMLU: Multilingual MMLU in 14 languages
  1. Select a dataset version Select Smoke for a 1-sample evaluation, Lite version for faster evaluation, and Full version for complete results.

Benchmark Evaluation


Verify all configurations, then click Start Evaluation. The evaluation progress appears in the list.


Eval-list


Step 3: View evaluation report​

After the evaluation completes, review the detailed report to analyze model performance.

Score overview​

Score Overview provides a summary of the evaluation results.

  • Quality Score: Represents the model's overall weighted score.
  • Evaluated Items: Shows the number of evaluated samples out of the total.
  • Metadata: Lists key information about the job.
  • Download Results: Click to download the complete evaluation results.

Quality-score


Performance metrics​

Performance Metrics provides visual insights into your evaluation results. Use these insights to identify where the model excels.

  • Performance Comparison: The radar chart compares model capabilities across all configured metrics.

  • Detailed Metrics Comparison: Lists the exact numerical scores for each evaluation metric across all models.


Performance Metrics


Case analysis​

Review individual evaluation cases for error analysis. This helps you understand specific model behaviors.


case-analysis


Next steps​

Fine-Tune Based on Results

Adapt a base model to your domain using SFT, DPO, or CPT.

Deploy the Best Model

Deploy a model from the gallery or upload your own model weights.