SupraLabs Research

Supra SLM Leaderboard

An open comparison of small language models (<150M parameters) across standardized benchmarks (zero-shot unless stated otherwise). Anyone can submit a model — community or SupraLabs — via a Space discussion, with benchmarks run locally and reviewed before publication.

Parameter range
Organization
# Model Params ARC-e HellaSwag PIQA Winogrande MMLU Average

Model comparison

How to submit a model

The leaderboard accepts community submissions. Run the benchmarks locally following the methodology below, open a discussion on the Space with your results, and a SupraLabs maintainer will validate and publish it.

  1. Run the 5 benchmarks (ARC-e, HellaSwag, PIQA, Winogrande, MMLU) using the harness noted in the methodology.
  2. Open a discussion on the Space with the template below filled in.
  3. Wait for review, reproducibility is checked before an entry is added to the table.
  4. Published, the model appears tagged Community or Supra, depending on its origin.
  5. REMEMBER! — "—" on the leaderboard means not evaluated on this benchmark. It is left out of the average, and the model is marked "partial".
### New model — [model name] repo: org/model-name params: 42M precision: fp16 arc_e: 0.00 hellaswag: 0.00 piqa: 0.00 winogrande: 0.00 mmlu: 0.00 harness: lm-eval-harness vX.X.X n_shot: 0 seed: 0 logs: link to raw output
Evaluation methodology

Scores are accuracy (or acc_norm where applicable), reported as a percentage. Evaluations are zero-shot unless the model row indicates otherwise: each model carries an n-shot tag, and a model evaluated with a different number of shots is not directly comparable with the others. Submissions must state their n_shot. The Average column is the plain mean of the benchmarks that were actually evaluated — no difficulty weighting. Models evaluated on fewer than 5 benchmarks are tagged partial and their average is not directly comparable with models evaluated on all 5. We recommend running evaluations with the lm-evaluation-harness to keep submissions comparable.

BenchmarkWhat it measuresMetric
ARC-eBasic scientific reasoning (AI2 Reasoning Challenge, easy set)acc_norm
HellaSwagCommonsense sentence completionacc_norm
PIQAEveryday physical reasoningacc_norm
WinograndeCoreference ambiguity resolutionacc
MMLUMultitask knowledge across 57 domainsacc

Community submissions get the community tag; models published under the SupraLabs organization get supra. Results that aren't reproducible or lack execution logs are not accepted.