An open comparison of small language models (<150M parameters) across standardized benchmarks (zero-shot unless stated otherwise). Anyone can submit a model — community or SupraLabs — via a Space discussion, with benchmarks run locally and reviewed before publication.
| # | Model | Params | ARC-e | HellaSwag | PIQA | Winogrande | MMLU | Average |
|---|
The leaderboard accepts community submissions. Run the benchmarks locally following the methodology below, open a discussion on the Space with your results, and a SupraLabs maintainer will validate and publish it.
Scores are accuracy (or acc_norm where applicable), reported as a percentage. Evaluations are
zero-shot unless the model row indicates otherwise: each model carries an
n-shot tag, and a model evaluated with a different number of shots is not directly comparable with the others.
Submissions must state their n_shot.
The Average column is the plain mean of the benchmarks that were actually evaluated — no difficulty weighting.
Models evaluated on fewer than 5 benchmarks are tagged partial
and their average is not directly comparable with models evaluated on all 5.
We recommend running evaluations with the
lm-evaluation-harness
to keep submissions comparable.
| Benchmark | What it measures | Metric |
|---|---|---|
| ARC-e | Basic scientific reasoning (AI2 Reasoning Challenge, easy set) | acc_norm |
| HellaSwag | Commonsense sentence completion | acc_norm |
| PIQA | Everyday physical reasoning | acc_norm |
| Winogrande | Coreference ambiguity resolution | acc |
| MMLU | Multitask knowledge across 57 domains | acc |
Community submissions get the community tag; models published under the SupraLabs organization get supra. Results that aren't reproducible or lack execution logs are not accepted.