Small language models, timed on an ordinary CPU
A tour through 57 benchmark runs: how quickly each model reads a prompt, how quickly it writes,
and how much it still makes sense while doing so. Everything here ran on ordinary CPU threads, no GPU.
Feel like wanting to have your model here too? Use CPU LM Benchmark
to evaluate your model on the same hardware like the rest of the models here and wait until I update this page :)
About the models on this page
The page covers 57 base models: the pretrained versions, before any instruction tuning or chat training. Most come from public releases on the Hugging Face Hub. Some are from large labs, such as Google's Gemma 3, Alibaba's Qwen3, Hugging Face's SmolLM2, LiquidAI's LFM2.5 and Tencent's Hunyuan, along with the RWKV and Mamba research models. Others are small community projects, with names like BananaMind, Supra, Ivme and Surjo. GPT-2 and DistilGPT-2 are included as older reference points.
The models range from a few thousand parameters to one billion. The page and the leaderboard hide anything under 10 million parameters, because tiny models look fast simply because there is so little of them. Fine-tuned, instruct and chat models are excluded, since this board compares base models.
Every model was timed the same way, on a CPU with 2 threads and no GPU. For each one, the benchmark ran prompts of 64, 256, 1024, 4096 and 11264 tokens, then generated 64 new tokens. Prefill is how fast the model reads the prompt, and decode is how fast it writes. Each speed is a median of several runs, so one slow run doesn't skew it.
Quality is the II, the Intelligence Index from the Open SLM Leaderboard. It combines HellaSwag, ARC, PIQA and ArithMark-3 scores, each adjusted so that random guessing scores 0 and a perfect score scores 100. Higher is better. It only exists for models on that leaderboard, so other rows show a dash. It's a rough signal, not a verdict on how useful a model is.
The Index column combines speed across the whole prompt ladder with how much of that ladder the model can run. A model that handles 11264 tokens gets credit for it, and a model whose context window ends at 1024 tokens doesn't. The small "ladder" number under each Index shows how many of the five prompt lengths were run.
Treat the numbers with some caution. The benchmarks ran on shared hardware, so speeds drift with other load, and results from different days aren't perfectly comparable. A few entries are noisy or odd. One model shows a negative decode speed at its longest prompt, for example. Those entries are left in, because the raw results are kept as they were recorded.
Data from FlameF0X/lm-cpu-benchmarks
Quickest to write
Decode speed: tokens generated per second, median across runs. Longer bar is faster.
Every model
Click a row to see how each prompt length behaved. Click a column heading to sort.
| Model | Params | Prefill tok/s | Decode tok/s | II | Quality | Index |
|---|---|---|---|---|---|---|
| Loading… | ||||||
Speed against quality
Each dot is one model. Further right is faster. Higher on the chart means a better II. Hover a dot to see its name.
About the numbers
Prefill is how fast the model processes the prompt, and decode is how fast it writes new tokens. Both are median throughput across the runs at each prompt length. II is higher-is-better; quality and speed index are the benchmark's own composite scores, so read them as relative rankings rather than absolute truths. Runs used the CPU threads recorded in each file, typically 2.
The overview shows the numbers. This tab asks what they mean: where speed comes from, what falls off with longer prompts, what looks strange, and what these measurements can't tell you.
The size trend
Decode speed against parameter count on log axes, with a straight-line fit. A model above the dashed line is faster than its size predicts.
Small-model speed, beaten by bigger models
Models that run at least 50% faster than the trend predicts for their size. Each one is listed with how many smaller models it outruns.
Which architectures are fastest
Median decode speed for each architecture family, with the number of models in brackets. A family can be quick at small sizes and slow at large ones, so read this next to the size chart.
Where speeds cluster
How many models fall into each decode-speed band, from slowest to fastest.
Quality against size
Bigger does not always mean better predictions. Look for models that sit well below the general slope.
How much speed survives long prompts
Decode speed at the longest prompt measured, as a share of its speed at 64 tokens. Lower means the model slows more as context grows.
Anomalies
Entries that look wrong or unreliable. They stay in the data as recorded, so treat them with care.
Limits of these numbers
Speed over sequence length
Search for a model to see how its prefill and decode speed change as the prompt gets longer. The dashed lines extend each model's own trend far past its longest measured prompt, out to one million tokens. That extension is a hypothetical prediction, not a measurement.
A running log of what changed on this page, newest first. Entries marked new arrived since you last opened this tab.
Read status is kept in memory for this visit only, so it resets when the page is reloaded.