CPU benchmark results

Small language models, timed on an ordinary CPU

A tour through 57 benchmark runs: how quickly each model reads a prompt, how quickly it writes, and how much it still makes sense while doing so. Everything here ran on ordinary CPU threads, no GPU.
Feel like wanting to have your model here too? Use CPU LM Benchmark to evaluate your model on the same hardware like the rest of the models here and wait until I update this page :)

About the models on this page

The page covers 57 base models: the pretrained versions, before any instruction tuning or chat training. Most come from public releases on the Hugging Face Hub. Some are from large labs, such as Google's Gemma 3, Alibaba's Qwen3, Hugging Face's SmolLM2, LiquidAI's LFM2.5 and Tencent's Hunyuan, along with the RWKV and Mamba research models. Others are small community projects, with names like BananaMind, Supra, Ivme and Surjo. GPT-2 and DistilGPT-2 are included as older reference points.

The models range from a few thousand parameters to one billion. The page and the leaderboard hide anything under 10 million parameters, because tiny models look fast simply because there is so little of them. Fine-tuned, instruct and chat models are excluded, since this board compares base models.

Every model was timed the same way, on a CPU with 2 threads and no GPU. For each one, the benchmark ran prompts of 64, 256, 1024, 4096 and 11264 tokens, then generated 64 new tokens. Prefill is how fast the model reads the prompt, and decode is how fast it writes. Each speed is a median of several runs, so one slow run doesn't skew it.

Quality is the II, the Intelligence Index from the Open SLM Leaderboard. It combines HellaSwag, ARC, PIQA and ArithMark-3 scores, each adjusted so that random guessing scores 0 and a perfect score scores 100. Higher is better. It only exists for models on that leaderboard, so other rows show a dash. It's a rough signal, not a verdict on how useful a model is.

The Index column combines speed across the whole prompt ladder with how much of that ladder the model can run. A model that handles 11264 tokens gets credit for it, and a model whose context window ends at 1024 tokens doesn't. The small "ladder" number under each Index shows how many of the five prompt lengths were run.

Treat the numbers with some caution. The benchmarks ran on shared hardware, so speeds drift with other load, and results from different days aren't perfectly comparable. A few entries are noisy or odd. One model shows a negative decode speed at its longest prompt, for example. Those entries are left in, because the raw results are kept as they were recorded.

Data from FlameF0X/lm-cpu-benchmarks

Quickest to write

Decode speed: tokens generated per second, median across runs. Longer bar is faster.

Loading the results…

Every model

Click a row to see how each prompt length behaved. Click a column heading to sort.

Loading…

Speed against quality

Each dot is one model. Further right is faster. Higher on the chart means a better II. Hover a dot to see its name.

About the numbers

Prefill is how fast the model processes the prompt, and decode is how fast it writes new tokens. Both are median throughput across the runs at each prompt length. II is higher-is-better; quality and speed index are the benchmark's own composite scores, so read them as relative rankings rather than absolute truths. Runs used the CPU threads recorded in each file, typically 2.