Leaderboard

Ranked exclusively by measured results submitted to the repository. No opaque composite score is used, and ranking is only ever within a comparison group — a set of results the comparison-safety classifier says may honestly be compared with one another.

Higher is better. Measured generation tokens per second.

4 comparison groups, of which 1 contains more than one result and can be ranked.

qwen2.5-0.5b-instruct-q4_k_m.gguf on llama.cpp (llama-server/cuda)

Single result — nothing to compare it against yet.

Generation throughput for qwen2.5-0.5b-instruct-q4_k_m.gguf on llama.cpp (llama-server/cuda)
1llamacpp-178739194512th Gen Intel(R) Core(TM) i9-12900HNVIDIA GeForce RTX 3080 Ti Laptop GPU360.9

qwen2.5-0.5b-instruct-q4_k_m.gguf q4_k_m on llama.cpp (llama-server/auto)

Single result — nothing to compare it against yet.

Generation throughput for qwen2.5-0.5b-instruct-q4_k_m.gguf q4_k_m on llama.cpp (llama-server/auto)
1llamacpp-1789031931-b98435ea12th Gen Intel(R) Core(TM) i9-12900HNVIDIA GeForce RTX 3080 Ti Laptop GPU338.395% CI 332.4344.2

qwen2.5:0.5b-instruct-q4_K_M q4_k_m on ollama (ollama-http-api/auto)

Generation throughput for qwen2.5:0.5b-instruct-q4_K_M q4_k_m on ollama (ollama-http-api/auto)
1ollama-1789031359-bfb6262812th Gen Intel(R) Core(TM) i9-12900HNVIDIA GeForce RTX 3080 Ti Laptop GPU297.195% CI 291.7302.4
2, statistically tied with rank 1ollama-1789031428-39f82fdf12th Gen Intel(R) Core(TM) i9-12900HNVIDIA GeForce RTX 3080 Ti Laptop GPU293.295% CI 289.5296.9

qwen2.5:0.5b-instruct-q4_K_M on ollama (ollama-http-api/auto)

Single result — nothing to compare it against yet.

Generation throughput for qwen2.5:0.5b-instruct-q4_K_M on ollama (ollama-http-api/auto)
1ollama-178738893012th Gen Intel(R) Core(TM) i9-12900HNVIDIA GeForce RTX 3080 Ti Laptop GPU110.9