Local AI Benchmarks

Median decode speed on home hardware · tokens per second

Runs
–
benchmark measurements
Models
–
distinct model names
Best prose
– tok/s
fastest prose run
Best code
– tok/s
fastest code run

Throughput per run prose tok/s, fastest first · hover a bar for details

Prose vs code one row per run · blue = prose, teal = code · hover for details

By memory class best prose and code per class

By inference engine best prose tok/s per engine

About this data

  • What is measured: median decode speed in tokens per second (tok/s) for prose and code prompts.
  • Hardware: consumer and prosumer GPUs, from 16 GB to 256 GB of memory.
  • Recipe: quantization, inference engine and context length used for each run.
  • Source:

All runs

All benchmark runs with machine, model, recipe, memory class and decode speed in tokens per second
Machine Model Recipe Class Prose tok/s Code tok/s Δ code−prose