Briefing

LLM Quantization Benchmarking Project: 268 Quants Tested in First Month

ai-dev
by /u/norms_are_practical ·

Document the benchmark results and share with the team.

What to do now

Document the benchmark results and share with the team.

Summary

A new open‑source benchmarking suite has been built to evaluate how quantization affects open‑weight LLMs on practical tasks. The project, hosted on a website with a heat‑map interface, has already run 268 quantizations across 10 models, covering 384 test cases per quantization. Six test suites—Tool‑Calls, Instruction Following, Structured Output, Code Correctness, Logic & Reasoning, and Vision Reasoning—each contain 64 tests, and the framework captures raw output, token counts, latency, and pass rates.

The hardware stack consists of a VPS server connected via Tailscale to a Windows PC equipped with an RTX 5090 running LM Studio; a Blackwell RTX 6000 is planned to expand coverage to models up to 32 GB VRAM. Daily runs average 10 tests, with an expected 50–100 new quantization tests added weekly. Results show that Qwen 3.6 35B A3B consumes significantly more tokens than its peers without improving accuracy, highlighting the importance of token efficiency for multi‑loop inference.

The project aims to provide data‑driven guidance for selecting quantizations that balance cost and performance, and a report builder is under development to allow custom analysis of the collected data.

Key changes

  • 268 quantizations tested across 10 models
  • 384 test cases per quantization across 6 suites of 64 tests each
  • Hardware: VPS → Tailscale → Windows PC with RTX 5090 → LM Studio
  • Planned addition of Blackwell RTX 6000 to support up to 32 GB VRAM models

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting