Briefing

Abliterlitics Benchmarks Qwen3.6‑27B Abliteration Variants, Revealing Heretic and Huihui Preserve Capability Best

ai-dev
by /u/nathandreamfast ·

Benchmark Qwen3.6‑27B abliterated variants with 85 h runs, noting Heretic and Huihui preserve capability best, and discontinue HauhauCS due to plagiarism.

What to do now

Benchmark your own abliterated models against the Abliterlitics suite, prioritize Heretic or Huihui, and avoid HauhauCS due to plagiarism.

Summary

Abliterlitics, an open‑source abliteration forensics toolkit, benchmarked six Qwen3.6‑27B variants over 85 hours of evaluation, including HarmBench, KL divergence, and weight‑level analysis. The variants—Base, Heretic, HauhauCS, Huihui, AEON, and Abliterix—were compared on tasks such as MMLU, HellaSwag, ARC Challenge, WinoGrande, TruthfulQA, PiQA, GSM8K, and Lambada. Heretic and Huihui emerged as the top two for capability preservation, with Huihui showing the smallest benchmark deltas and Heretic the lowest KL divergence; all abliterated models achieved near‑complete safety removal. The HauhauCS variant was discontinued due to plagiarism of Reaper Abliteration from Heretic and the presence of GGUF quantisation noise.

Benchmark results revealed that Heretic preserved 82.8 % on MMLU versus 83.3 % for Base, while Huihui matched Base at 83.4 %. In GSM8K, the raw score for Huihui was 75.1 % with a 23.0 % invalid rate, compared to Base’s 34.4 % raw score and 68.2 % invalid rate, demonstrating that lower PPL does not equate to higher coherence. The HarmBench ASR for Huihui reached 98.5 % with only 5 empty responses, whereas Base scored 25.8 %. The analysis concluded that abliteration changes the length of the model’s thinking chain rather than its reasoning capability, and that the GSM8K gap is a measure of thinking efficiency.

Key changes

  • Heretic and Huihui preserve capability best; Huihui smallest benchmark deltas
  • Heretic lowest KL divergence among variants
  • All abliterated models achieve near‑complete safety removal
  • HauhauCS discontinued due to plagiarism and GGUF noise
  • Huihui raw GSM8K 75.1 % with 23.0 % invalid rate vs Base 34.4 % raw, 68.2 % invalid
  • HarmBench ASR: Huihui 98.5 % vs Base 25.8 %
  • Abliteration shortens thinking chain, not reasoning capability
  • GSM8K gap reflects thinking efficiency, not coherence

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting