Briefing

Qwen3.6-27B Uncensored Heretic v2 Native MTP Preserved Models Released

ai-dev
by /u/LLMFan46 · Llama

Download the Qwen3.6‑27B‑uncensored‑heretic‑v2 Native MTP Preserved GGUF models from HuggingFace and test them with your inference engine.

What to do now

Download the models from HuggingFace and integrate them into your local inference pipeline.

Summary

Qwen3.6-27B-uncensored-heretic-v2 Native MTP Preserved models have been released on HuggingFace. The release includes several GGUF variants: standard, NVFP4, NVFP4-MLP-Only, and others, all retaining the full 15 Multi‑Token Prediction (MTP) capabilities. A benchmark is provided to compare performance across quantization settings. The models are available under the llmfan46 organization and include a benchmark that shows inference speed and memory usage. The main model is a 27‑billion‑parameter Qwen3.6 variant, with the "heretic v2" designation indicating a second iteration of the uncensored model with updated training data and safety mitigations. The Native MTP Preserved tag means the models can be used directly with inference engines that support MTP, such as llama.cpp. Users can download the models from the provided HuggingFace URLs and integrate them into their pipelines.

Key changes

  • Release of Qwen3.6‑27B‑uncensored‑heretic‑v2 Native MTP Preserved models
  • Multiple GGUF variants: standard, NVFP4, NVFP4‑MLP‑Only, etc.
  • All variants retain full 15 MTPs
  • Benchmark included for performance comparison
  • Models hosted under llmfan46 on HuggingFace
  • Uncensored heretic v2 indicates updated training and safety mitigations
  • Native MTP Preserved allows direct use with MTP‑capable inference engines

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting