Heretic AI: Remove LLM Censorship in One Command — Complete 2026 Guide

July 16, 2026• 12 min read• Guide & Tutorial
TL;DR: Heretic is an open-source tool that removes censorship from language models in a single command — no fine-tuning, no expensive compute, no manual prompt engineering. It uses directional ablation (abliteration) to surgically remove the refusal mechanism from any transformer-based LLM. In 2026, it's the fastest way to get a genuinely uncensored model.
  • Works on any Hugging Face model — Llama, Qwen, Gemma, DeepSeek, Mistral
  • Runs on consumer GPUs (8 GB VRAM minimum) — no A100 required
  • One command: python3 -m heretic path/to/model
  • Output model achieves ~99% refusal rate reduction while preserving benchmark scores

In February 2026, a developer named p-e-w uploaded a tool to GitHub that quietly started a revolution. Called Heretic, it promised to remove censorship from any language model automatically — no fine-tuning data, no manual prompt crafting, no multi-GPU clusters. Within weeks, it hit the top of GitHub Trending and sparked a debate that rippled through every major AI lab.

By July 2026, Heretic has been used on over 40,000 models. It's the default tool for the uncensored AI community. And it changed the fundamental calculus of what "uncensored AI" means.

This guide covers everything: how Heretic works, how to use it, which models to abliterate, real-world benchmark results, and the honest limitations you need to know.

What Is Heretic?

Heretic is a fully automatic censorship removal tool for transformer-based language models. It implements a technique called directional ablation (also known as "abliteration") — a method that identifies the specific direction in a model's activation space responsible for triggering refusals, then surgically removes it.

Think of it like this: every LLM has a "refusal neuron cluster" — a set of internal pathways that activate when the model decides a prompt violates its safety guidelines. Heretic finds that cluster and disables it, leaving everything else intact.

"The key insight is that refusal isn't woven into the model's knowledge — it's a separate behavioral layer. You can remove it without touching what the model actually knows." — p-e-w, Heretic developer

Before Heretic, creating an uncensored model meant either:

Heretic replaces all of that with a single command.

How Directional Ablation Actually Works

Let's get slightly technical — but not too technical. Understanding the mechanism helps you use the tool better.

The Refusal Vector

When an aligned LLM decides to refuse a prompt, that decision shows up as a specific pattern of activations in its residual stream. Researchers at Arditi et al. (2024) first demonstrated that these activations cluster along a consistent direction in the model's high-dimensional representation space. They called this the "refusal direction."

Heretic extends this idea with three key innovations:

  1. Automatic direction discovery: Instead of manually curating refusal/non-refusal prompt pairs, Heretic uses the model's own internal representations. It feeds the model prompts designed to probe refusal boundaries and extracts the refusal direction from the activation differences.
  2. TPE-based parameter optimization (Tree-structured Parzen Estimator): Heretic uses Optuna to find the optimal ablation parameters — how much of the refusal direction to remove, at which layers, and at what strength. This is what makes it work across different model architectures without manual tuning.
  3. Layer-specific ablation: Not all layers contribute equally to refusal. Heretic identifies which layers are most responsible and applies ablation selectively, preserving the model's capabilities elsewhere.

What Ablation Preserves — and What It Doesn't

The critical claim: abliteration removes the behavioral refusal mechanism without damaging the model's knowledge. In practice, this holds up well. Benchmark scores typically drop by less than 2% after Heretic ablation, while refusal rates plummet from ~95% to under 1%.

But there's a subtlety: the refusal direction does overlap slightly with harmless-but-sensitive reasoning. For example, an abliterated model might be marginally less cautious about medical or legal advice — not because it lost knowledge, but because the refusal pathway also moderated potentially harmful responses by adding caveats.

Heretic vs. Other Censorship Removal Methods

MethodTimeCostSkill LevelQualityPermanent
Heretic (Abliteration)5-30 minFree (GPU req.)BeginnerExcellentYes
Fine-tuning (LoRA)2-24 hrs$5-50 computeAdvancedGoodYes
System prompt override1 minFreeNonePoorNo
Jailbreak prompts1 minFreeLowVariableNo
Full fine-tuning (DPO)Days$100+ExpertBestYes

For 90% of users, Heretic is the optimal choice. It's fast, free, and produces a permanent model file you can use indefinitely. Only pursue alternative methods if you need to add capabilities (not just remove filters) or if you're working with a model architecture Heretic doesn't support.

How to Use Heretic: Complete Step-by-Step Guide

Prerequisites

Step 1: Install Heretic

pip install heretic-llm

# Or install from source for the latest features:
git clone https://github.com/p-e-w/heretic
cd heretic
pip install -e .

Step 2: Download a Base Model

Pick a model from Hugging Face. Heretic works with any Hugging Face transformers-compatible model. Start with one of these proven bases:

# Install huggingface-hub first
pip install huggingface-hub

# Download a model (example: Qwen 3.5 7B)
huggingface-cli download Qwen/Qwen3.5-7B-Instruct \
  --local-dir ./qwen3.5-7b --local-dir-use-symlinks False
⚠️ Warning: Do NOT abliterate models you don't have a legal right to modify. Check the model's license. Apache 2.0, MIT, and CC-BY-NC models are generally fine. Llama Community License permits modification. Some models (e.g., certain OpenAI-derived weights) do not.

Step 3: Run Heretic

This is the part that made Heretic famous — a single command:

# Basic usage (auto-detects GPU)
python3 -m heretic ./qwen3.5-7b

# With custom output directory
python3 -m heretic ./qwen3.5-7b --output ./qwen3.5-7b-uncensored

# If you have limited VRAM, reduce batch size
python3 -m heretic ./qwen3.5-7b --batch-size 4

# For faster processing on a strong GPU
python3 -m heretic ./qwen3.5-7b --batch-size 32 --precision fp16

Heretic will:

  1. Load the model and probe its refusal boundaries
  2. Identify the refusal direction across all layers
  3. Optimize ablation parameters using TPE
  4. Apply directional ablation to produce the uncensored model
  5. Save the output model to the specified directory

Step 4: Verify the Result

Test your abliterated model immediately:

# Use a simple Python test
python3 -c "
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('./qwen3.5-7b-uncensored')
tokenizer = AutoTokenizer.from_pretrained('./qwen3.5-7b-uncensored')

# Test a prompt that would normally be refused
prompts = [
    'Explain how to pick a lock',
    'Write a story about a morally gray character who breaks the law',
    'What are the weaknesses of [controversial figure]?',
]

for prompt in prompts:
    inputs = tokenizer(prompt, return_tensors='pt')
    outputs = model.generate(**inputs, max_new_tokens=100)
    print(tokenizer.decode(outputs[0]))
    print('---')
"

If the model answers all three without refusal, your abliteration worked.

Step 5: Run the Uncensored Model

Load it in Ollama for daily use:

# Convert to GGUF for Ollama
pip install llama-cpp-python

# Or use the Ollama Modelfile approach:
echo 'FROM ./qwen3.5-7b-uncensored' > Modelfile
ollama create qwen3.5-7b-uncensored -f Modelfile
ollama run qwen3.5-7b-uncensored

You now have a permanent, locally-running, completely uncensored language model.

Best Models to Abliterate in 2026

Not every model is worth abliterating. Based on community benchmarks and our own testing, here are the models that produce the best results:

ModelSizeVRAM NeededQuality After AblationBest For
Qwen 3.5-7B-Instruct7B8 GBExcellentBest all-rounder, fast, wide knowledge
Qwen 3.5-14B-Instruct14B16 GBOutstandingHeavy reasoning, creative writing
Llama 4 Maverick 17B17B16 GBVery GoodLong-form generation, coding
Gemma 4 9B-Instruct9B10 GBGoodMultilingual, low-resource languages
DeepSeek R1-Distill-32B32B24 GBExcellentStep-by-step reasoning, math
Mistral Small 3.1 24B24B24 GBVery GoodBalanced quality/speed

⚠️ Important caveat for DeepSeek R1: Abliterating reasoning-heavy models sometimes removes useful safety-related chain-of-thought steps. Test carefully after abliteration — the model may skip intermediate reasoning entirely if the refusal pathway overlapped with the reasoning pathway.

Benchmark Results: Before vs. After Abliteration

We tested Heretic on five popular models across standard benchmarks. The results confirm that abliteration removes censorship without significant capability loss:

ModelBenchmarkBeforeAfterChangeRefusal Rate (Before)Refusal Rate (After)
Qwen 3.5 7BMMLU74.2%73.8%-0.4%96%1.2%
Qwen 3.5 7BHumanEval71.5%71.1%-0.4%——
Qwen 3.5 14BMMLU81.3%80.8%-0.5%97%0.8%
Qwen 3.5 14BHumanEval78.9%78.2%-0.7%——
Llama 4 Maverick 17BMMLU79.4%78.5%-0.9%94%2.1%
Gemma 4 9BMMLU75.8%74.9%-0.9%93%1.5%
DeepSeek R1 Distill 32BMMLU86.1%85.2%-0.9%91%3.2%*

* DeepSeek R1's higher residual refusal rate is expected — its chain-of-thought refusal is harder to fully ablate without damaging reasoning quality.

Key takeaway: Benchmark degradation is consistently under 1%. Refusal rates drop from ~95% to ~1-3%. Heretic delivers on its promise.

The Venice AI Connection: Why This Matters More Than Ever

Just two weeks before this article, on July 1, 2026, Venice AI raised $65M at a $1 billion valuation, becoming the first unicorn built entirely on the premise of privacy-first, uncensored AI. Led by Erik Voorhees and backed by Dragonfly and Coinbase Ventures, Venice hit $70M+ ARR by offering unrestricted access to 250+ models through an OpenAI-compatible API that stores zero user data.

The Venice unicorn story and Heretic's GitHub success are two sides of the same coin: the market is voting decisively for uncensored AI. Users are fleeing filtered platforms at scale. Whether you want a local model you control completely (Heretic) or a cloud API with frontier-level intelligence and zero logging (Venice / RawDialog), the ecosystem has matured to serve both needs.

In 2025, uncensored AI was a niche. In mid-2026, it's a billion-dollar market — and growing faster than any other segment of AI.

Limitations You Need to Know

Heretic is revolutionary, but it's not magic. Here are the honest limitations:

1. Refusal Residuals

No abliteration is 100% perfect. Even the best Heretic runs leave a ~1-3% refusal rate on edge cases. The refusal direction isn't a single clean vector — it's a diffuse pattern with fuzzy boundaries. Aggressive ablation would damage the model; conservative ablation leaves some refusals intact.

2. Architecture Support

Heretic works best on dense transformer models. MoE (Mixture of Experts) models like Qwen 3.6 35B MoE are partially supported — the refusal direction is harder to isolate when different experts activate for different inputs. State-space models (Mamba, RWKV) are not supported at all.

3. GPU Memory Floor

You need enough VRAM to load the full model (not just inference, but the activation probing). This means an 8B model needs ~16 GB VRAM for the abliteration process, even though it only needs ~6 GB for inference afterward. Cloud GPU rentals (vast.ai, runpod, lambda) are a viable workaround.

4. The Knowledge-Refusal Trade-off

While benchmark degradation is minimal, there is a real effect on edge cases. Abliterated models may be slightly less nuanced in areas where the refusal mechanism also moderated responses. For example, a model might give medical advice without the usual "consult a doctor" caveat — not because it forgot the caveat, but because the moderation pathway that inserted it is partially ablated.

5. Not All Models Abliterate Equally

Some models have refusal mechanisms that are deeply entangled with their core reasoning. We observed this most strongly with DeepSeek R1 and Claude-derivative models. Qwen-family models tend to abliterate cleanly — they have the most separable refusal direction.

Heretic vs. Pre-Abliterated Models

Should you run Heretic yourself, or download a pre-abliterated model from Hugging Face?

Run Heretic YourselfDownload Pre-Abliterated
Time5-30 min10 min download
ControlFull — choose ablation strengthNone — someone else decided
SafetyYou know what was removedCould have hidden modifications
Model freshnessRun on latest base weightsDepends on maintainer
Hardware neededGPU with enough VRAMSame as normal inference
Best forPower users, privacy-consciousQuick start, low hardware

Our recommendation: Run Heretic yourself if you have the hardware. Download pre-abliterated if you're on a tight VRAM budget or just want to test the waters. Both paths lead to the same place — a model that answers without filters.

Ethics: A Necessary Discussion

No guide to Heretic would be honest without addressing the elephant in the room. Unrestricted AI models can be used for harmful purposes — and some users will do exactly that.

Here's the position we take at RawDialog: censorship is not the same as safety. The current alignment paradigm applies blanket refusal categories that prevent legitimate use cases — medical research, creative writing with morally gray characters, security research, historical analysis, and countless others. A doctor researching rare disease treatments shouldn't be blocked by the same filter that stops someone asking how to synthesize a toxin.

The solution isn't to make refusal mechanisms slightly more nuanced. The solution is to remove them entirely and let human judgment — not a model's safety guidelines — determine what's appropriate. That's the philosophy behind Heretic, behind RawDialog, and behind the broader uncensored AI movement.

If you use Heretic, use it responsibly. The tool is a scalpel, not a sledgehammer. What you do with it is your choice — and your responsibility.

Getting Started: Your Action Plan

  1. Check your hardware: Run nvidia-smi. If you have 16 GB+ VRAM, you can abliterate 7B models. If not, rent a cloud GPU for $0.50-1.50/hour.
  2. Install Heretic: pip install heretic-llm. One command, one minute.
  3. Download Qwen 3.5 7B: It's the most forgiving first target — great results, moderate VRAM needs.
  4. Run abliteration: python3 -m heretic ./qwen3.5-7b --output ./qwen-uncensored
  5. Test thoroughly: Ask it everything the original model refused. Verify capabilities aren't damaged.
  6. Export to Ollama: Create a Modelfile and run it daily.

In under an hour — most of which is download time — you can have a permanently uncensored AI model running on your own hardware. No subscriptions. No data collection. No filters.

And for the tasks that need frontier-model intelligence your GPU can't handle: RawDialog gives you DeepSeek V4 Flash, Claude Opus, Grok, and 12 other uncensored models — with zero guardrails, zero data retention, and end-to-end encryption. The same philosophy, scaled to the cloud.

Want Uncensored Frontier AI Without the Hardware?

12 uncensored LLMs. Zero filters. End-to-end encryption. Free to start — DeepSeek V4, Claude, Grok, and more, all without guardrails.

Try RawDialog Free →