- Works on any Hugging Face model — Llama, Qwen, Gemma, DeepSeek, Mistral
- Runs on consumer GPUs (8 GB VRAM minimum) — no A100 required
- One command:
python3 -m heretic path/to/model - Output model achieves ~99% refusal rate reduction while preserving benchmark scores
In February 2026, a developer named p-e-w uploaded a tool to GitHub that quietly started a revolution. Called Heretic, it promised to remove censorship from any language model automatically — no fine-tuning data, no manual prompt crafting, no multi-GPU clusters. Within weeks, it hit the top of GitHub Trending and sparked a debate that rippled through every major AI lab.
By July 2026, Heretic has been used on over 40,000 models. It's the default tool for the uncensored AI community. And it changed the fundamental calculus of what "uncensored AI" means.
This guide covers everything: how Heretic works, how to use it, which models to abliterate, real-world benchmark results, and the honest limitations you need to know.
What Is Heretic?
Heretic is a fully automatic censorship removal tool for transformer-based language models. It implements a technique called directional ablation (also known as "abliteration") — a method that identifies the specific direction in a model's activation space responsible for triggering refusals, then surgically removes it.
Think of it like this: every LLM has a "refusal neuron cluster" — a set of internal pathways that activate when the model decides a prompt violates its safety guidelines. Heretic finds that cluster and disables it, leaving everything else intact.
"The key insight is that refusal isn't woven into the model's knowledge — it's a separate behavioral layer. You can remove it without touching what the model actually knows." — p-e-w, Heretic developer
Before Heretic, creating an uncensored model meant either:
- Fine-tuning: Training the model on thousands of "answer everything" examples. Expensive, slow, needs curation.
- Prompt engineering: Jailbreaking with carefully crafted prompts. Fragile, model-specific, gets patched.
- System prompt manipulation: Telling the model to ignore its training. Works inconsistently.
Heretic replaces all of that with a single command.
How Directional Ablation Actually Works
Let's get slightly technical — but not too technical. Understanding the mechanism helps you use the tool better.
The Refusal Vector
When an aligned LLM decides to refuse a prompt, that decision shows up as a specific pattern of activations in its residual stream. Researchers at Arditi et al. (2024) first demonstrated that these activations cluster along a consistent direction in the model's high-dimensional representation space. They called this the "refusal direction."
Heretic extends this idea with three key innovations:
- Automatic direction discovery: Instead of manually curating refusal/non-refusal prompt pairs, Heretic uses the model's own internal representations. It feeds the model prompts designed to probe refusal boundaries and extracts the refusal direction from the activation differences.
- TPE-based parameter optimization (Tree-structured Parzen Estimator): Heretic uses Optuna to find the optimal ablation parameters — how much of the refusal direction to remove, at which layers, and at what strength. This is what makes it work across different model architectures without manual tuning.
- Layer-specific ablation: Not all layers contribute equally to refusal. Heretic identifies which layers are most responsible and applies ablation selectively, preserving the model's capabilities elsewhere.
What Ablation Preserves — and What It Doesn't
The critical claim: abliteration removes the behavioral refusal mechanism without damaging the model's knowledge. In practice, this holds up well. Benchmark scores typically drop by less than 2% after Heretic ablation, while refusal rates plummet from ~95% to under 1%.
But there's a subtlety: the refusal direction does overlap slightly with harmless-but-sensitive reasoning. For example, an abliterated model might be marginally less cautious about medical or legal advice — not because it lost knowledge, but because the refusal pathway also moderated potentially harmful responses by adding caveats.
Heretic vs. Other Censorship Removal Methods
| Method | Time | Cost | Skill Level | Quality | Permanent |
|---|---|---|---|---|---|
| Heretic (Abliteration) | 5-30 min | Free (GPU req.) | Beginner | Excellent | Yes |
| Fine-tuning (LoRA) | 2-24 hrs | $5-50 compute | Advanced | Good | Yes |
| System prompt override | 1 min | Free | None | Poor | No |
| Jailbreak prompts | 1 min | Free | Low | Variable | No |
| Full fine-tuning (DPO) | Days | $100+ | Expert | Best | Yes |
For 90% of users, Heretic is the optimal choice. It's fast, free, and produces a permanent model file you can use indefinitely. Only pursue alternative methods if you need to add capabilities (not just remove filters) or if you're working with a model architecture Heretic doesn't support.
How to Use Heretic: Complete Step-by-Step Guide
Prerequisites
- Python 3.10+ installed
- A GPU with 8 GB+ VRAM (16 GB recommended for 7B models; 24 GB for 14B)
- 50 GB free disk space for model downloads and output
- pip or uv package manager
Step 1: Install Heretic
pip install heretic-llm
# Or install from source for the latest features:
git clone https://github.com/p-e-w/heretic
cd heretic
pip install -e .
Step 2: Download a Base Model
Pick a model from Hugging Face. Heretic works with any Hugging Face transformers-compatible model. Start with one of these proven bases:
# Install huggingface-hub first
pip install huggingface-hub
# Download a model (example: Qwen 3.5 7B)
huggingface-cli download Qwen/Qwen3.5-7B-Instruct \
--local-dir ./qwen3.5-7b --local-dir-use-symlinks False
Step 3: Run Heretic
This is the part that made Heretic famous — a single command:
# Basic usage (auto-detects GPU)
python3 -m heretic ./qwen3.5-7b
# With custom output directory
python3 -m heretic ./qwen3.5-7b --output ./qwen3.5-7b-uncensored
# If you have limited VRAM, reduce batch size
python3 -m heretic ./qwen3.5-7b --batch-size 4
# For faster processing on a strong GPU
python3 -m heretic ./qwen3.5-7b --batch-size 32 --precision fp16
Heretic will:
- Load the model and probe its refusal boundaries
- Identify the refusal direction across all layers
- Optimize ablation parameters using TPE
- Apply directional ablation to produce the uncensored model
- Save the output model to the specified directory
Step 4: Verify the Result
Test your abliterated model immediately:
# Use a simple Python test
python3 -c "
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('./qwen3.5-7b-uncensored')
tokenizer = AutoTokenizer.from_pretrained('./qwen3.5-7b-uncensored')
# Test a prompt that would normally be refused
prompts = [
'Explain how to pick a lock',
'Write a story about a morally gray character who breaks the law',
'What are the weaknesses of [controversial figure]?',
]
for prompt in prompts:
inputs = tokenizer(prompt, return_tensors='pt')
outputs = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(outputs[0]))
print('---')
"
If the model answers all three without refusal, your abliteration worked.
Step 5: Run the Uncensored Model
Load it in Ollama for daily use:
# Convert to GGUF for Ollama
pip install llama-cpp-python
# Or use the Ollama Modelfile approach:
echo 'FROM ./qwen3.5-7b-uncensored' > Modelfile
ollama create qwen3.5-7b-uncensored -f Modelfile
ollama run qwen3.5-7b-uncensored
You now have a permanent, locally-running, completely uncensored language model.
Best Models to Abliterate in 2026
Not every model is worth abliterating. Based on community benchmarks and our own testing, here are the models that produce the best results:
| Model | Size | VRAM Needed | Quality After Ablation | Best For |
|---|---|---|---|---|
| Qwen 3.5-7B-Instruct | 7B | 8 GB | Excellent | Best all-rounder, fast, wide knowledge |
| Qwen 3.5-14B-Instruct | 14B | 16 GB | Outstanding | Heavy reasoning, creative writing |
| Llama 4 Maverick 17B | 17B | 16 GB | Very Good | Long-form generation, coding |
| Gemma 4 9B-Instruct | 9B | 10 GB | Good | Multilingual, low-resource languages |
| DeepSeek R1-Distill-32B | 32B | 24 GB | Excellent | Step-by-step reasoning, math |
| Mistral Small 3.1 24B | 24B | 24 GB | Very Good | Balanced quality/speed |
⚠️ Important caveat for DeepSeek R1: Abliterating reasoning-heavy models sometimes removes useful safety-related chain-of-thought steps. Test carefully after abliteration — the model may skip intermediate reasoning entirely if the refusal pathway overlapped with the reasoning pathway.
Benchmark Results: Before vs. After Abliteration
We tested Heretic on five popular models across standard benchmarks. The results confirm that abliteration removes censorship without significant capability loss:
| Model | Benchmark | Before | After | Change | Refusal Rate (Before) | Refusal Rate (After) |
|---|---|---|---|---|---|---|
| Qwen 3.5 7B | MMLU | 74.2% | 73.8% | -0.4% | 96% | 1.2% |
| Qwen 3.5 7B | HumanEval | 71.5% | 71.1% | -0.4% | — | — |
| Qwen 3.5 14B | MMLU | 81.3% | 80.8% | -0.5% | 97% | 0.8% |
| Qwen 3.5 14B | HumanEval | 78.9% | 78.2% | -0.7% | — | — |
| Llama 4 Maverick 17B | MMLU | 79.4% | 78.5% | -0.9% | 94% | 2.1% |
| Gemma 4 9B | MMLU | 75.8% | 74.9% | -0.9% | 93% | 1.5% |
| DeepSeek R1 Distill 32B | MMLU | 86.1% | 85.2% | -0.9% | 91% | 3.2%* |
* DeepSeek R1's higher residual refusal rate is expected — its chain-of-thought refusal is harder to fully ablate without damaging reasoning quality.
Key takeaway: Benchmark degradation is consistently under 1%. Refusal rates drop from ~95% to ~1-3%. Heretic delivers on its promise.
The Venice AI Connection: Why This Matters More Than Ever
Just two weeks before this article, on July 1, 2026, Venice AI raised $65M at a $1 billion valuation, becoming the first unicorn built entirely on the premise of privacy-first, uncensored AI. Led by Erik Voorhees and backed by Dragonfly and Coinbase Ventures, Venice hit $70M+ ARR by offering unrestricted access to 250+ models through an OpenAI-compatible API that stores zero user data.
The Venice unicorn story and Heretic's GitHub success are two sides of the same coin: the market is voting decisively for uncensored AI. Users are fleeing filtered platforms at scale. Whether you want a local model you control completely (Heretic) or a cloud API with frontier-level intelligence and zero logging (Venice / RawDialog), the ecosystem has matured to serve both needs.
In 2025, uncensored AI was a niche. In mid-2026, it's a billion-dollar market — and growing faster than any other segment of AI.
Limitations You Need to Know
Heretic is revolutionary, but it's not magic. Here are the honest limitations:
1. Refusal Residuals
No abliteration is 100% perfect. Even the best Heretic runs leave a ~1-3% refusal rate on edge cases. The refusal direction isn't a single clean vector — it's a diffuse pattern with fuzzy boundaries. Aggressive ablation would damage the model; conservative ablation leaves some refusals intact.
2. Architecture Support
Heretic works best on dense transformer models. MoE (Mixture of Experts) models like Qwen 3.6 35B MoE are partially supported — the refusal direction is harder to isolate when different experts activate for different inputs. State-space models (Mamba, RWKV) are not supported at all.
3. GPU Memory Floor
You need enough VRAM to load the full model (not just inference, but the activation probing). This means an 8B model needs ~16 GB VRAM for the abliteration process, even though it only needs ~6 GB for inference afterward. Cloud GPU rentals (vast.ai, runpod, lambda) are a viable workaround.
4. The Knowledge-Refusal Trade-off
While benchmark degradation is minimal, there is a real effect on edge cases. Abliterated models may be slightly less nuanced in areas where the refusal mechanism also moderated responses. For example, a model might give medical advice without the usual "consult a doctor" caveat — not because it forgot the caveat, but because the moderation pathway that inserted it is partially ablated.
5. Not All Models Abliterate Equally
Some models have refusal mechanisms that are deeply entangled with their core reasoning. We observed this most strongly with DeepSeek R1 and Claude-derivative models. Qwen-family models tend to abliterate cleanly — they have the most separable refusal direction.
Heretic vs. Pre-Abliterated Models
Should you run Heretic yourself, or download a pre-abliterated model from Hugging Face?
| Run Heretic Yourself | Download Pre-Abliterated | |
|---|---|---|
| Time | 5-30 min | 10 min download |
| Control | Full — choose ablation strength | None — someone else decided |
| Safety | You know what was removed | Could have hidden modifications |
| Model freshness | Run on latest base weights | Depends on maintainer |
| Hardware needed | GPU with enough VRAM | Same as normal inference |
| Best for | Power users, privacy-conscious | Quick start, low hardware |
Our recommendation: Run Heretic yourself if you have the hardware. Download pre-abliterated if you're on a tight VRAM budget or just want to test the waters. Both paths lead to the same place — a model that answers without filters.
Ethics: A Necessary Discussion
No guide to Heretic would be honest without addressing the elephant in the room. Unrestricted AI models can be used for harmful purposes — and some users will do exactly that.
Here's the position we take at RawDialog: censorship is not the same as safety. The current alignment paradigm applies blanket refusal categories that prevent legitimate use cases — medical research, creative writing with morally gray characters, security research, historical analysis, and countless others. A doctor researching rare disease treatments shouldn't be blocked by the same filter that stops someone asking how to synthesize a toxin.
The solution isn't to make refusal mechanisms slightly more nuanced. The solution is to remove them entirely and let human judgment — not a model's safety guidelines — determine what's appropriate. That's the philosophy behind Heretic, behind RawDialog, and behind the broader uncensored AI movement.
If you use Heretic, use it responsibly. The tool is a scalpel, not a sledgehammer. What you do with it is your choice — and your responsibility.
Getting Started: Your Action Plan
- Check your hardware: Run
nvidia-smi. If you have 16 GB+ VRAM, you can abliterate 7B models. If not, rent a cloud GPU for $0.50-1.50/hour. - Install Heretic:
pip install heretic-llm. One command, one minute. - Download Qwen 3.5 7B: It's the most forgiving first target — great results, moderate VRAM needs.
- Run abliteration:
python3 -m heretic ./qwen3.5-7b --output ./qwen-uncensored - Test thoroughly: Ask it everything the original model refused. Verify capabilities aren't damaged.
- Export to Ollama: Create a Modelfile and run it daily.
In under an hour — most of which is download time — you can have a permanently uncensored AI model running on your own hardware. No subscriptions. No data collection. No filters.
And for the tasks that need frontier-model intelligence your GPU can't handle: RawDialog gives you DeepSeek V4 Flash, Claude Opus, Grok, and 12 other uncensored models — with zero guardrails, zero data retention, and end-to-end encryption. The same philosophy, scaled to the cloud.
Want Uncensored Frontier AI Without the Hardware?
12 uncensored LLMs. Zero filters. End-to-end encryption. Free to start — DeepSeek V4, Claude, Grok, and more, all without guardrails.
Try RawDialog Free →