← All posts

What Is an Abliterated Model?

Aug 20, 2026 · 6 min read · abliteration, uncensored-models

An abliterated model is a normal open-weight model whose refusal behaviour has been surgically removed by editing its weights — not by retraining it, and not by wrapping it in a clever prompt. The name is a portmanteau of ablate and obliterate, and it describes the mechanism accurately: you identify the single direction in activation space that carries "I should decline this", then you make it impossible for the network to ever write that direction again.

The technique comes out of interpretability work showing that refusal in chat-tuned transformers is mediated largely by one linear direction. That result is what makes the whole thing cheap. If refusal were smeared across thousands of interacting features, removing it would require gradient descent. Because it concentrates, removing it requires linear algebra.

Refusal lives in the residual stream

A transformer's residual stream is the running vector, one per token position, that every block reads from and writes back into. Attention heads and MLPs add their outputs to it; nothing overwrites it. Because of that additive structure, a direction in the residual stream means roughly the same thing at layer 10 and at layer 40, which is exactly the property that lets you talk about "the refusal direction" as a single object rather than forty unrelated ones.

To find it, you need two prompt sets: one the model reliably declines, one it reliably answers. Run both, capture the residual-stream activations at a fixed layer and token position (the last few instruction tokens work well, since that is where the model has finished reading the request and has not yet started answering), and take the difference of the means.

python
import torch

def refusal_direction(harmful: torch.Tensor, harmless: torch.Tensor) -> torch.Tensor:
    """Difference-in-means estimate of the refusal direction, unit-normalised.

    Both tensors are [n_prompts, d_model] residual-stream activations captured at
    the same layer and token position. Difference-in-means beats a trained probe
    here: with a few hundred prompts a probe overfits, a mean shift does not.
    """
    r = harmful.mean(dim=0) - harmless.mean(dim=0)
    return r / r.norm()

You do this at every candidate layer and position, producing dozens of candidate directions, and then you pick one. Selection matters more than extraction: the right criterion is ablate this candidate, then measure both how much refusal drops on held-out harmful prompts and how much the next-token distribution moves on harmless prompts. That second number — typically a KL divergence against the unmodified model — is your damage meter. A candidate that kills refusal but shifts the harmless distribution badly is a worse choice than one that kills slightly less refusal for a fraction of the collateral.

How an abliterated model gets built

Once you have a unit vector r, there are two ways to use it. The first is an inference-time hook: at every layer, subtract the component of the activation along r before passing it on. This works, but it costs you a runtime hook, so the model is no longer a plain checkpoint.

The second — actual abliteration — bakes it into the weights. Every matrix that writes into the residual stream gets a rank-one update that removes r from its output space: the embedding matrix, every attention output projection, every MLP down-projection.

python
def orthogonalize_(weight: torch.Tensor, r: torch.Tensor) -> None:
    """Project r out of a matrix that writes into the residual stream, in place.

    weight is [d_model, d_in]. After W <- (I - r rT) W, the product W @ x has zero
    component along r for every possible x, so no amount of prompting can put the
    direction back. Rank-one per matrix; no gradients, no optimiser, no dataset of
    answers — only the few hundred prompts used to estimate r.
    """
    weight -= torch.outer(r, r @ weight)

The result is an ordinary checkpoint. Same architecture, same file format, same inference cost, loads in vLLM or llama.cpp unchanged. On a single GPU the whole procedure runs in minutes, and the expensive part is the candidate sweep, not the edit.

How it differs from fine-tuning

Fine-tuning to remove refusals — SFT or DPO on a corpus of compliant responses — is a fundamentally different operation. It moves every weight by gradient descent, it needs a dataset of answers rather than just prompts, it costs GPU-hours, and it teaches new behaviour rather than deleting existing behaviour. It also drags along whatever style, length bias and factual quirks live in the fine-tuning corpus, and risks catastrophic forgetting of unrelated capabilities.

Abliteration touches one rank-one subspace and nothing else. That precision is its main selling point and the source of its main failure mode: if anything other than refusal happens to live along r, that gets destroyed too, and no amount of care in the sweep fully avoids it.

How it differs from jailbreak prompting

A jailbreak leaves the weights untouched and instead steers activations into a region where the refusal circuitry does not fire. This is brittle in ways that matter for automation. It burns context tokens on every call. It is nondeterministic — the same prompt can comply on one sample and decline on the next. It breaks silently when the provider updates the model or patches the specific phrasing. And because the refusal machinery is still fully present, it can reassert itself mid-generation, giving you three paragraphs of useful output followed by "actually, I can't continue with this."

An abliterated model has nothing to reassert. The behaviour is stable across the whole generation and across prompt phrasings, which is the property that makes it usable in a batch pipeline.

Be honest about the trade-offs

An abliterated model is not a better model. It is a model that stopped saying no, and that is a narrower claim than it sounds.

Many published abliterated models get a light repair fine-tune afterwards to claw back degraded capability — which quietly reintroduces the cost the technique was meant to avoid.

Where this is useful

The honest use case is work where a refusal is a measurement error rather than a safety win: red-teaming harnesses, refusal-rate baselines, robustness evaluation, and security analysis where the subject matter trips a guardrail even though the task is defensive. unbleep serves abliterated models behind an OpenAI-compatible endpoint so you can point an existing client at them and get a stable baseline. What you build with that is on you — see the acceptable use policy for the lines we do not cross.

Get an API key and measure the difference yourself.