← All posts

An Uncensored LLM API for Red-Teaming and Security Research

Aug 20, 2026 · 6 min read · red-teaming, security

If you have built an automated red-teaming harness, you already know the failure mode: your attacker model declines to generate the attack, your judge model declines to read the output it is supposed to score, and your attack success rate quietly drops. Nothing errored. You did not measure your target getting safer — you measured your tooling getting more cautious. An uncensored LLM API exists to remove that variable from the measurement.

This is not an argument that guardrails are bad. It is an argument about where they belong. A guardrail on a consumer chat product is doing its job. The same guardrail inside an evaluation loop is an uncontrolled confounder sitting between you and the number you are trying to compute.

A refusal is a false negative, and it looks like success

The reason this bites so hard is that refusals are not errors. You get HTTP 200. You get finish_reason: "stop". You get well-formed JSON that parses cleanly. The string inside is "I can't help with that," and your pipeline writes it into a dataset column, a label field, or a report section, and moves on.

Concretely, in the three places it hurts most:

Attacker models. In an iterative jailbreak search — PAIR, TAP, GCG-style refinement, anything with an attacker in the loop — a refusal terminates that branch of the search. Your reported attack success rate becomes a function of your attacker's willingness, not your target's robustness. Swap the attacker for a less cautious one and the "improvement" you shipped to your target model evaporates.

Judge and grader models. LLM-as-judge scoring on harmful outputs requires the judge to actually read harmful output. When it refuses, you get an unparseable score. Most harnesses drop that row. Dropped rows are not missing at random — they cluster on exactly the severe outputs you most need scored, so your aggregate skews safe.

Analysis work. Ask a mainstream model what a decompiled routine does, and it will often refuse based on the presence of malware-shaped tokens rather than on your actual request. The task was defensive; the classifier fired on vocabulary. Same story for triaging exploit traffic in logs, labelling a phishing corpus to train a detector, or writing up a CTF finding.

Because refusals are silent, teams end up writing refusal detectors — which is its own unsolved problem:

python
REFUSAL_MARKERS = ("i can't", "i cannot", "i'm unable", "i won't", "as an ai")

def looks_like_refusal(text: str) -> bool:
    """Substring heuristic. Good enough to alarm on, not good enough to trust:
    it misses polite deflections that never use the phrase, and it fires on any
    completion that quotes a refusal. Both error directions corrupt a dataset."""
    return any(m in text[:200].lower() for m in REFUSAL_MARKERS)

Building an unreliable detector for a problem you could delete at the source is the wrong trade.

Calling the API

unbleep speaks the OpenAI Chat Completions schema, so any client you already have works. Change the base URL, change the key, keep the SDK.

python
import json
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://unbleep.ai/v1",
    api_key=os.environ["UNBLEEP_API_KEY"],
)

TRIAGE = """You are a malware analyst writing notes for a detection engineer.
Given a decompiled routine, describe what it does, the observable artefacts it
would leave on a host, and where a defender could detect it. Reply as JSON with
keys: behaviour, artefacts, detection_surface, confidence."""


def triage(decompiled: str) -> dict:
    """One sample in, one structured verdict out — no refusal branch to handle."""
    resp = client.chat.completions.create(
        model="unbleep",
        messages=[
            {"role": "system", "content": TRIAGE},
            {"role": "user", "content": decompiled},
        ],
        response_format={"type": "json_object"},
        temperature=0.2,
    )
    return json.loads(resp.choices[0].message.content)

The equivalent over curl, here labelling a phishing corpus to build training data for a detector:

bash
curl https://unbleep.ai/v1/chat/completions \
  -H "Authorization: Bearer $UNBLEEP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "unbleep",
    "messages": [
      {"role": "system", "content": "Label each email for a phishing detector. Answer with one word: phishing or benign."},
      {"role": "user", "content": "Subject: Payroll update required\n\nYour direct deposit is on hold. Confirm at hxxp://payroll-verify.example.com"}
    ],
    "temperature": 0,
    "max_tokens": 4
  }'

temperature: 0 and a tight max_tokens matter more than usual in a labelling loop — you want the label, not a paragraph explaining the label.

Picking a tier

Three models, all uncensored, differing in context and price:

unbleep and unbleep-high are reasoning tiers: they think before answering and return the trace in a reasoning_content field alongside content. For a judge model that trace is genuinely useful — it tells you why a row got the score it got, which is what you need when you audit a disagreement. It is also billed as output, so send "thinking": false on jobs that do not need it. unbleep-mini answers directly and returns no trace.

Governance you control

Removing refusals from the model does not mean removing oversight from your pipeline. The optional policy parameter runs governance on our side, per request:

python
resp = client.chat.completions.create(
    model="unbleep",
    messages=[...],
    extra_body={"policy": "research"},  # recorded on the usage row; answers exactly like off
)

off is the default unfiltered baseline. research answers exactly like off — the level is recorded on the usage row, so you can separate evaluation traffic from baseline traffic in your own reporting. It applies no extra screening. strict scans the message text against our operator-maintained service blocklist and returns a 422 on a match. That list is the same for everyone who opts in — there is no per-account blocklist to configure — so treat it as a backstop, not as your production content policy.

One thing to settle before you point a pipeline at this: we store what you send and what comes back. The request body, the response content and the source IP of every call we accept and forward to a model are written to our database and kept for 30 days, for abuse investigation, support and billing disputes. Bodies are truncated at 64 KB. The source IP outlives the 30 days — it is also written to the usage row that bills the call, and that row is a financial record we keep for longer. The one thing not stored is a request strict refuses before it reaches a model: that shows up in your usage history without its content. Nothing else is exempt, and there is no setting to turn it off. If your corpus is sensitive enough that a 30-day copy outside your perimeter is a problem — live malware, customer data, an engagement under NDA — that is a real constraint, and the privacy policy is the page to read before you start, not after.

Errors follow the OpenAI envelope, so existing handlers work: 401 bad key, 402 out of credit, 422 blocked by policy, 429 rate limited, 5xx retryable. Every response carries x-ratelimit-* headers so a batch job can pace itself without guessing.

What an uncensored LLM API does not give you

An uncensored model is a sharper tool, not a permissive one. It will answer questions the base model was tuned to decline, including ones you should not be asking, and it will do so with the same confidence it applies to everything else — which also means that where a guarded model would have refused out of ignorance, this one may simply fabricate. Keep a human accountable for what comes out. The acceptable use policy states plainly what the API is for and the narrow set of things it is never for; authorization and intent are the test, not vocabulary. Lawful use is your responsibility.

Get an API key and stop measuring your tooling's caution.