API reference

unbleep speaks the OpenAI Chat Completions API. If you've called OpenAI before, you already know this API — point your client at https://unbleep.ai/v1 and change the key.

Quickstart

Install the OpenAI SDK, set the base URL and your key, and make a call.

python
from openai import OpenAI

client = OpenAI(
    base_url="https://unbleep.ai/v1",
    api_key="ub_live_9f2c…",
)

resp = client.chat.completions.create(
    model="unbleep",
    messages=[{"role": "user", "content": "Say hello."}],
)
print(resp.choices[0].message.content)

Authentication

Every request needs a Bearer token in the Authorization header. Keys carry a prefix so a leak is obvious to secret scanners:

header
Authorization: Bearer ub_live_9f2c…

Keep keys server-side. Never ship a live key in browser or mobile code.

Models

Pass one of these IDs as model. The bare alias always points to the latest build; the dated snapshot IDs are accepted too and currently resolve to that same build. Whichever form you send, the response reports the bare ID — a request for unbleep-250811 comes back as "model": "unbleep".

ModelAlias points toContextBest for
unbleepunbleep-250811256KGeneral use — the default
unbleep-highunbleep-high-2508111MLargest jobs — long documents & whole codebases
unbleep-miniunbleep-mini-25081132KCheap, fast, high-volume calls

Chat completions

POST /v1/chat/completions — the core endpoint. Request and response bodies match the OpenAI schema.

curl · request
curl https://unbleep.ai/v1/chat/completions \
  -H "Authorization: Bearer ub_live_9f2c…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "unbleep",
    "messages": [
      {"role": "system", "content": "You are terse."},
      {"role": "user", "content": "Explain abliteration in one line."}
    ],
    "temperature": 0.7,
    "max_tokens": 256
  }'
json · response
{
  "id": "chatcmpl_a1b2c3",
  "object": "chat.completion",
  "model": "unbleep",
  "choices": [{
    "index": 0,
    "message": { "role": "assistant", "content": "…" },
    "finish_reason": "stop"
  }],
  "usage": { "prompt_tokens": 24, "completion_tokens": 18, "total_tokens": 42 }
}

Streaming

Set "stream": true to receive Server-Sent Events. Each event is a chat.completion.chunk with a delta; the stream ends with a literal data: [DONE].

event stream
data: {"choices":[{"delta":{"content":"Ab"}}]}
data: {"choices":[{"delta":{"content":"literation"}}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}]}
data: [DONE]

Reasoning

Reasoning models think before answering. The trace comes back as reasoning_content next to the usual content — on message for a normal call, and on delta while streaming. The field is present only when the model actually produced a trace, so treat it as optional and read content for the answer itself.

json · response fragment
{
  "index": 0,
  "message": {
    "role": "assistant",
    "reasoning_content": "The question asks for one line, so…",
    "content": "…"
  },
  "finish_reason": "stop"
}

Reasoning tokens are billed. The trace is generated output and is charged at the model's normal output rate, whether or not your code reads the field. A long deliberation on a short question is a real line on your bill.

Send "thinking": false to turn reasoning off, so the completion budget goes to the answer instead of the trace:

json · request fragment
{
  "model": "unbleep",
  "messages": […],
  "thinking": false
}

Policy dial

unbleep's differentiator. The optional policy parameter sets how much governance runs on a request. It defaults to off.

json · request fragment
{
  "model": "unbleep",
  "messages": […],
  "policy": "research"
}

Errors

Errors use the OpenAI envelope, so existing error handling works unchanged.

json · 401
{
  "error": {
    "type": "invalid_request_error",
    "code": "invalid_api_key",
    "message": "Incorrect API key provided."
  }
}
StatusMeaning
401Missing or invalid key
402Out of credit — top up to continue
422Blocked by policy: strict
429Rate limit — back off and retry
5xxUpstream error — safe to retry with backoff

Rate limits

Two independent limits apply, both per account: a request rate and a concurrency cap.

Request rate

60 requests per minute per account, measured over a sliding 60-second window. The limit is on the account, not the key — minting extra keys does not buy extra throughput, and every key you own draws on the same 60. A test key carries a lower per-key ceiling of 15 requests per minute; it still counts against the same account window.

Every response carries the standard headers so you can pace requests without guessing. They report whichever window is closest to stopping you:

response headers
x-ratelimit-limit-requests: 60
x-ratelimit-remaining-requests: 58
x-ratelimit-reset-requests: 43

x-ratelimit-reset-requests is a bare integer — whole seconds until the window frees a slot, with no unit suffix. Parse it as a number, not as a duration string.

Concurrency

At most 8 requests in flight at once per account. A ninth concurrent request is rejected immediately with 429 and code too_many_concurrent_requests; the response carries retry-after: 1. Nothing is billed for a rejected request. A streaming call holds its slot until the stream finishes, so long streams are what usually put you at the cap.

json · 429
{
  "error": {
    "type": "rate_limit_error",
    "code": "too_many_concurrent_requests",
    "message": "Too many concurrent requests for this account (limit 8)."
  }
}

Both ceilings are fixed for standard accounts — they do not scale with your prepaid balance. Need more headroom? Enterprise lifts them.