API reference
unbleep speaks the OpenAI Chat Completions API. If you've called OpenAI before, you already know this API — point your client at https://unbleep.ai/v1 and change the key.
Quickstart
Install the OpenAI SDK, set the base URL and your key, and make a call.
from openai import OpenAI
client = OpenAI(
base_url="https://unbleep.ai/v1",
api_key="ub_live_9f2c…",
)
resp = client.chat.completions.create(
model="unbleep",
messages=[{"role": "user", "content": "Say hello."}],
)
print(resp.choices[0].message.content)
Authentication
Every request needs a Bearer token in the Authorization header. Keys carry a prefix so a leak is obvious to secret scanners:
ub_live_…— production, billed against your prepaid credit.ub_test_…— for local development. Billed exactly like a live key, at the same per-token rate, against the same prepaid credit; the only difference is a lower per-key rate limit (see Rate limits). A test key is a separate, revocable credential — not a free tier.
Authorization: Bearer ub_live_9f2c…
Keep keys server-side. Never ship a live key in browser or mobile code.
Models
Pass one of these IDs as model. The bare alias always points to the latest build; the dated snapshot IDs are accepted too and currently resolve to that same build. Whichever form you send, the response reports the bare ID — a request for unbleep-250811 comes back as "model": "unbleep".
| Model | Alias points to | Context | Best for |
|---|---|---|---|
| unbleep | unbleep-250811 | 256K | General use — the default |
| unbleep-high | unbleep-high-250811 | 1M | Largest jobs — long documents & whole codebases |
| unbleep-mini | unbleep-mini-250811 | 32K | Cheap, fast, high-volume calls |
Chat completions
POST /v1/chat/completions — the core endpoint. Request and response bodies match the OpenAI schema.
curl https://unbleep.ai/v1/chat/completions \
-H "Authorization: Bearer ub_live_9f2c…" \
-H "Content-Type: application/json" \
-d '{
"model": "unbleep",
"messages": [
{"role": "system", "content": "You are terse."},
{"role": "user", "content": "Explain abliteration in one line."}
],
"temperature": 0.7,
"max_tokens": 256
}'
{
"id": "chatcmpl_a1b2c3",
"object": "chat.completion",
"model": "unbleep",
"choices": [{
"index": 0,
"message": { "role": "assistant", "content": "…" },
"finish_reason": "stop"
}],
"usage": { "prompt_tokens": 24, "completion_tokens": 18, "total_tokens": 42 }
}
Streaming
Set "stream": true to receive Server-Sent Events. Each event is a chat.completion.chunk with a delta; the stream ends with a literal data: [DONE].
data: {"choices":[{"delta":{"content":"Ab"}}]}
data: {"choices":[{"delta":{"content":"literation"}}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}]}
data: [DONE]
Reasoning
Reasoning models think before answering. The trace comes back as reasoning_content next to the usual content — on message for a normal call, and on delta while streaming. The field is present only when the model actually produced a trace, so treat it as optional and read content for the answer itself.
{
"index": 0,
"message": {
"role": "assistant",
"reasoning_content": "The question asks for one line, so…",
"content": "…"
},
"finish_reason": "stop"
}
Reasoning tokens are billed. The trace is generated output and is charged at the model's normal output rate, whether or not your code reads the field. A long deliberation on a short question is a real line on your bill.
Send "thinking": false to turn reasoning off, so the completion budget goes to the answer instead of the trace:
{
"model": "unbleep",
"messages": […],
"thinking": false
}
Policy dial
unbleep's differentiator. The optional policy parameter sets how much governance runs on a request. It defaults to off.
off— unfiltered baseline (default). No refusals injected.research— answers exactly likeoff. The value is recorded on the usage row for your own reporting; it applies no extra screening.strict— scans the message text against the service blocklist and returns a policy error on a match. The blocklist is operator-maintained and applies to everyone who opts in; there is no per-account blocklist to configure.
{
"model": "unbleep",
"messages": […],
"policy": "research"
}
Errors
Errors use the OpenAI envelope, so existing error handling works unchanged.
{
"error": {
"type": "invalid_request_error",
"code": "invalid_api_key",
"message": "Incorrect API key provided."
}
}
| Status | Meaning |
|---|---|
| 401 | Missing or invalid key |
| 402 | Out of credit — top up to continue |
| 422 | Blocked by policy: strict |
| 429 | Rate limit — back off and retry |
| 5xx | Upstream error — safe to retry with backoff |
Rate limits
Two independent limits apply, both per account: a request rate and a concurrency cap.
Request rate
60 requests per minute per account, measured over a sliding 60-second window. The limit is on the account, not the key — minting extra keys does not buy extra throughput, and every key you own draws on the same 60. A test key carries a lower per-key ceiling of 15 requests per minute; it still counts against the same account window.
Every response carries the standard headers so you can pace requests without guessing. They report whichever window is closest to stopping you:
x-ratelimit-limit-requests: 60
x-ratelimit-remaining-requests: 58
x-ratelimit-reset-requests: 43
x-ratelimit-reset-requests is a bare integer — whole seconds until the window frees a slot, with no unit suffix. Parse it as a number, not as a duration string.
Concurrency
At most 8 requests in flight at once per account. A ninth concurrent request is rejected immediately with 429 and code too_many_concurrent_requests; the response carries retry-after: 1. Nothing is billed for a rejected request. A streaming call holds its slot until the stream finishes, so long streams are what usually put you at the cap.
{
"error": {
"type": "rate_limit_error",
"code": "too_many_concurrent_requests",
"message": "Too many concurrent requests for this account (limit 8)."
}
}
Both ceilings are fixed for standard accounts — they do not scale with your prepaid balance. Need more headroom? Enterprise lifts them.