sondahub / Anthropic sandbox
A Claude API mock with a scripted model: the same request, the same answer
The Claude API, answered by sondahub: Messages with streaming, tool use, extended and adaptive thinking with signed blocks, structured outputs, stop sequences and prompt caching, plus count_tokens, Message Batches and Models — every request checked the way the API checks it. The model is a script, so tests are deterministic, instant and free, and a phrase in the message picks the outcome: a tool call, a refusal, a 529, a stream that breaks halfway.
An independent imitation for testing. Not affiliated with, or endorsed by, Anthropic. No model runs and nothing is billed.
Connect
- Instead of
- https://api.anthropic.com
- Use
- https://api.sondahub.com
- API key
- x-api-key: any key that starts with sk-ant- — sk-ant-sondahub
- Version
- anthropic-version: 2023-06-01
- Also at
- https://api.sondahub.com/sandbox/anthropic
- OpenAPI 3
- https://api.sondahub.com/sandbox/anthropic/openapi.json
The quickest way: set ANTHROPIC_BASE_URL=https://api.sondahub.com and ANTHROPIC_API_KEY=sk-ant-sondahub, and code that uses the official SDK talks to the sandbox unchanged. That is enough for anything stateless — messages, streams, tool loops, thinking, count_tokens.
For state — batches and the prompt cache — the SDK also has to carry X-Sondahub-Session from each answer to the next request, as below; there is no database, so the token is where those live.
Python (anthropic)
from anthropic import Anthropic, DefaultHttpxClient
# carry the sandbox session (batches, the prompt cache) from each answer to the next request
session = {}
def put(request):
request.headers.update(session)
def keep(response):
if 'X-Sondahub-Session' in response.headers:
session['X-Sondahub-Session'] = response.headers['X-Sondahub-Session']
client = Anthropic(
base_url='https://api.sondahub.com',
api_key='sk-ant-sondahub',
http_client=DefaultHttpxClient(event_hooks={'request': [put], 'response': [keep]}),
)
message = client.messages.create(
model='claude-opus-5-5',
max_tokens=1024,
messages=[{'role': 'user', 'content': 'Say this is a test'}],
)
print(message.content[0].text) # This is a test.
TypeScript (@anthropic-ai/sdk)
import Anthropic from '@anthropic-ai/sdk'
// carry the sandbox session (batches, the prompt cache) from each answer to the next request
let session = null
const client = new Anthropic({
baseURL: 'https://api.sondahub.com',
apiKey: 'sk-ant-sondahub',
fetch: async (url, init = {}) => {
const headers = new Headers(init.headers)
if (session) headers.set('X-Sondahub-Session', session)
const res = await fetch(url, { ...init, headers })
session = res.headers.get('X-Sondahub-Session') ?? session
return res
},
})
const stream = client.messages.stream({ model: 'claude-sonnet-5-5', max_tokens: 1024, messages: [{ role: 'user', content: '[[tokens:200]]' }] })
stream.on('text', (text) => process.stdout.write(text))
const final = await stream.finalMessage()
Import it. In Sonda: Import → From a URL, paste the OpenAPI address. Every operation comes with its anthropic-version header and example bodies — a tool call, extended thinking, prompt caching, a batch, an error on purpose. Set the auth: API key sk-ant-sondahub in x-api-key.
Try it here
curl https://api.sondahub.com/v1/messages \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello, Claude"}]
}'
curl -N https://api.sondahub.com/v1/messages \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"max_tokens": 1024,
"stream": true,
"messages": [{"role": "user", "content": "[[tokens:60]] Tell me a story."}]
}'
curl https://api.sondahub.com/v1/messages \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"tools": [
{
"name": "get_weather",
"description": "Get the current weather in a given location",
"input_schema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA"
}
},
"required": ["location"]
}
}
],
"messages": [{"role": "user", "content": "What'\''s the weather like in San Francisco?"}]
}'
curl https://api.sondahub.com/v1/messages \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5-5",
"max_tokens": 16000,
"thinking": {"type": "enabled", "budget_tokens": 10000},
"messages": [{"role": "user", "content": "What is (17 * 23) + 4?"}]
}'
DOC=$(printf 'You answer questions about the sondahub handbook. The handbook says: every API is a mock, every write is kept in a session token the client carries, and nothing is stored on the server. %.0s' $(seq 30))
curl https://api.sondahub.com/v1/messages \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"max_tokens": 1024,
"system": [{"type": "text", "text": "'"$DOC"'", "cache_control": {"type": "ephemeral"}}],
"messages": [{"role": "user", "content": "Where is my data kept?"}]
}'
curl https://api.sondahub.com/v1/messages/count_tokens \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5-5",
"messages": [{"role": "user", "content": "Hello, world"}]
}'
curl https://api.sondahub.com/v1/messages/batches \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"requests": [
{
"custom_id": "first",
"params": {
"model": "claude-opus-5-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Say batch one"}]
}
},
{
"custom_id": "second",
"params": {
"model": "claude-haiku-4-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "What is 6 * 7?"}]
}
}
]
}'
curl https://api.sondahub.com/v1/messages \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"tools": [
{
"name": "get_weather",
"description": "Get the current weather in a given location",
"input_schema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA"
}
},
"required": ["location"]
}
}
],
"messages": [
{"role": "user", "content": "What'\''s the weather in Oslo?"},
{
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "toolu_01Sx3sondahubExample9QzT",
"name": "get_weather",
"input": {"location": "Oslo"}
}
]
},
{"role": "user", "content": "Never mind."}
]
}'
curl https://api.sondahub.com/v1/messages \
-H "x-api-key: sk-ant-sondahub" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "[[error:529]] hello"}]
}'
From the shell, keep the token yourself: curl -i shows X-Sondahub-Session; send it back with -H "X-Sondahub-Session: …". How sessions work.
The scripted model
No model runs here. A script answers, and it reads the conversation the way a test wants a model to: the same request always gets the same answer, instantly, for nothing. It still does a few useful things on its own — “Say this is a test” answers This is a test., an arithmetic question gets its answer, a message that names a tool (by a word of its name: “weather” for get_weather) calls it with arguments made from its schema and your words (the city after “in”, the numbers in order), a turn that brings tool results back gets a reply that reports them, and a JSON schema gets an object that fits it, required fields, enums and formats included. Anything else gets a short note saying it is Claude’s stand-in and how to steer it:
| Put this in the message | What comes back |
|---|---|
| [[tokens:N]] | A reply exactly N tokens long (up to 32,000), cut by the output limit like any reply — the way to test a long stream, a full buffer or stop_reason: "max_tokens". |
| [[echo]] | The message back, as written. |
| [[tool]] | Call the first tool. [[tool:name]] calls that one; [[tools:all]] calls every one at once (unless disable_parallel_tool_use is set). |
| [[refuse]] | A refusal: stop_reason: "refusal" with stop_details; [[refuse:cyber]] sets its category (cyber, bio, frontier_llm, reasoning_extraction, general_harms). |
| [[error:CODE]] | That error in Anthropic’s shape: any status from 400 to 599 (429 rate_limit_error, 529 overloaded_error…), or quota (billing_error) or context (prompt is too long). |
| [[flaky:CODE]] | That error on the first try only: the SDK retries (it sends x-stainless-retry-count) and the retry gets the answer — the way to watch your retry settings work. Without that header, every try fails. |
| [[break]] | A stream that dies halfway with an error event (overloaded_error). A plain request gets a 529. |
| [[slow]] | A stream paced like a slow model: over a second before the first token, then 140 ms between deltas. |
With tools and tool_choice: auto a short text block comes before the tool_use (“I’ll use get_weather for this.”), as Claude tends to write one; a forced choice goes straight to the call. With thinking on, a thinking block comes first, its text saying what the script decided and why. A final assistant message is a prefill, and the reply continues it.
Checked like the real API
Requests are refused where the API refuses them, in its words — {"type": "error", "error": {"type", "message"}, "request_id"}, with field paths like messages.1.content.0 — so the mistakes surface here, not in production:
- Missing and unknown fields (“max_tokens: Field required”, “foo: Extra inputs are not permitted”), wrong types, a missing or unknown
anthropic-version,max_tokensabove the model’s output limit, a prompt over its context window. - A
tool_usenot answered by atool_resultin the next message, or atool_resultanswering nothing; duplicate tool names; aninput_schemathat isn’t an object. - Thinking:
budget_tokensunder 1,024 or not belowmax_tokens, a forcedtool_choicewith thinking on, a temperature other than 1, a changed thinking block, a tool loop without its thinking block. - Caching: more than four
cache_controlblocks, a 1-hour breakpoint after a 5-minute one,cache_controlon empty text. - Models: unknown ones and retired ones (Claude 3, Claude 4.0, Mythos Preview) answer 404
not_found_error; an alias answers as its dated id.
Streaming, thinking and caching
A stream is the API’s event sequence: message_start (with input usage), then per block content_block_start, its deltas — text_delta, input_json_delta partial JSON, thinking_delta then one signature_delta — and content_block_stop, a ping after the first block starts, message_delta with the stop reason and final usage, and message_stop. The SDK’s stream helpers rebuild the same Message a plain request returns. Everything is decided before the first byte: the stream is paced, not computed as it goes.
Stop reasons are the API’s: end_turn, max_tokens, stop_sequence (naming the one hit), tool_use and refusal. Thinking spends tokens inside max_tokens — with enabled up to its budget, with adaptive by output_config.effort (none at low) — counted in output_tokens and in output_tokens_details.thinking_tokens.
Message Batches
POST /v1/messages/batches takes up to 50 requests (16 KB in all — the batch lives in your session token; the real API takes 100,000). The batch is in_progress for a few seconds (twenty with [[slow]] in a request), then ended with its request_counts and a results_url; the results are JSONL, one line per custom_id: succeeded with the message (service tier batch), or errored with the error that request would have got on its own. Cancel one in progress and it goes canceling, then ends with its requests canceled; delete it once it has ended. The session keeps the latest five.
Models
GET /v1/models lists these newest first, with limit, after_id and before_id, each with its display name, release date, limits and capabilities:
| Model | Name | Context / output | Thinking | Min. cached prefix |
|---|---|---|---|---|
| claude-sonnet-5-5 | Claude Sonnet 5.5 | 1M / 128K | enabled, adaptive | 1,024 |
| claude-fable-5-1 | Claude Fable 5.1 | 1M / 128K | enabled, adaptive | 1,024 |
| claude-opus-5-5 | Claude Opus 5.5 | 1M / 128K | enabled, adaptive | 4,096 |
| claude-mythos-5-1 | Claude Mythos 5.1 | 1M / 128K | enabled, adaptive | 1,024 |
| claude-opus-5 | Claude Opus 5 | 1M / 128K | enabled, adaptive | 4,096 |
| claude-sonnet-5 | Claude Sonnet 5 | 1M / 128K | enabled, adaptive | 1,024 |
| claude-fable-5 | Claude Fable 5 | 1M / 128K | enabled, adaptive | 1,024 |
| claude-mythos-5 | Claude Mythos 5 | 1M / 128K | enabled, adaptive | 1,024 |
| claude-opus-4-8 | Claude Opus 4.8 | 1M / 128K | enabled, adaptive | 4,096 |
| claude-opus-4-7 | Claude Opus 4.7 | 1M / 128K | enabled, adaptive | 4,096 |
| claude-sonnet-4-6 | Claude Sonnet 4.6 | 1M / 64K | enabled, adaptive | 1,024 |
| claude-opus-4-6 | Claude Opus 4.6 | 1M / 128K | enabled, adaptive | 4,096 |
| claude-opus-4-5-20251101 alias claude-opus-4-5 | Claude Opus 4.5 | 200K / 64K | enabled | 4,096 |
| claude-haiku-4-5-20251001 alias claude-haiku-4-5 | Claude Haiku 4.5 | 200K / 64K | enabled | 4,096 |
| claude-sonnet-4-5-20250929 alias claude-sonnet-4-5 | Claude Sonnet 4.5 | 200K / 64K | enabled | 1,024 |
What it answers
Answers carry request-id and the anthropic-ratelimit-* headers (a simulated quota; the hub’s X-Sondahub-RateLimit control makes real 429s). Errors on purpose carry retry-after, so the SDK’s retries are quick.
Questions
Is this Anthropic’s API?
No — an independent imitation of the Claude API for testing, not affiliated with or endorsed by Anthropic. No model runs and nothing is billed: a scripted model answers, so it is for testing what your code does with answers — streaming, tool loops, thinking blocks, batches, retries, errors — not for testing prompts.
Will the official SDKs work against it?
Yes. The Python SDK (anthropic) and the TypeScript SDK (@anthropic-ai/sdk) were both run against it: plain and streamed messages with the stream helpers, tool loops by hand and with client.beta.messages.tool_runner, extended and adaptive thinking carried across tool turns, messages.parse() with a pydantic model, count_tokens, batches and their JSONL results, models with pagination, and the SDK’s retries on 429 and 529. For calls that keep no state, ANTHROPIC_BASE_URL and ANTHROPIC_API_KEY are all it takes.
Are thinking signatures checked?
Yes. Each thinking block carries a signature over its text; send the block back changed and the request is refused (“Invalid signature in thinking block”), and with thinking.type: "enabled" a tool loop whose last assistant message doesn’t start with its thinking block gets the error the API gives for that. display: "omitted" returns the block with empty thinking and a signature, billed the same.
How does prompt caching work here?
As the API does it, with the session as the cache: each cache_control breakpoint (up to four, tools then system then messages, a 1-hour TTL never after a 5-minute one) writes its prefix when it is long enough — 1,024 tokens, 4,096 for Opus and Haiku models — and a later request that starts the same way reads it: cache_creation_input_tokens (split into 5-minute and 1-hour) the first time, cache_read_input_tokens after. The top-level cache_control marks the last block. Without the session token every request is a cold cache.
How are tokens counted?
With an approximate tokenizer — a word with its leading space, long words in pieces, numbers in threes, each punctuation mark — so counts land near the real ones without the real vocabulary. What matters is that it is consistent: usage, the streamed deltas and count_tokens all come from the same cut, and [[tokens:N]] is exactly N.
What is not modelled?
Server and client tools (web search, web fetch, code execution, bash, text editor, computer use, memory), MCP servers, the Files API, Skills, Agents and the admin API — those answer saying so — and citations, which are accepted and never produced. Images and documents in are accepted and counted. Temperature, top_p and top_k are accepted for older code.