sondahub

sondahub / OpenAI sandbox

An OpenAI mock API with a scripted model: the same request, the same answer

OpenAI’s API, answered by sondahub: Chat Completions and the Responses API with streaming, tool calls, structured outputs and reasoning tokens, embeddings, moderations and models — every request checked the way the API checks it. The model is a script, so tests are deterministic, instant and free, and a phrase in the message picks the outcome: a tool call, a refusal, a 429, a stream that breaks halfway.

An independent imitation for testing. Not affiliated with, or endorsed by, OpenAI. No model runs and nothing is billed.

Connect

Instead of
https://api.openai.com/v1
Use
https://api.sondahub.com/v1
API key
any key that starts with sk- — sk-sondahub
Also at
https://api.sondahub.com/sandbox/openai/v1
OpenAPI 3
https://api.sondahub.com/sandbox/openai/openapi.json

The quickest way: set OPENAI_BASE_URL=https://api.sondahub.com/v1 and OPENAI_API_KEY=sk-sondahub, and code that uses the official SDK talks to the sandbox unchanged. That is enough for anything stateless — chat completions, streams, tool calls, embeddings.

For state — stored responses, previous_response_id, background responses, the prompt cache — the SDK also has to carry X-Sondahub-Session from each answer to the next request, as below; there is no database, so the token is where those live.

Python (openai)

from openai import OpenAI, DefaultHttpxClient

# carry the sandbox session (stored responses, the prompt cache) from each answer to the next request
session = {}
def put(request):
    request.headers.update(session)
def keep(response):
    if 'X-Sondahub-Session' in response.headers:
        session['X-Sondahub-Session'] = response.headers['X-Sondahub-Session']

client = OpenAI(
    base_url='https://api.sondahub.com/v1',
    api_key='sk-sondahub',
    http_client=DefaultHttpxClient(event_hooks={'request': [put], 'response': [keep]}),
)

first = client.responses.create(model='gpt-5.5', input='Say this is a test')
print(first.output_text)                       # This is a test.
second = client.responses.create(model='gpt-5.5', input='And again?', previous_response_id=first.id)

Node (openai)

import OpenAI from 'openai'

// carry the sandbox session (stored responses, the prompt cache) from each answer to the next request
let session = null
const client = new OpenAI({
  baseURL: 'https://api.sondahub.com/v1',
  apiKey: 'sk-sondahub',
  fetch: async (url, init = {}) => {
    const headers = new Headers(init.headers)
    if (session) headers.set('X-Sondahub-Session', session)
    const res = await fetch(url, { ...init, headers })
    session = res.headers.get('X-Sondahub-Session') ?? session
    return res
  },
})

const stream = await client.chat.completions.create({ model: 'gpt-4o', stream: true, messages: [{ role: 'user', content: '[[tokens:200]]' }] })
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? '')

Import it. In Sonda: Import → From a URL, paste the OpenAPI address. Every operation comes with example bodies — a tool call, structured output, a stream, an error on purpose — so each works as it is. Set the auth to Bearer sk-sondahub.

Try it here

These examples share one session: run them in order and each sees what the one before it did.
Say this is a test
curl https://api.sondahub.com/v1/chat/completions \
  -H "Authorization: Bearer sk-sondahub" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "gpt-5.5",
  "messages": [
    {"role": "developer", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Say this is a test"}
  ]
}'
A stream, usage last — watch it arrive
curl -N https://api.sondahub.com/v1/chat/completions \
  -H "Authorization: Bearer sk-sondahub" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "gpt-4o",
  "stream": true,
  "stream_options": {"include_usage": true},
  "messages": [{"role": "user", "content": "[[tokens:60]] Tell me a story."}]
}'
A tool call, arguments from the schema and the message
curl https://api.sondahub.com/v1/chat/completions \
  -H "Authorization: Bearer sk-sondahub" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "gpt-5.5",
  "messages": [{"role": "user", "content": "What'\''s the weather in Paris?"}],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "The weather in a city",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {"type": "string"},
            "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
          },
          "required": ["location", "unit"],
          "additionalProperties": false
        },
        "strict": true
      }
    }
  ]
}'
Reasoning eats the budget: empty content, finish_reason length
curl https://api.sondahub.com/v1/chat/completions \
  -H "Authorization: Bearer sk-sondahub" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "o4-mini",
  "max_completion_tokens": 200,
  "reasoning_effort": "high",
  "messages": [{"role": "user", "content": "What is (17 * 23) + 4?"}]
}'
Structured output with the Responses API (stored in the session)
curl https://api.sondahub.com/v1/responses \
  -H "Authorization: Bearer sk-sondahub" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "gpt-5.5",
  "input": "Extract the event: Ada and Grace meet on Friday in London.",
  "text": {
    "format": {
      "type": "json_schema",
      "name": "event",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {
          "name": {"type": "string"},
          "date": {"type": "string", "format": "date"},
          "location": {"type": "string"},
          "participants": {"type": "array", "items": {"type": "string"}}
        },
        "required": ["name", "date", "location", "participants"],
        "additionalProperties": false
      }
    }
  }
}'
Embeddings: similar sentences, close vectors
curl https://api.sondahub.com/v1/embeddings \
  -H "Authorization: Bearer sk-sondahub" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "text-embedding-3-small",
  "input": [
    "The food was delicious and the waiter was friendly.",
    "The meal was tasty and the staff were kind."
  ],
  "dimensions": 8
}'
Strict mode without additionalProperties: the 400 the API gives
curl https://api.sondahub.com/v1/chat/completions \
  -H "Authorization: Bearer sk-sondahub" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "gpt-5.5",
  "messages": [{"role": "user", "content": "Extract: Ada, 36"}],
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "person",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {"name": {"type": "string"}, "age": {"type": "integer"}}
      }
    }
  }
}'
A 429 on purpose
curl https://api.sondahub.com/v1/chat/completions \
  -H "Authorization: Bearer sk-sondahub" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "gpt-5.5",
  "messages": [{"role": "user", "content": "[[error:429]] hello"}]
}'

From the shell, keep the token yourself: curl -i shows X-Sondahub-Session; send it back with -H "X-Sondahub-Session: …". How sessions work.

The scripted model

No model runs here. A script answers, and it reads the conversation the way a test wants a model to: the same request always gets the same answer, instantly, for nothing. It still does a few useful things on its own — “Say this is a test” answers This is a test., an arithmetic question gets its answer, a message that names a tool (by a word of its name: “weather” for get_weather) calls it with arguments made from its schema and your words (the city after “in”, the numbers in order), a turn that brings tool results back gets a reply that reports them, and a JSON schema gets an object that fits it, required fields, enums and formats included. Anything else gets a short note saying it is the model you asked for’s stand-in and how to steer it:

Put this in the messageWhat comes back
[[tokens:N]]A reply exactly N tokens long (up to 32,000), cut by the output limit like any reply — the way to test a long stream, a full buffer or finish_reason: "length".
[[echo]]The message back, as written.
[[tool]]Call the first tool. [[tool:name]] calls that one; [[tools:all]] calls every one at once (unless parallel_tool_calls is false).
[[refuse]]A refusal: message.refusal set and content null (a refusal part in the Responses API) — what .parse() hands back as refusal with no parsed object.
[[error:CODE]]That error in OpenAI’s shape: any status from 400 to 599 (429 is a rate limit, 500 a server error…), or quota (insufficient_quota) or context (context_length_exceeded).
[[flaky:CODE]]That error on the first try only: the SDK retries (it sends x-stainless-retry-count) and the retry gets the answer — the way to watch your retry settings work. Without that header, every try fails.
[[break]]A stream that dies halfway: an error chunk (Responses: error and response.failed) and no [DONE]. A plain request gets a 500.
[[slow]]A stream paced like a slow model: over a second before the first token, then 140 ms between deltas.
[[flag]]Moderations flag the text: [[flag:violence]] in that category (any of 13: harassment, harassment/threatening, hate, hate/threatening, illicit, illicit/violent, self-harm, self-harm/intent, self-harm/instructions, sexual, sexual/minors, violence, violence/graphic).

Checked like the real API

Requests are refused where OpenAI refuses them, in its words — {"error": {"message", "type", "param", "code"}} — so the mistakes surface here, not in production:

  • Unknown fields (“Unrecognized request argument supplied: foo”), wrong types and out-of-range values (invalid_type, decimal_above_max_value…), more than four stop sequences, stream_options without stream.
  • An assistant message with tool_calls not followed by a tool message for each id, or a tool message answering nothing; in the Responses API, a function_call_output with an unknown call_id, a continued response whose calls were never answered, and tools in the chat shape ("function": {…}, where the Responses API wants them flat).
  • Strict schemas without additionalProperties: false or with properties missing from required; json_object when no message says “json”.
  • Reasoning models: max_tokens instead of max_completion_tokens, a temperature or top_p other than 1, an effort the model doesn’t take; the GPT-5.1-and-later rule that sampling settings only go with reasoning_effort: "none".
  • Models: a Responses-only model on chat completions, an embedding model asked to chat, a retired one (the deprecation 404), a limit above the model’s output cap, a prompt over its context window.

Streaming

Chat completions stream as the API streams them: a first chunk with the role, content deltas a token or a few at a time, a chunk with the finish_reason, the usage chunk (empty choices) when stream_options.include_usage is set, and data: [DONE]; tool calls arrive as an id and name, then argument fragments. The Responses API streams typed events with sequence_number — response.created, response.in_progress, response.output_item.added, response.content_part.added, response.output_text.delta, the reasoning-summary and function-call-arguments events, the .done events, response.completed (or response.incomplete) — so the SDK’s stream helpers rebuild the same object a plain request returns. Deltas carry obfuscation padding unless include_obfuscation is false. Everything is decided before the first byte: the stream is paced (a short wait, then a few milliseconds per delta), not computed as it goes.

Models

Any of these answers (and GET /v1/models lists them, with the embedding and moderation models); an alias answers as its dated snapshot, as the real API does. Context window and output limit as the sandbox applies them:

ModelContext / outputKind
gpt-6.1-sol1.05M / 128Kreasoning, effort none by default
gpt-6-astra1.05M / 128Kreasoning, effort none by default
gpt-6-sol1.05M / 128Kreasoning, effort none by default
gpt-6-luna400K / 128Kreasoning, effort none by default
gpt-5.6-sol1.05M / 128Kreasoning, effort none by default
gpt-5.6-terra1.05M / 128Kreasoning, effort none by default
gpt-5.6-luna400K / 128Kreasoning, effort none by default
gpt-5.5
→ gpt-5.5-2026-04-23
1.05M / 128Kreasoning, effort none by default
gpt-5.5-pro
→ gpt-5.5-pro-2026-04-23
1.05M / 128Kreasoning, effort high by default; Responses API only
gpt-5.4-mini
→ gpt-5.4-mini-2026-03-17
400K / 128Kreasoning, effort none by default
gpt-5.4-nano
→ gpt-5.4-nano-2026-03-17
400K / 128Kreasoning, effort none by default
gpt-5.41.05M / 128Kreasoning, effort none by default
gpt-5.3-chat-latest128K / 16Kchat
gpt-5.2
→ gpt-5.2-2025-12-11
400K / 128Kreasoning, effort none by default
gpt-5.2-pro
→ gpt-5.2-pro-2025-12-11
400K / 128Kreasoning, effort high by default
gpt-5.2-chat-latest128K / 16Kchat
gpt-5.1-codex-max400K / 128Kreasoning, effort medium by default; Responses API only
gpt-5.1
→ gpt-5.1-2025-11-13
400K / 128Kreasoning, effort none by default
gpt-5.1-mini400K / 128Kreasoning, effort none by default
gpt-5.1-codex400K / 128Kreasoning, effort medium by default
gpt-5.1-chat-latest128K / 16Kchat
gpt-5-pro
→ gpt-5-pro-2025-10-06
400K / 128Kreasoning; Responses API only
gpt-5-codex400K / 128Kreasoning; Responses API only
gpt-5
→ gpt-5-2025-08-07
400K / 128Kreasoning
gpt-5-mini
→ gpt-5-mini-2025-08-07
400K / 128Kreasoning
gpt-5-nano
→ gpt-5-nano-2025-08-07
400K / 128Kreasoning
gpt-5-chat-latest128K / 16Kchat
o3-pro
→ o3-pro-2025-06-10
200K / 100Kreasoning (o-series); Responses API only
codex-mini-latest200K / 100Kreasoning (o-series); Responses API only
o4-mini
→ o4-mini-2025-04-16
200K / 100Kreasoning (o-series)
o3
→ o3-2025-04-16
200K / 100Kreasoning (o-series)
gpt-4.1
→ gpt-4.1-2025-04-14
1.048M / 33Kchat
gpt-4.1-mini
→ gpt-4.1-mini-2025-04-14
1.048M / 33Kchat
gpt-4.1-nano
→ gpt-4.1-nano-2025-04-14
1.048M / 33Kchat
o1-pro
→ o1-pro-2025-03-19
200K / 100Kreasoning (o-series); Responses API only
gpt-4o-search-preview128K / 16Kchat
gpt-4o-mini-search-preview128K / 16Kchat
o3-mini
→ o3-mini-2025-01-31
200K / 100Kreasoning (o-series)
o1
→ o1-2024-12-17
200K / 100Kreasoning (o-series)
chatgpt-4o-latest128K / 16Kchat
gpt-4o-mini
→ gpt-4o-mini-2024-07-18
128K / 16Kchat
gpt-4o
→ gpt-4o-2024-08-06
128K / 16Kchat
gpt-4-turbo
→ gpt-4-turbo-2024-04-09
128K / 4Kchat
gpt-3.5-turbo-012516K / 4Kchat
gpt-4
→ gpt-4-0613
8K / 8Kchat
gpt-4-06138K / 8Kchat
gpt-3.5-turbo
→ gpt-3.5-turbo-0125
16K / 4Kchat

Embeddings: text-embedding-3-small (1,536 dimensions), text-embedding-3-large (3,072) and text-embedding-ada-002 (1,536), dimensions on the v3 models. The vectors come from feature hashing — words, letter trigrams, word pairs — normalised to length 1: the same text always gives the same vector, and sentences that share words land close together, so a small semantic search built on them behaves sensibly. Moderations: omni-moderation-latest by default, all 13 categories with scores, nothing flagged without [[flag]].

What it answers

POST/v1/chat/completionsStreaming, n, tools and the legacy functions, response_format (json_schema, json_object), logprobs, stop, reasoning_effort, prompt caching.
POST/v1/responsesStreaming events, function and custom tools, text.format, reasoning summaries, previous_response_id, store, background, max_output_tokens, include.
GET/v1/responses/{id}A stored response (?stream=true plays it back as events); DELETE removes it; POST …/cancel stops a background one; GET …/input_items lists its input.
POST/v1/responses/input_tokensWhat an input would count as.
POST/v1/embeddingsfloat or base64, dimensions, up to 2,048 inputs.
POST/v1/moderationsText, or text and image parts.
GET/v1/modelsThe models above; GET /v1/models/{id} one of them.

Answers carry x-request-id, openai-processing-ms, openai-version and the x-ratelimit-* headers (a simulated quota; the hub’s X-Sondahub-RateLimit control makes real 429s). Errors on purpose carry retry-after-ms, so the SDK’s retries are quick.

Questions

Is this OpenAI?

No — an independent imitation of OpenAI’s API for testing, not affiliated with or endorsed by OpenAI. No model runs and nothing is billed: a scripted model answers, so it is for testing what your code does with answers — parsing, streaming, tool loops, retries, errors — not for testing prompts.

Will the official SDKs work against it?

Yes. The Python SDK (openai) and the Node SDK (openai) were both run against it: plain and streamed chat completions, .stream() and .parse() with pydantic models and pydantic_function_tool, tool-call loops, the Responses API with previous_response_id, responses.stream(), responses.parse(), background mode, input items, embeddings (the SDK’s default base64 included), moderations and models. For calls that keep no state, OPENAI_BASE_URL and OPENAI_API_KEY are all it takes.

Why is content empty with finish_reason "length"?

Because a reasoning model spends part of max_completion_tokens thinking before it writes, and the sandbox counts it the same way: reasoning_tokens by effort (minimal 24, low about 96, medium about 256, high about 768, xhigh about 1,536), the reply from what is left. Set the limit too low and nothing is left — the real gotcha, reproduced on purpose. The fourth example above does it.

How are tokens counted?

With an approximate tokenizer — a word with its leading space, long words in pieces, numbers in threes, each punctuation mark — so counts land near the real ones without the real vocabulary. What matters is that it is consistent: usage, the streamed deltas and input_tokens all come from the same cut, and [[tokens:N]] is exactly N.

Does prompt caching work?

Yes, with the session. A prompt of 1,024 tokens or more is remembered in 128-token steps for ten minutes (a day with prompt_cache_retention: "24h"), per model and prompt_cache_key; the next request that starts the same way reports the shared part as cached_tokens. Without the session token every request is a cold cache.

What is not modelled?

Audio and images out, built-in Responses tools (web search, file search, code interpreter, computer use, MCP), stored prompts and the Conversations API, stored chat completions (store is accepted, nothing is listed back), and files, batches, fine-tuning, assistants, vector stores, realtime and the rest of the platform — those paths answer a 404 saying so. Images and files in are accepted and counted.