sondahub

sondahub / Chaos and rate limits

Chaos, rate limits and the rest of real life

A client that only ever met a fast, healthy server has not been tested. Put one header on any request here and the answer is slow, failing, rate-limited or cut short — at the rate you choose — and test the retries, the backoff and the idempotency that production will need.

Chaos

One header makes any endpoint slow, flaky or broken, at the rate you choose:

X-Sondahub-Chaos: latency=800,jitter=200,fail=0.2,status=503,truncate=0.1
KeyDoes
latencyWait this many ms before answering (up to 10,000).
jitter± this many ms, at random (up to 5,000).
failFail this share of requests, 0 to 1. Nothing runs; the answer is an error.
statusThe failure's status, 400–599 (503 by default). 429 and 503 carry Retry-After.
truncateEnd this share of bodies halfway, as a dropped connection would — for parsers that must not trust what they got.

From a browser address bar: ?_chaos=latency=800,fail=0.2, or the shortcuts ?_delay=800, ?_fail=0.2, ?_status=500 (a status alone always fails). What was done is in X-Sondahub-Chaos-Applied.

Slow and flaky: 300–700 ms, half of them 503
curl -i https://api.sondahub.com/v1/store/products/1 -H "X-Sondahub-Chaos: latency=500,jitter=200,fail=0.5,status=503"
Always a 500
curl -i "https://api.sondahub.com/v1/store/products/1?_status=500"

Rate limits

X-Sondahub-RateLimit: 10/60        (or ?_ratelimit=10/60 — requests / seconds)

Answers carry the limit in both shapes clients look for — the IETF RateLimit-Policy and RateLimit fields and the classic X-RateLimit-Limit, -Remaining and -Reset — and once the window's quota is spent, 429 Too Many Requests with Retry-After.

Five per ten seconds — run it a few times
curl -i https://api.sondahub.com/v1/store/products/1 -H "X-Sondahub-RateLimit: 5/10"

Idempotency keys

Send Idempotency-Key with a POST, PUT, PATCH or DELETE, the way payment APIs ask you to. Inside a session the first answer is kept, and a retry with the same key gets it back unchanged, marked Idempotent-Replayed: true, instead of creating a second record. The same key with a different method, path or body answers 422 idempotency_key_reused.

The examples on this page share one session: run them in order and each sees what the one before it wrote.
Run it twice: one post, two identical answers
# the first try: keep the X-Sondahub-Session token it answers with
curl -i -X POST https://api.sondahub.com/v1/social/posts \
  -H "Content-Type: application/json" -H "Idempotency-Key: order-attempt-42" \
  -d '{"author_id":1,"body":"Posted once, whatever the retries."}'

# the retry, with that token: the first answer again, Idempotent-Replayed: true
curl -i -X POST https://api.sondahub.com/v1/social/posts \
  -H "Content-Type: application/json" -H "Idempotency-Key: order-attempt-42" \
  -H "X-Sondahub-Session: $TOKEN" \
  -d '{"author_id":1,"body":"Posted once, whatever the retries."}'

Cursor paging

Every list pages by number, and by cursor when asked: ?paging=cursor adds has_more, next_cursor and prev_cursor to meta; pass one back as ?cursor=. ?starting_after={id} and ?ending_before={id} page from a record, Stripe-style. A cursor is tied to its query — reuse it with other filters and it answers 400 bad_cursor.

The first page, by cursor
curl "https://api.sondahub.com/v1/store/products?paging=cursor&limit=3&fields=id,name"

Formats and errors

Any JSON answer to a GET comes in other shapes too — ?format=csv, xml, yaml, ndjson or msgpack, or the matching Accept type. And errors come as RFC 9457 problem details when the client asks with Accept: application/problem+json (or ?_errors=problem), each type an address that describes that error, at https://api.sondahub.com/v1/problems/{code}.

Products as CSV
curl "https://api.sondahub.com/v1/store/products?limit=5&fields=id,sku,name,price&format=csv"
Orders as XML
curl -H "Accept: application/xml" "https://api.sondahub.com/v1/store/orders?limit=2&fields=id,number,total"
A 404 as problem+json
curl -H "Accept: application/problem+json" https://api.sondahub.com/v1/store/products/999999

Questions

Which endpoints take these controls?

Every HTTP endpoint of the hub: the REST APIs, GraphQL, gRPC-Web and Connect, MCP, SCIM, the OpenAPI mock and the utilities. They run before the request is routed, so a chaos failure means nothing ran — retrying is always safe.

Is the rate limit per client?

It is per moment. The hub keeps no counters, so the quota drains with the clock: through the first 80% of each window the remaining count falls from the limit to zero, and the last 20% answers 429 with Retry-After until the window resets. Every client sees the same numbers at the same time — which makes a backoff test repeatable.

How do I test that my client honours Retry-After?

Two ways: X-Sondahub-RateLimit: 5/10 and a loop (429s arrive in the last two seconds of every ten), or X-Sondahub-Chaos: fail=0.5,status=503 — every 429 and 503 from chaos carries Retry-After: 2. /v1/utils/flaky and /v1/utils/status/{code} fail on a single endpoint, with a Retry-After of their own.

What does an Idempotency-Key do without a session?

The same key always produces the same new id, so even without a session two POSTs with one key describe the same record. With a session the first answer is remembered and replayed exactly — status, body and Idempotent-Replayed: true — and the same key on a different request answers 422.