Rate limits
A sliding window per key, the headers that carry the numbers, and why the numbers are not printed here.
Limits are applied per key over a sliding 60-second window. Not per organization, not per IP for authenticated traffic — per key. Two agents in one organization do not consume each other's budget, and rotating a key does not carry its recent history over.
The numbers travel in headers
Every authenticated response carries the current state for the key that made the request:
RateLimit-Limit: 600
RateLimit-Remaining: 594
RateLimit-Reset: 41RateLimit-Reset is seconds until the window has fully rolled off. A 429 adds Retry-After, also in seconds:
{
"error": {
"code": "rate_limited",
"message": "Over the limit agreed for this key; retry after Retry-After",
"request_id": "req_8a21"
}
}Read the headers rather than hard-coding a rate. They are on every authenticated response, so a client that watches RateLimit-Remaining never needs to guess and never needs a redeploy when a limit changes.
The headers are absent on unauthenticated responses, because without a valid key there is no key to report a limit for.
Why no number is printed on this page
Limits are set per key at onboarding, against your measured traffic, and stated in your agreement. A published figure would be either wrong for you or a ceiling nobody agreed to.
There is a separate, cruder limit at the network edge, applied per IP, which exists to keep unauthenticated traffic from reaching the application at all. It is not the limit your integration works against and it is not tuned per customer.
Request bodies are capped at 64 KB
A larger body is refused before it is parsed. This matters for one endpoint in practice — bulk metadata on a mandate — and it is a hard cap, not a per-key setting.

