Environments, limits and quotas
Sandbox and live credentials, the base URL convention, the rate budgets and concurrency compartments the gateway enforces, and what a 429 or a 503 means.
The base URL is per deployment
There is no single api.walletd.com. Every partner runs its own stack, so the host you call is issued with your credentials rather than published here. Throughout this site it is written as an environment variable:
export WALLETD_API="https://apigw.your-deployment.example"
export WALLETD_API_KEY="sk_sandbox_..."
curl -sS "$WALLETD_API/v1/users" -H "Authorization: Bearer $WALLETD_API_KEY"Two conventions hold everywhere. The version lives in the path, not in the host, so every route begins /v1. And the host never appears in a response body: identifiers are opaque strings, and nothing the API returns needs to be concatenated onto a base URL to be usable.
Sandbox and live
A credential carries its environment in the token itself.
| Sandbox | sk_sandbox_<key-id>_<secret> |
| Live | sk_live_<key-id>_<secret> |
The environment segment is not decoration. It is stored on the key record as well as encoded in the token, and verification compares the two: a token whose segment does not match its record fails, and the failure is the same opaque {"active": false} as a wrong secret or a revoked key. A sandbox key pointed at a production deployment does not find its record there at all, so it simply fails to authenticate rather than doing something expensive.
What differs between the two is the money, and only the money. Same routes, same request and response shapes, same problem codes, same webhook signatures, same idempotency semantics. Money in runs through the processor's test mode: a test-mode top-up produces a genuine topup.succeeded webhook with a real signature, so your verification code is exercised for real. See Testing for the test cards and the failure modes worth provoking.
A non-production environment is a separate deployment named in your agreement, not a flag on a request. Within a single deployment the environment segment on a key is checked, but a sandbox key and a live key issued against that same deployment reach the same data. Splitting sandbox and live into distinct data planes is a tracked gap, not a shipped isolation boundary, which is why development work belongs in the non-production deployment rather than behind a sandbox key in the production one.
Limits
Two different things are enforced at the edge, and they answer two different questions.
A rate budget answers "how often may this caller arrive". A concurrency compartment answers "how many requests in this traffic class may be in flight in one gateway process". They are independent: a caller comfortably inside its budget can still find a compartment full, because a compartment fills when the backend behind it slows down rather than when the caller speeds up. Compartments exist so that one slow money path cannot consume the capacity every other route needs.
The tables are regenerated at developer time from the staging gateway profile and committed with this site. They describe that source snapshot, not a live measurement or a partner deployment guarantee. Rate-budget identity depends on trusted network information; if the gateway falls back to a shared proxy address, multiple callers share a bucket. A compartment is shared by its traffic class within a gateway process.
Rate budgets
Every policy whose match applies is enforced, not just the most specific one, so
an endpoint budget and the API-wide budget are both spent by the same request.
Burst is the bucket the caller may drain at once; sustained is the rate it
refills at. When the store is unreachable is what the edge does if it cannot
reach its counter store: money-moving paths refuse, reads continue.
| Policy | Applies to | Burst | Sustained | When the store is unreachable |
|---|---|---|---|---|
staff_login | /v1/auth/login and below, POST | 10 | 12/minute | refused |
auth_token_exchange | /v1/auth/user_tokens and below, POST | 50 | 5/second | allowed |
topup_user | /v1/topups exactly, POST | 5 | 3/minute | refused |
topup_status | /v1/topups/ and below, POST | 60 | 1/second | allowed |
p2p_user | /v1/transfers and below, POST | 10 | 10/minute | refused |
reads_user | /v1/users and below, GET | 30 | 5/second | allowed |
gateway_webhooks | /v1/gateways/ and below, POST | 100 | 20/second | allowed |
payments_client | /v1/payments and below, POST | 200 | 50/second | refused |
discovery_ip | /v1/explore and below, GET | 60 | 10/second | allowed |
refunds_client | /v1/refunds exactly, POST | 30 | 1/second | refused |
rewards_convert | /v1/rewards/convert exactly, POST | 10 | 12/minute | refused |
client_onboard | /v1/clients exactly, POST | 5 | 6/minute | refused |
client_manage | /v1/clients/ and below, POST | 30 | 1/second | refused |
user_admin | /v1/users/ and below, POST | 30 | 1/second | refused |
purchase_subscribe | /v1/explore/ and below, POST | 30 | 1/second | refused |
tenant_global | every request | 200 | 50/second | allowed |
Anything matching no policy falls to the plugin defaults: burst 100, sustained 25/second, allowed when the store is unreachable.
The caller's identity for all of these is its source address, read in order from client-ip-header:CF-Connecting-IP, forwarded-for, remote-addr. It is not the API key: one backend behind one egress address is one identity no matter how many keys it holds.
Concurrency compartments
A compartment caps how many requests of its class may be in flight to a backend
at once. Unlike the budgets above, exactly one compartment applies to a request:
the first whose match fits, then default. Wait is how long a request waits
for a free slot before it is shed.
| Compartment | Applies to | In flight at once | Wait |
|---|---|---|---|
payments | /v1/payments and below, POST | 150 | 50 ms |
transfers | /v1/transfers and below, POST | 40 | 50 ms |
topups | /v1/topups and below, POST | 40 | 50 ms |
refunds | /v1/refunds and below, POST | 40 | 50 ms |
reads | GET | 120 | 0 ms |
default | everything not matched above | 80 | 0 ms |
The published figures come from the staging settings profile, which is the gateway image's default. A partner deployment may be configured differently, and your agreement governs, not this page. The ceilings are sized from each backend's healthy concurrency and are meant to be tuned against measured capacity. Ask for the profile your deployment runs before you size a client around these numbers.
Two things that surprise people
The rate-limit identity is the source address, not the API key. One backend behind one egress address is one identity however many keys it holds, and a busy load generator on one machine is also one identity. If you are sizing a batch job, size it against one budget, not against one per worker.
Every matching budget is spent, not just the most specific one. A POST /v1/transfers spends the p2p_user budget and the API-wide tenant_global budget on the same request. Compartments work the other way: exactly one applies, the first whose match fits.
When you are refused
Both refusals are application/problem+json and both are the gateway's, not walletd's. They are the reason the error code table lists rate_limited and bulkhead_full separately from every other code: nothing in walletd emits them, and no amount of correcting your request body prevents them.
429, you are asking too often
{
"type": "about:blank",
"title": "Too Many Requests",
"status": 429,
"code": "rate_limited",
"retry_after_sec": 3
}Note the shape. The gateway's problem body carries retry_after_sec and no detail, because it is written before any service has seen the request. Alongside it:
| Header | Meaning |
|---|---|
Retry-After | Whole seconds to wait. Always at least 1 |
X-RateLimit-Limit | The budget's capacity |
X-RateLimit-Remaining | What is left in it |
X-RateLimit-Used | What has been spent |
X-RateLimit-Reset | Unix seconds at which the budget is whole again |
Repeated denials from the same identity are short-circuited for a few seconds without touching the counter store. The Retry-After on a short-circuited response is aged by the time since the denial was cached, so it never tells you to wait longer than you actually must.
503, the class is momentarily full
{
"type": "about:blank",
"title": "Service Unavailable",
"status": 503,
"code": "bulkhead_full",
"retry_after_sec": 1
}With Retry-After: 1 and X-Bulkhead naming the saturated compartment. This is not an outage and it is not your budget. It means the compartment your request classified into had no free slot, and after waiting its configured window the request was shed rather than queued, so that the overload stayed inside that class. X-Bulkhead: payments while reads keep flowing is the system working as designed.
Retrying either
Both are retryable, and both must be retried with the same Idempotency-Key. A 429 or a 503 from the gateway means the request never reached a money path, but your client cannot know that from the status code alone, and reusing the key makes the question irrelevant. Back off exponentially with jitter on top of Retry-After. The per-language retry helpers on Idempotency already do this.
What the edge does when its own store is unavailable
The rate limiter keeps its counters in a shared store so budgets hold across gateway replicas. When that store cannot be reached, each policy falls one way or the other, and the table above says which.
Money-moving paths fail closed: transfers, top-ups, refunds, payments, conversions, client and user administration are refused rather than allowed through unmetered. Reads and the signature-verified processor callbacks fail open, because availability wins where no money moves. A circuit breaker stops the gateway hammering a store that is already down.
Next
- Errors for the full problem code table and what is retryable.
- Idempotency for the retry patterns in six languages.
- Security for integrators for key lifecycle, scopes and rotation.
- Support, SLAs and status if a limit is refusing traffic you believe is legitimate.