Error handling and debugging
Every Gate error carries a machine-readable code and source so you can branch on stable values, see who failed, and know exactly what to do next.
How to read any Gate error
Every error response is JSON with this envelope:
{ "error": { "code": "insufficient_balance", "message": "Insufficient PAYG balance for this request....", "source": "gateway", "balanceCents": 12, "availableCents": 0, "estimatedCostCents": 47, "billingUrl": "/billing" }}Two fields carry the essential triage information:
code: the machine-readable error identifier. Branch on this, not the prose message.source: which side was responsible. This is usually more useful than the HTTP status, because a503could be Gate’s or the provider’s, and a401could be a bad Gate key or a bad upstream credential.
The same source value also rides on the x-gate-error-source response header, so you can read it without parsing the body.
Error source reference
source | Who failed | Retry the same request? |
|---|---|---|
gateway | Gate stopped the request early (balance, rate limit, security, internal error). | Only after fixing the cause. |
pricing | Gate could not price the request and refused to serve it unpaid. | Not until pricing is available. |
customer_config | A provider account Gate manages is misconfigured (missing credentials, wrong model, etc.). | Not until the configuration is fixed. |
byok_credentials | The provider key you supplied was rejected by the upstream. | Not until you fix the key. |
provider | The upstream provider’s infrastructure failed. | Usually yes, with backoff. |
vendor | The model API returned an error (rate limit, bad input, overloaded). | Depends on the vendor error. |
validation | Your request shape was invalid (bad JSON, wrong method, unsupported header). | Not until you fix the request. |
The request ID, your fastest debugging path
Every response, including error responses, carries X-Gate-Request-Id. Paste that value into the Messages page on the dashboard to see the full record: prompt, response, tokens, cost, cost status, security verdict, and upstream provider. It is the single fastest path from “something went wrong” to “here is exactly what happened.”
Complete error code reference
The codes a model request on a gateway API key can return, with HTTP status and source. Free-model and X-Gate-Authorization sign-in codes are not listed; security scan codes are in Security scan.
| Code | HTTP | Source | Short description |
|---|---|---|---|
invalid_key | 401 | validation | The API key is missing, invalid, or revoked. |
auth_unavailable | 503 | gateway | Gate could not check the API key right now. Retry after Retry-After. |
unknown_route | 404 | validation | This Gate deployment does not serve that path. |
no_route | 400 | validation | No provider account can serve the requested model. |
invalid_upstream | 400 | validation | The X-Gate-Upstream-Url header is missing or malformed. |
missing_upstream_header | 400 | validation | Passthrough tokens require an explicit upstream URL header. |
rate_limit_exceeded | 429 | gateway | The API key or org has exceeded its rate limit. |
usage_limit_exceeded | 429 | gateway | A configured cost/token/request spending cap was hit. |
budget_exceeded | 429 | gateway | A hard org or team budget was hit. |
insufficient_balance | 402 | gateway | Your prepaid balance can’t cover this request. |
payg_reserved | 429 | gateway | Your other in-flight requests hold the balance right now. Retry shortly. |
payg_busy | 429 | gateway | Billing for this workspace is briefly busy. Retry shortly. |
payg_disabled | 402 | gateway | Paying through Gate isn’t enabled for this workspace. |
org_not_provisioned | 402 | gateway | The workspace hasn’t completed paying through Gate setup. |
pricing_unavailable | 503 | pricing | Gate cannot price the model for paying through Gate. |
security_blocked | 403 | gateway | The request was blocked by Gate’s security policy. |
ssrf_blocked | 403 | gateway | The upstream URL was blocked by Gate’s network policy. |
circuit_open | 503 | provider | The upstream is temporarily unavailable. |
upstream_error | 502 | provider | The upstream request failed. |
upstream_timeout | 504 | provider | The upstream request timed out. |
upstream_error_in_2xx | 502 | vendor | The upstream answered 2xx but the payload was an error. |
upstream_response_too_large | 502 | provider | The upstream response exceeded the size Gate can relay. |
provider_auth_failed | 502 | provider | The provider refused Gate’s own credential. Your key is fine. |
missing_credentials | 503 | customer_config | The provider account has no credentials configured. |
bedrock_sdk_unavailable | 503 | customer_config | The Bedrock SDK is not available in this gateway build. |
invalid_byok_credentials | 401 | byok_credentials | The upstream rejected the caller-supplied your own keys credential. |
control_plane_credential_required | 403 | validation | Claude Code managed policy needs your own Anthropic credential on the request. |
control_plane_busy | 503 | gateway | Too many concurrent control-plane reads on this gateway task. |
request_too_large | 413 | validation | The request body exceeds the maximum allowed size. |
invalid_request | 400 | validation | The request shape is invalid (bad JSON, wrong method, etc.). |
not_found | 404 | validation | The requested resource does not exist. |
forbidden | 403 | gateway | Access denied. |
method_not_allowed | 405 | validation | The HTTP method is not allowed on this endpoint. |
duplicate_entry | 409 | validation | A record with this value already exists. |
internal_error | 500 | gateway | An unexpected internal error occurred. |
Error details by category
Authentication errors
auth_unavailable, the API key could not be checked (HTTP 503)
{ "error": { "code": "auth_unavailable", "message": "API key validation is temporarily unavailable. Retry shortly.", "source": "gateway", "retryable": true }}Gate could not look your key up at all (for example, its key store was briefly unreachable), so the key was never judged. This is not a verdict on your credential: a missing, invalid, or revoked key always gets 401 invalid_key. Nothing was served or charged, and the request never reached a provider.
What to do. Retry after Retry-After (1 second), with jittered backoff. The response is retryable: true. Do not rotate or discard the key because of this code.
Routing errors
unknown_route, this deployment does not serve that path (HTTP 404)
{ "error": { "code": "unknown_route", "message": "This path is not served by this Gate deployment. Check the path; if the route is newer than the deployment, check its release tag at GET /version.", "source": "validation" }}The path is inside one of Gate’s own API namespaces (/v1/me, /v1/security, /v1/audit - the namespace root itself, or anything under it) and no endpoint there matches it. Almost always one of two things: a typo, or a route that exists in the docs but has not reached the deployment you are calling yet.
What to do. Check the path first, then check the deployment: GET /version is public, needs no key, and returns the release tag serving your requests. If the endpoint you want is newer than that tag, you are ahead of the deployment and there is nothing to fix on your side.
GET /version is also how you probe for a route without credentials. An unrouted path under those namespaces returns this 404 unauthenticated, while a path that does exist returns 401 invalid_key - so the two are distinguishable before you have a key.
BYOK passthrough is exempt. A request carrying X-Gate-Upstream-Url is forwarded to the upstream you named, verbatim path and all, even inside these namespaces - Gate never manufactures a 404 in place of your own server’s answer. If you get this error on a passthrough request, the header is missing or empty.
Two consequences of that exemption are worth knowing if you forward under these three prefixes:
- It costs double against the IP rate limit. Gate’s 100 req/60s per-IP limit is evaluated once when the path matches the 404 handler and again when it reaches the proxy, so the effective ceiling for these paths is about 50 req/min. Your per-key and per-org limits are unaffected. Forwarding under any other path (
/v1/messagesand friends) is unchanged and not subject to this. - A passthrough token with no
X-Gate-Upstream-Urlreadsunknown_routehere, not themissing_upstream_header400 you would get on any other path. The fix is the same one that error names: set the header.
Note that unknown_route means the path is not served. not_found (also 404) means the endpoint exists but the specific resource behind it does not - an unknown scan jobId, for example. And no_route below, despite the name, is not about paths at all: it means no provider can serve your model.
no_route, model not available on any configured provider (HTTP 400)
{ "error": { "code": "no_route", "message": "Model \"claude-sonnet-4-5\" is not available on any provider …", "source": "validation", "param": "model", "requested_model": "claude-sonnet-4-5", "suggested_models": ["claude-3-5-sonnet-20241022"], "available_provider_types": ["bedrock", "openrouter"], "available_account_names": ["my-bedrock", "my-openrouter"] }}Gate could not find a provider that can serve the model you requested.
The error body includes useful extras:
requested_model: the model ID you sent.suggested_models: similar models Gate can serve (if any).available_provider_types: the provider types available to your organization.available_account_names: the provider accounts available to serve your request.
What to do. Use the X-Gate-Provider header or a "provider" body field to target a specific provider, pick one of the suggested_models, or use your own provider key for the model you want. Model IDs must match what the provider exposes, so the suggested_models list is a good starting point if you have a typo.
invalid_upstream / missing_upstream_header, your own keys upstream URL issue (HTTP 400)
These arise when you use the passthrough path:
invalid_upstream: theX-Gate-Upstream-Urlheader is present but the value is missing or malformed.missing_upstream_header: a passthrough token requires an explicit upstream URL, but noX-Gate-Upstream-Urlwas provided.
What to do. Supply a valid X-Gate-Upstream-Url value pointing at the provider’s API endpoint (for example, https://api.anthropic.com). See Billing mode.
Rate and usage limit errors
rate_limit_exceeded, API key or org rate limit (HTTP 429)
{ "error": { "code": "rate_limit_exceeded", "message": "Rate limit exceeded. Please retry after the reset window.", "source": "gateway", "retryAfter": 1749200460000 }}Gate sets three headers alongside a 429 so you can implement backoff precisely:
| Header | Meaning |
|---|---|
X-RateLimit-Limit | The configured limit for this window. |
X-RateLimit-Remaining | How many requests remain in the current window. |
X-RateLimit-Reset | Unix timestamp (seconds) when the window resets. |
X-RateLimit-Scope | Which limit fired: org or passthrough-ip. |
retryAfter in the error body echoes X-RateLimit-Reset as a convenience for clients that parse the body but not headers.
X-RateLimit-* headers describe rate limiting only. They appear on a throttling 429 and on allowed requests, never on a balance refusal (402, payg_reserved 429, payg_busy 429). payg_reserved and payg_busy share the 429 status but are not throttling: tell them apart by the code, or by the absence of X-RateLimit-*.
What to do. Back off until X-RateLimit-Reset and retry. If you hit org-level limits regularly, contact support to discuss a higher plan. See Limits and retention.
usage_limit_exceeded, configured spending cap hit (HTTP 429)
{ "error": { "code": "usage_limit_exceeded", "message": "Usage limit exceeded: \"Monthly budget\" (cost limit of 100 p…", "source": "gateway", "limitType": "cost", "threshold": 100, "currentUsage": 100.4, "timeWindow": "month" }}Your organization (or the specific API key) has a configured spending cap, a rolling cost, token, or request budget, and this request would exceed it.
The extras tell you exactly which limit fired:
limitType:cost,tokens, orrequests.threshold: the cap value.currentUsage: what you’ve used so far in the window.timeWindow: the rolling window (day,week,month).
What to do. Wait for the window to reset, raise the cap in Dashboard → Limits, or switch to your own provider key that is not subject to the same cap. See Limits and retention.
Security errors
security_blocked, request blocked by security policy (HTTP 403)
For Anthropic-compatible endpoints:
{ "type": "error", "error": { "type": "permission_error", "message": "Request blocked by security policy." }}For all other endpoints:
{ "error": { "code": "security_blocked", "message": "Request blocked by security policy.", "source": "gateway", "type": "permission_error" }}Gate’s prompt-injection detection decided this request should be blocked. The x-gate-error-source: gateway header is always present on blocked responses.
No tokens were generated and nothing was billed.
What to do.
- If this is a false positive, open the request in the dashboard (Messages page, paste
X-Gate-Request-Id) to see the security verdict and the flagged content. - Review the prompt for patterns that resemble injection attempts.
- On Pro, lower the prompt-injection sensitivity on the Policies page if your use case consistently triggers false positives. The same card sets whether a detection flags or blocks.
Note: Nothing you can put in a request skips the scan — no header, no marker, no “system” text in the body. See Prompt-injection defense.
ssrf_blocked, upstream URL blocked (HTTP 403)
{ "error": { "code": "ssrf_blocked", "message": "Upstream URL blocked by network policy.", "source": "gateway" }}When you use the passthrough path with X-Gate-Upstream-Url, Gate checks the target URL against its network policy before making the outbound call. This error means the URL resolved to an address that is not permitted.
What to do. Make sure X-Gate-Upstream-Url points at a legitimate, publicly routable provider API endpoint (for example, https://api.openai.com). Internal IP ranges, localhost, and link-local addresses are not permitted.
Provider and infrastructure errors
circuit_open, upstream temporarily unavailable (HTTP 503)
{ "error": { "code": "circuit_open", "message": "Upstream \"my-bedrock\" is temporarily unavailable.", "source": "provider" }}When Gate sees sustained failures from a provider account, it stops sending traffic to that account for a short time. This protects your traffic from hammering an unresponsive upstream, and it clears automatically after a short delay.
What to do. Retry with exponential backoff. If you have multiple provider accounts, add a provider hint (X-Gate-Provider) to route to a healthy alternative while the failing account recovers. Check the upstream provider’s status page for outage information.
upstream_error / upstream_timeout, provider infrastructure failure (HTTP 502 / 504)
{ "error": { "code": "upstream_error", "message": "Upstream request failed. Please try again later.", "source": "provider" }}The upstream provider returned a 5xx error or the connection timed out. The x-gate-upstream-status response header pins the provider’s original HTTP status code when Gate relays the failure.
What to do. Retry with backoff. These errors come from the provider side, so a retry usually resolves them. Check the upstream provider’s status page if they persist.
upstream_error_in_2xx, error payload inside a 2xx response (HTTP 502)
{ "error": { "code": "upstream_error_in_2xx", "message": "Upstream returned an error payload despite a 2xx HTTP status.", "source": "vendor" }}The upstream answered with HTTP 200 but the response body contained an error envelope (for example, an Anthropic {"type": "error",...} or a stream that closed without a terminal event). Gate reclassifies these so the dashboard surfaces the failure rather than a misleading green badge.
Provider account configuration errors
missing_credentials, provider account has no credentials (HTTP 503)
{ "error": { "code": "missing_credentials", "message": "Provider account is missing credentials.", "source": "customer_config" }}The provider account Gate would use to serve this request has no credentials configured. This is a platform-side configuration issue, not something you can change from your dashboard.
What to do. Use a different model or provider, or use your own provider key for this request. If it keeps happening, contact support.
invalid_byok_credentials, upstream rejected your provider key (HTTP 401)
{ "error": { "code": "invalid_byok_credentials", "message": "BYOK credential rejected by the upstream provider.", "source": "byok_credentials" }}The key you passed (for example via Authorization: Bearer <your-key>) was rejected by the upstream provider. The key may be expired, revoked, or scoped to a different resource.
What to do. Verify the key is valid with a direct call to the provider, then update it in your client.
control_plane_credential_required, Claude Code managed policy needs your own credential (HTTP 403)
{ "error": { "code": "control_plane_credential_required", "message": "Claude Code managed policy and account settings are read from your own Anthropic account, so this request must carry your own credential: send your Gate key as ANTHROPIC_AUTH_TOKEN and your Anthropic key as ANTHROPIC_API_KEY. The gateway will not answer with the platform account's policy.", "source": "validation" }}Claude Code reads two things from Anthropic that are not model calls: your organization’s admin-set managed policy (GET /api/claude_code/policy_limits) and your account settings (GET /api/oauth/account/settings). Gate relays both to your own Anthropic account when the request carries a credential of yours: your request is forwarded with your credential, and the upstream’s status and body come back verbatim. Response headers are limited to content-type, rate-limit telemetry, request-id and, on a redirect, location. Gate does not follow redirects on these paths, and marks the response no-store.
This error means the request carried no such credential, so Gate had nothing to ask with. Gate will not answer these two paths using a platform provider account: the payload describes what that account’s administrator decided, and returning it would make your Claude Code enforce another organization’s restrictions and compliance settings.
What to do. Send both credentials in their native slots, which needs no custom header: your Gate key as ANTHROPIC_AUTH_TOKEN (Gate authenticates and attributes the request with it) and your Anthropic key as ANTHROPIC_API_KEY (Gate forwards that one upstream). Requests paid through Gate cannot read managed policy, by design; if your organization does not set managed policy, nothing is lost by the refusal.
Gate answers 403 rather than 404 deliberately. Claude Code treats a 404 on this path as a successful “no restrictions configured” answer and caches it, so answering 404 would have Gate assert something about your organization that it cannot know. A 403 says only what is true: the policy could not be retrieved.
Note that Claude Code only requests managed policy when it is pointed at
api.anthropic.comdirectly. Behind any custom base URL, including Gate, the client skips the fetch entirely, so today these paths are reached only by tooling that calls them explicitly.
Paying through Gate errors
Everything in this section applies to traffic paid through Gate only. Requests on your own keys are billed by the provider on your own credential, so the balance and pricing errors below do not apply to them. See Billing mode.
How a request paid through Gate is billed (so the errors make sense)
A request paid through Gate is billed in two steps:
-
Before the call (a hold, not a charge). Before contacting any provider, Gate checks two things: can it price this model at all, and can your balance cover a worst-case estimate? If either check fails, the request is refused before any provider call, so you are never charged for a refusal. If both pass, Gate places a temporary hold on your balance.
-
After the response (the real charge). Once the response completes, Gate releases the hold and debits the exact amount the provider charged. You are billed what the provider charged, one for one. See Billing mode.
A hold is never a charge. It is released as soon as the charge is final, and the only money that ever moves is the provider-exact debit.
Which refusals to retry
| Code | HTTP | Retry? | What it means |
|---|---|---|---|
insufficient_balance | 402 | No. retryable: false, no Retry-After. | Your balance can’t cover this request even with nothing else in flight. |
payg_reserved | 429 | Yes, after Retry-After. retryable: true. | Your balance covers it, but your other in-flight requests hold it now. |
payg_busy | 429 | Yes, after Retry-After. retryable: true. | Too many requests are billing against the workspace at the same moment. |
payg_disabled | 402 | No. | Paying through Gate isn’t active for this workspace. |
Balance headers on successful responses
Every admitted request paid through Gate carries the balance it was admitted against, on streaming and non-streaming responses alike, so a client can warn or top up before it hits a refusal:
| Header | Meaning |
|---|---|
X-Gate-Payg-Balance-Cents | Your balance in whole cents, rounded down as on refusals: at admission if streaming, after this charge if not. |
X-Gate-Payg-Available-Cents | What was still spendable at admission, after this request’s hold and the holds of your other in-flight requests. |
X-Gate-Payg-Debited-Cents | What this request was charged (non-streaming responses only). |
pricing_unavailable, Gate can’t price the model (HTTP 503)
{ "error": { "code": "pricing_unavailable", "message": "Model \"some/model\" is not in our pricing catalog for provider \"my-bedrock\"....", "source": "pricing" }}Gate refuses to serve a request paid through Gate it cannot bill, because serving one would mean giving you the request for free and silently absorbing the provider cost. There are two sub-reasons:
- Catalog miss. Discovery has never seen this
(provider account, model)pair, or the model ID doesn’t match what discovery recorded. Gate has no row to price against. - Pricing missing. Discovery did record the model, but neither the upstream nor the rate index produced a usable number for it.
Both surface as the same 503 pricing_unavailable code, so your client only needs to handle one case. The difference matters for fixing the problem, not catching it.
What to do.
- Wait for the next scheduled catalog refresh, or switch that model to your own provider key (the provider bills you directly, so Gate skips pricing).
- Double-check the model ID you sent matches what the provider exposes.
- Switch that model to your own provider key. Requests on your own keys skip pricing entirely because the provider bills you directly.
insufficient_balance, balance can’t cover the request (HTTP 402)
{ "error": { "code": "insufficient_balance", "message": "Insufficient PAYG balance for this request. Worst-case cost…", "source": "gateway", "balanceCents": 3, "availableCents": 3, "estimatedCostCents": 47, "billingUrl": "/billing", "retryable": false }}The hold couldn’t be placed because your balance, on its own, can’t cover this request’s worst-case estimate. No retry can fix that, so the response is retryable: false and carries no Retry-After. Read the three numbers together:
balanceCents: your total balance.availableCents: your balance minus the holds placed by other in-flight requests. This is what’s spendable right now.estimatedCostCents: the worst-case estimate this request needs to reserve.
You get this code exactly when estimatedCostCents > balanceCents (balanceCents is rounded down, so 3.9999¢ reads 3). When your balance would cover the request but your other in-flight requests hold too much of it, you get payg_reserved (below) instead.
How the worst-case estimate is computed. Gate reserves a conservative worst-case estimate of the request’s cost with a safety margin. The hold is released once the charge is final, and you are billed the exact provider cost.
What to do. Top up at /billing, or lower the request’s output cap (max_tokens / max_output_tokens) so its worst-case estimate is smaller. See Plans.
payg_reserved, balance held by your other in-flight requests (HTTP 429)
{ "error": { "code": "payg_reserved", "message": "This workspace's PAYG balance is fully reserved by other in-flight requests. Nothing was charged; retry after they finish.", "source": "gateway", "retryable": true, "retryAfter": 5, "balanceCents": 19, "availableCents": 2, "estimatedCostCents": 4, "billingUrl": "/billing" }}Your balance covers this request (balanceCents >= estimatedCostCents), but the holds of your other in-flight requests leave less than it needs (availableCents < estimatedCostCents). Each request holds its own worst-case estimate until it finishes, so a burst of concurrent requests on a small balance can reserve all of it.
This is transient. The holds are released as those requests finish, so the response is retryable: true and carries a Retry-After header (seconds). Nothing was reserved or charged for the refused request.
What to do. Retry after Retry-After, with backoff. If you see this often, your concurrency is high for your balance: keep the balance (or auto-topup threshold) well above your in-flight worst case, or reduce in-flight concurrency. Once the remaining balance can’t cover the request even with nothing in flight, the refusal becomes insufficient_balance.
payg_busy, billing briefly busy (HTTP 429)
{ "error": { "code": "payg_busy", "message": "Too many concurrent requests are billing against this workspace right now. The request was not served and nothing was charged — retry shortly.", "source": "gateway", "retryable": true, "retryAfter": 1 }}Gate places one workspace’s holds one at a time. When many requests arrive at the same moment, a request that would wait too long for its turn is refused rather than queued. Nothing was reserved or charged. It is a 429 but not throttling: it carries no X-RateLimit-* headers.
What to do. Retry after Retry-After (1 second), with jittered backoff.
insufficient_balance with reason: "cost_unbounded" (HTTP 402)
{ "error": { "code": "insufficient_balance", "message": "This request's maximum cost could not be bounded, so it cannot be safely billed against your PAYG balance. Set an explicit max output (max_tokens / max_output_tokens) or use a BYOK key.", "source": "gateway", "reason": "cost_unbounded", "billingUrl": "/billing" }}Note the "reason": "cost_unbounded". This is a different failure from a plain balance refusal. Gate could not compute any worst-case ceiling for the request (no catalog estimate and no observed cost to scale from), so it refused rather than serve a request whose cost it can’t bound.
What to do. Set an explicit output cap (max_tokens / max_output_tokens) so the estimate has an upper bound, or use your own provider key for that request.
payg_disabled / org_not_provisioned (HTTP 402)
{ "error": { "code": "payg_disabled", "message": "PAYG is not enabled for this workspace. Enable it in /billing or use a BYOK key.", "source": "gateway" }}payg_disabled: Paying through Gate isn’t switched on for this workspace. Enable it under Billing, or send the request as your own keys.org_not_provisioned: the workspace hasn’t completed paying through Gate setup.
An overdraft after a successful response
A request can pass the up-front check and still fail to charge cleanly, and this case is easy to miss because you do not get an HTTP error. The request already succeeded and the response was already delivered to your client. Instead, Gate records the gap in your audit log as a payg.debit.insufficient_balance event:
{ "event_type": "payg.debit.insufficient_balance", "message": "PAYG balance insufficient at charge time, overdraft after the up-front check", "data": { "balanceCents": 0, "requiredCents": 5, "costUsd": "0.000512" }}Why the up-front check passes but the charge fails. The hold reserves a conservative estimate. By the time the response is complete, the real cost is known, and your spendable balance may have changed in between:
- Concurrent burn-down. Several requests pass the check at nearly the same time on a near-empty balance. As each one’s real cost is charged, the balance drains, and a later one finds nothing left to charge.
- A balance edit in between. Something changed the balance (a manual adjustment, a refund) after the hold was placed but before the charge landed.
What happens to the request. Gate already served it, so your response is real and complete. Gate does not fail the request mid-stream. It records the uncharged amount and moves on. The balance is not pushed negative.
How the money is recovered. Gate automatically re-applies any debit for a request that was served (2xx), produced a real positive cost, and has no matching ledger entry. It uses an idempotency key so it can never double-charge, it debits the exact provider cost rather than an estimate, and it skips anything the balance still can’t cover. An overdraft like this is a timing gap, not lost money. Your next top-up squares the books.
Telling the two “insufficient balance” cases apart in your audit log.
Audit reason/event_typeWhen Was the request served? insufficient_balance_reservationUp-front, hold couldn’t be placed. No. Refused with 402. payg_reserved_by_inflightUp-front, balance held by in-flight requests. No. Refused with 429. payg.debit.insufficient_balanceCharge time, real charge couldn’t be applied. Yes, then recovered later. The up-front refusal is the normal case: Gate compares your worst-case estimate against your available balance (total minus other in-flight holds) and refuses early. The charge-time event is the only one where the request actually ran.
Understanding cost: estimate vs. catalog vs. what you’re billed
Three different numbers can appear around a single request:
| Number | Where you see it | What it is |
|---|---|---|
| Worst-case estimate | The 402 refusal message / hold | A conservative upper bound used only to size the hold. |
| Catalog estimate | Cost-divergence reporting (informational) | Gate’s own computed cost from the pricing catalog. |
| Billed cost | X-Gate-Cost-Usd, cost_usd, ledger | The exact provider cost, the only number that moves money. |
The rule that resolves all confusion: you are billed exactly what the provider charged, one for one. The catalog is used for estimating the hold. It never decides what you pay. Per-request cost is calculated against the provider catalog and we verify accuracy.
A note on
_cents. Your balance and ledger are tracked at sub-cent precision. The*_centsfigures in headers, the API, and the dashboard are a rounded display mirror of those amounts. They are correct for display but not for reconciliation. For exact reconciliation, sum the 6-decimalcost_usdvalues, not the rounded cents.
Why a request’s cost couldn’t be calculated
If X-Gate-Cost-Usd is absent or cost_usd is null, read X-Gate-Cost-Calculation-Status (the cost_status field on the request record) to learn why. A 0.000000 is never ambiguous. The status tells you whether it’s a genuine zero or an uncalculable one.
| Status | Meaning |
|---|---|
calculated | Priced from token usage. Normal case; may be a legitimately tiny or zero amount. |
upstream_reported_cost | Priced from the provider’s own per-request cost (the authoritative, provider-exact figure). |
cache_hit | Served from Gate’s cache, never reached a provider, so cost is zero by definition. |
byok | Your own keys: the provider bills you directly, so Gate has no cost to report. |
unknown_model | Gate had no pricing for the resolved model; cost is null. |
pricing_gap | Model was found but a billing dimension has no usable rate; cost is null. |
missing_usage | The upstream returned no usage data (e.g. no include_usage on a streamed response). |
streaming_no_usage | A streamed response closed without a usage event. |
free_request | A non-token-billed modality (embeddings, images, etc.), so no cost applies. |
non_billable | A non-completion endpoint (e.g. model listing), so no tokens exchanged. |
The full set of values is also documented in Dashboard tour.
Catalog vs. provider divergence is informational
When Gate’s catalog estimate differs from the provider’s actual reported cost, it records the difference so you can see it, but it never changes your bill. A background verification sweep independently confirms what Gate billed against what the provider actually charged and flags any mismatch beyond a small tolerance. If the catalog and provider figures disagree, that is expected, it is surfaced on purpose, and the amount you paid already matches the provider. See Billing mode.
A debugging checklist
When a request behaves unexpectedly, work through this in order:
- Grab the request ID. Read
X-Gate-Request-Idfrom the response (or find the request in the dashboard) and open its full record. Nearly every question is answered there. - Check
source(orx-gate-error-source) first. It tells you whether you, Gate, the provider, or the vendor is responsible, and whether a plain retry can help. - Refused or charged? A
402/429/503/403means the request was refused before the provider call, so nothing was billed. A successful response that shows a charge-time gap (payg.debit.insufficient_balancein the audit log) means it was served and will be recovered later. - For a balance refusal, read the code.
insufficient_balance(402) means top up.payg_reserved(429) means other in-flight requests are holding funds: the gap betweenavailableCentsandbalanceCentsfrees up as they finish, so retry afterRetry-After. - For a cost question, read the cost status.
X-Gate-Cost-Calculation-Statusdistinguishes a genuine$0from an uncalculable one, and explains an absentX-Gate-Cost-Usd. - For a
no_routeerror, checksuggested_modelsandavailable_provider_types. These tell you what Gate can route to, and whether the issue is a typo or a missing provider account. - For a security block, open the message in the dashboard. The security verdict and flagged content are visible on the Messages page, where Mark false positive is also available; sensitivity is set per organization on the Policies page.
Related
- Limits and retention: rate limits, usage caps, request size, and retention.
- Audit trail: where billing events are recorded.
- Prompt-injection defense: how security evaluation works and what triggers a block.