Build
Model routing and fallbacks.
How a request picks a route, what changes its price, and what happens when every route fails.
On this page
Append a suffix to any model id, or send the same ask as a provider preference in the request body:
<model>:cheapest <model>:fastest <model>:private
{
"model": "<model>",
"messages": [ ... ],
"provider": { "private": true }
}:cheapest
Or provider.sort: "price". Eligible routes ordered by upstream list price. The model's own sheet price still applies, whichever route serves.
:fastest
Or provider.sort: "latency". Ordered by the latency this gateway has observed on each route; a route with no observations goes last.
:private
Or provider.private: true. Attested-hardware routes only, charged the model's separately published private rate. See Private inference.
A response and its receipt always name the model without the suffix. The :batch suffix is not a routing order. It submits an asynchronous job instead; see Batch.
Send a provider object in the request body; the gateway honours these fields and ignores the rest:
| Field | Type | Effect |
|---|---|---|
zdr | Boolean | Zero data retention: only routes whose published policy keeps no request content. |
data_collection | "allow" or "deny" | "deny": only routes whose policy says content is not used for training. |
private | Boolean | Only attested-hardware routes, at the model's private price. |
region | String | Only routes whose upstream publishes that processing region, as a lowercase code such as "us". |
only, ignore | Array of provider ids | Providers to allow or exclude, as GET /v1/pricing names them per route. |
sort | "price", "latency" or "throughput" | The order routes are tried in. "throughput" is the same as "latency". |
allow_fallbacks | Boolean | false: the first route only. |
max_price | { prompt, completion } | A price ceiling. See Price ceiling. |
Ask for a policy
curl https://api.usdf.fi/v1/chat/completions \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model>",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 256,
"provider": {"zdr": true, "data_collection": "deny", "sort": "price"}
}'A route whose policy is not verified is never used for a request that asks. Each model's policy_asks in GET /v1/pricing says whether such a request has a route for it. A route with no published region is never used for a request that names one. On a multipart endpoint, send provider as a JSON string form field.
provider.max_price sets the most a request may be charged per 1M tokens, in USD: prompt for input, completion for output.
It is compared with every rate the request can be billed at. The input, cached-input and cache-write rates count against prompt; the output rate against completion. Past a model's long-context threshold, its long-context rates count instead.
Set a price ceiling
curl https://api.usdf.fi/v1/chat/completions \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model>",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 256,
"provider": {"max_price": {"prompt": 0.5, "completion": 1.5}}
}'A model above the ceiling is refused with 404 no_route_for_policy, naming the rates above it. Nothing is held. The ceiling applies to token-priced endpoints only: chat, completions, embeddings and rerank. Speech, transcription, images and video refuse it with 400 max_price_unsupported.
By default, a request tries its eligible routes in order until one of them serves it. Set allow_fallbacks to false and it runs on the first route only: if that one fails, the request fails instead of moving on to the next.
{
"model": "<model>",
"messages": [ ... ],
"provider": { "allow_fallbacks": false }
}A request with no eligible route is refused before anything is tried, with 404 no_route_for_policy. Nothing is held. When every eligible route is tried and fails, the answer is 502 upstream_unavailable, or 504 upstream_timeout when no provider answered before the deadline. The hold is released first.
Set any of these once for every key on the account with PATCH /v1/account/settings, under the session: zdr, data_collection, sort, provider_allow, provider_deny, allow_fallbacks and max_price. On the dashboard: Settings.
Set account defaults
curl -X PATCH https://api.usdf.fi/v1/account/settings \
-H "Authorization: Bearer $SESSION" \
-H "Content-Type: application/json" \
-d '{
"zdr": true,
"sort": "latency",
"provider_deny": ["<provider id>"],
"allow_fallbacks": true,
"max_price": {"input": 0.5, "output": 1.5}
}'In the settings, max_price takes input and output. The gateway applies provider_allow and provider_deny as the request's own only and ignore. A request's own provider field overrides the routing ones (sort, the provider lists, allow_fallbacks, max_price), and an empty only or ignore list clears the account's. The account's zdr: true and data_collection: "deny" hold for every request: a request can add them, never turn the account's off. Defaults apply to requests made with a key; a request paid per request carries only its own asks.
Paid with a key
A request that serves nothing is not charged. One that fails after its hold is taken is marked failed_refunded: cost zero, and the hold back in the balance. The row gives the reason: no provider produced anything, the stream ended empty, or the gateway itself failed. It stays in the usage log at zero cost, so a refund is as checkable as a charge.
One exception: some providers bill reasoning they do not stream. On such a model, a stream you leave, or a call cut at its deadline, after the provider accepted it is billed its held output budget, as the provider bills it. The receipt says partial.
Paid with x402 or MPP
If the gateway or the provider fails after payment, before any output, the full amount is refunded in the asset it was paid in. If the provider rejects the request itself as invalid, the minimum charge is kept as the settlement fee and the rest is refunded.
A caller who disconnects from a stream before any output is not refunded. An MCP call cut at its deadline after the provider accepted it keeps the settlement fee. On a model whose provider bills reasoning it does not stream, the payment stands once the provider has accepted the request. See Refunds and refusals.