Skip to main content
60db Judge is pay-as-you-go, billed on the model’s own unit: input tokens. Every charge is logged in the same wallet audit trail as the rest of the platform and readable via GET /judge/usage.
Output tokens are structurally zero — Judge is a classifier and emits no text. You are only ever billed for what goes in.
Judge bills to the same workspace wallet as every other 60db service, through your existing API key. Nothing separate to set up or top up.

Rates

There is no separate per-question or per-call fee. The token count already accounts for everything, because of how the model is fed.

Why a 10-question rubric costs ten times the transcript

The upstream builds one row per question, and every row carries the shared content again. Its own specification puts it plainly: “shared state is counted again for each question.” So the token total is:
Not content + questions. Ten questions about a 20,000-character document encode that document ten times. This is the single most important thing to know when designing a rubric: keep rubrics tight when the content is long.

What that works out to

Note the jump on the last rubric row. That is the same model — just asked ten questions about twenty thousand characters.
Scoring every call rather than a 2% sample is the point of pricing like this. Ten thousand three-minute calls a day costs roughly $4.50/month.

Failed runs cost nothing

The wallet is debited before the call, then refunded automatically if anything goes wrong — a busy queue, an unavailable model, a rubric the service rejects. Both the charge and its refund appear in the ledger, netting to zero. You are never charged for an answer you did not receive.

Estimation and reconciliation

Because the wallet is debited before the call, the token count at charge time is an estimate. The upstream’s real usage.input_tokens comes back with the answer and is stored on the run, but it never re-prices a charge the caller already saw. The estimator is script-aware: roughly 4 characters per token for Latin text, but close to 1 for Devanagari, CJK and other dense scripts. Without that, Hindi content would be under-billed about fourfold — and, worse, could slip past the context check into a failure you were charged for.

Context limit

The model’s context is 8,000 tokens per question, not per request. A rubric with many questions over a short piece of content is fine; a single piece of content longer than the window is not, however few questions you ask. Requests that cannot fit are rejected with 400 before anything is charged.

Running out

When the wallet cannot cover a run, the request returns 402 before any inference happens:
Unbilled endpoints — rubrics, history, usage, models — keep working on an empty wallet.

Watching spend

Every billed response carries three headers: They are reported to 8 decimal places — token-priced charges routinely land below a microcent. GET /judge/usage aggregates the same ledger by period, with refunds already netted off.