> ## Documentation Index
> Fetch the complete documentation index at: https://docs.60db.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Pricing & Billing

> Pay-as-you-go pricing for 60db Judge — billed per input token, with automatic refunds on failure

60db Judge is **pay-as-you-go**, billed on the model's own unit: **input tokens**. Every charge is logged in the same wallet audit trail as the rest of the platform and readable via [`GET /judge/usage`](/api-reference/judge/usage).

<Info>
  Output tokens are structurally **zero** — Judge is a classifier and emits no text. You are only ever billed for what goes in.
</Info>

<Note>
  Judge bills to the **same workspace wallet** as every other 60db service, through your existing API key. Nothing separate to set up or top up.
</Note>

## Rates

| Direction  | Rate                                            |
| ---------- | ----------------------------------------------- |
| **Input**  | **\$0.010** per 1M tokens                       |
| **Output** | **\$0** — a classifier generates no text tokens |

There is no separate per-question or per-call fee. The token count already accounts for everything, because of how the model is fed.

## Why a 10-question rubric costs ten times the transcript

The upstream builds **one row per question**, and every row carries the shared content again. Its own specification puts it plainly: *"shared state is counted again for each question."*

So the token total is:

```
input_tokens ≈ questions × content + every question's own text
```

Not `content + questions`. Ten questions about a 20,000-character document encode that document **ten times**. This is the single most important thing to know when designing a rubric: **keep rubrics tight when the content is long.**

## What that works out to

| Scenario                                  | Tokens   | Per 1,000 runs |
| ----------------------------------------- | -------- | -------------- |
| Ticket triage — 2 questions, short ticket | \~70     | **\$0.0007**   |
| Call QA — 3 questions, 3-minute call      | \~1,500  | **\$0.0150**   |
| Call QA — 3 questions, 15-minute call     | \~7,500  | **\$0.0750**   |
| Deep audit — 10 questions, long document  | \~50,000 | **\$0.5000**   |
| Extract — one conversational turn         | \~36     | **\$0.0004**   |

Note the jump on the last rubric row. That is the same model — just asked ten questions about twenty thousand characters.

<Tip>
  Scoring every call rather than a 2% sample is the point of pricing like this. Ten thousand three-minute calls a day costs roughly **\$4.50/month**.
</Tip>

## Failed runs cost nothing

The wallet is debited **before** the call, then refunded automatically if anything goes wrong — a busy queue, an unavailable model, a rubric the service rejects. Both the charge and its refund appear in the ledger, netting to zero.

You are never charged for an answer you did not receive.

| Outcome                    | Charged                                 |
| -------------------------- | --------------------------------------- |
| Answers returned           | ✅ yes                                   |
| `429` queue full           | refunded                                |
| `503` service unavailable  | refunded                                |
| `400` invalid rubric       | refunded                                |
| `402` insufficient credits | never debited — the run does not happen |

## Estimation and reconciliation

Because the wallet is debited before the call, the token count at charge time is an **estimate**. The upstream's real `usage.input_tokens` comes back with the answer and is stored on the run, but it never re-prices a charge the caller already saw.

The estimator is **script-aware**: roughly 4 characters per token for Latin text, but close to 1 for Devanagari, CJK and other dense scripts. Without that, Hindi content would be under-billed about fourfold — and, worse, could slip past the context check into a failure you were charged for.

## Context limit

The model's context is **8,000 tokens per question**, not per request. A rubric with many questions over a short piece of content is fine; a single piece of content longer than the window is not, however few questions you ask.

Requests that cannot fit are rejected with `400` **before anything is charged**.

## Running out

When the wallet cannot cover a run, the request returns `402` before any inference happens:

```json theme={null}
{
  "success": false,
  "error_code": "INSUFFICIENT_CREDITS",
  "details": { "required": 0.00000161, "available": 0.0, "shortfall": 0.00000161 }
}
```

Unbilled endpoints — rubrics, history, usage, models — keep working on an empty wallet.

## Watching spend

Every billed response carries three headers:

| Header             | Meaning                           |
| ------------------ | --------------------------------- |
| `x-credit-charged` | What this request cost            |
| `x-credit-balance` | Wallet balance afterwards         |
| `x-billing-tx`     | Ledger row id, for reconciliation |

They are reported to **8 decimal places** — token-priced charges routinely land below a microcent.

[`GET /judge/usage`](/api-reference/judge/usage) aggregates the same ledger by period, with refunds already netted off.
