> ## Documentation Index
> Fetch the complete documentation index at: https://docs.60db.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Judge

> CLI commands for rubric evaluation and turn extraction — score content against questions you define, and parse conversational turns into intent and entities

# Judge

Judge asks a pinned classifier a question about content you supply. It never
generates text: every answer is a probability distribution over an answer space
**you** defined, which is what makes the results averageable, sortable and
safe to feed straight into other code.

<Info>
  **Uses your existing 60db credentials.** Judge wraps the 60db Jev model behind
  the same API key as TTS, STT and Memory — there is no separate Judge key, and
  keys you already issued keep working. `60db config` and `X60DB_API_KEY` apply
  unchanged.

  The two inference commands are billed per input token. A failed run is refunded
  automatically — you are never charged for an answer you did not get. Every
  billed command prints the charge and remaining balance in a footer.
</Info>

## Commands

| Command                    | Billed | Purpose                                    |
| -------------------------- | ------ | ------------------------------------------ |
| `60db judge:evaluate`      | ✓      | Run a rubric over content                  |
| `60db judge:extract`       | ✓      | Classify a turn and pull entity spans      |
| `60db judge:models`        | —      | Model names, plus a starter rubric to copy |
| `60db judge:rubrics`       | —      | List saved rubrics                         |
| `60db judge:create-rubric` | —      | Save a rubric so it can be re-run by id    |
| `60db judge:delete-rubric` | —      | Delete a saved rubric                      |
| `60db judge:runs`          | —      | Run history, including the review queue    |
| `60db judge:run`           | —      | One run in full                            |
| `60db judge:delete-run`    | —      | Remove a run from history                  |
| `60db judge:usage`         | —      | Spend and run counts                       |
| `60db judge:health`        | —      | Upstream readiness (owner/admin)           |

## The three question types

A rubric is a JSON map of *answer key* → *question*. Each question is one of:

| Type     | Asks                   | Returns                                                                       |
| -------- | ---------------------- | ----------------------------------------------------------------------------- |
| `choice` | Pick one of my options | The chosen option + a probability for every option                            |
| `score`  | Rate on my ladder      | A **weighted position** — `1.87`, not an index — plus per-level probabilities |
| `noul`   | True or false          | A raw probability of true. No confidence field, by design                     |

```json rubric.json theme={null}
{
  "tone": {
    "type": "choice",
    "instructions": "How did the agent come across?",
    "criteria": {
      "professional": "Calm, courteous, takes ownership",
      "dismissive": "Brushes the caller off or hides behind policy",
      "unknown": "Not enough of the call to tell"
    }
  },
  "satisfaction": {
    "type": "score",
    "instructions": "How satisfied does the caller sound by the end?",
    "criteria": ["Angry", "Unhappy", "Neutral", "Satisfied"]
  },
  "resolved": {
    "type": "noul",
    "instructions": "Was the caller's problem actually solved?"
  }
}
```

<Warning>
  Always include an `unknown` option on a `choice`. Without an escape hatch the
  model has to pick one of your real labels even on off-topic content.
</Warning>

The **descriptions are what the model reads** — they are the part to tune. A
label with a vague description produces vague answers.

## Evaluate content

```bash theme={null}
60db judge:evaluate \
  --state-file call.txt \
  --rubric-file rubric.json \
  --label "support-qa"
```

**Options:**

* `-s, --state <text>` — The content to judge
* `-f, --state-file <path>` — Read the content from a file instead
* `-r, --rubric <id>` — Run a saved rubric by id
* `--rubric-file <path>` — Run an inline rubric from a JSON file
* `-m, --model <name>` — Model to use (see `judge:models`)
* `-l, --label <text>` — Tag the run for later filtering
* `--no-save` — Bill the run but keep it out of history

Pass **either** `--rubric` or `--rubric-file`, never both.

**Output:**

```
Answers  60db-decision-model-v1  ·  8ms  ·  111 input tokens  ·  2 escalated

tone  54% confident  ⚠ below the 0.85 review line
  dismissive
  ████████████░░░░░░░░░░░░   49% dismissive
  ████████░░░░░░░░░░░░░░░░   35% unknown
  ████░░░░░░░░░░░░░░░░░░░░   16% professional

satisfaction  48% confident  ⚠ below the 0.85 review line
  1.87 of 3
  ██████████░░░░░░░░░░░░░░   43% 3 · Satisfied
  █████████░░░░░░░░░░░░░░░   39% 1 · Unhappy

resolved  probability
  28% true
  ███████░░░░░░░░░░░░░░░░░

⚠ Lowest confidence 48% — worth a second look.
ℹ run 6d4f97ec-48f3-42be-80b2-6a1cc1c7b5ab
  charged $0.00000161  ·  balance $761.512548  ·  tx a5797951...
```

The distribution is the answer. A choice with `0.86` on the winner and one with
`0.34` read identically if you only print the label.

## Extract from a turn

Classify one conversational turn and pull the values out of it:

```bash theme={null}
60db judge:extract \
  --text "move my 10am appointment to friday" \
  --schema-file schema.json
```

```json schema.json theme={null}
{
  "intents":    { "booking": "Wants to arrange, change or cancel a booking" },
  "operations": { "reschedule": "Move an existing booking to a new time" },
  "entities":   { "date": "A calendar date", "time": "A clock time" }
}
```

**Options:**

* `-t, --text <text>` — The turn (required, max 2,048 characters)
* `--schema-file <path>` — JSON file with the label schema (required)
* `--profile <name>` — `generic` (default) or `medical`
* `--budget <ms>` — Total request budget, 1–30000
* `--no-save` — Bill the run but keep it out of history

Adding a `responsePaths` map to the schema selects the v2 model, which also
returns how the agent should reply.

**Output:**

```
v1  4ms
  intent        booking  94%
  operation     reschedule  61%

Entities  (offsets are Unicode code points)
  date          friday  [28,34)  80%
  time          10am  [8,12)  80%

  charged $0.00000054  ·  balance $761.512547  ·  tx ad2898f5...
```

<Warning>
  Entity `start`/`end` are **Unicode code point** offsets, not byte or UTF-16
  indices. An emoji earlier in the turn shifts every naive slice after it.
</Warning>

## Save and re-run a rubric

Build the rubric once, then run it by id from then on:

```bash theme={null}
60db judge:create-rubric \
  --name "Call QA" \
  --file rubric.json \
  --description "Tone, satisfaction and outcome on inbound calls"

# → id 8b0da171-352b-4af3-8790-a5efb18eaac9

60db judge:evaluate --rubric 8b0da171-... --state-file call.txt
```

**Options:**

* `-n, --name <name>` — Unique per workspace
* `-f, --file <path>` — JSON file containing the questions map
* `-d, --description <text>` — What the rubric is for
* `-m, --model <name>` — Default model for this rubric
* `--shared` — Publish to the whole workspace (owner/admin only)

Every run of a saved rubric files itself under it, so
`60db judge:runs --rubric <id>` gives you the whole series.

```bash theme={null}
60db judge:rubrics                    # list
60db judge:delete-rubric --id <id>    # delete (past runs stay readable)
```

## The review queue

The single most useful command. Anything the judge was under 85% sure about —
the upstream's own escalation line:

```bash theme={null}
60db judge:runs --needs-review
```

```
┌──────────┬──────────┬────────────────────────────────────────┬──────┬─────────────┐
│ id       │ kind     │ summary                                │ conf │ cost        │
├──────────┼──────────┼────────────────────────────────────────┼──────┼─────────────┤
│ 6d4f97ec │ evaluate │ tone: dismissive · satisfaction: 1.87  │ 48%  │ $0.00000161 │
│ 03a928b0 │ extract  │ intent: booking · entities: 2          │ 61%  │ $0.00000054 │
└──────────┴──────────┴────────────────────────────────────────┴──────┴─────────────┘
```

**Options:**

* `--needs-review` — Only low-confidence runs
* `--kind <kind>` — `evaluate` or `extract`
* `-r, --rubric <id>` — Only runs of one rubric
* `--limit <n>` / `--offset <n>` — Paging (max 100 per page)

```bash theme={null}
60db judge:run --id <id>          # one run in full, with what was sent and returned
60db judge:delete-run --id <id>   # remove it from history
```

<Info>
  History is private to whoever made the run. Workspace owners and admins see
  everyone's — a run stores the content it judged.
</Info>

## Models and spend

```bash theme={null}
60db judge:models   # authoritative list, plus a starter rubric to copy
60db judge:usage    # --period current_month | last_30_days | all_time
60db judge:health   # upstream readiness + circuit-breaker state (owner/admin)
```

Do not hardcode model names — `judge:models` is the source of truth and
changes when the deployment does.

## Pricing

Billed on **input tokens**, the model's own unit.

The upstream builds one row per question, each carrying the shared content
again, so a 10-question rubric encodes the transcript **ten times**. Question
count and content length both land in the token total on their own.

| Example                                   | Per 1,000 runs |
| ----------------------------------------- | -------------- |
| Ticket triage — 2 questions, short ticket | `$0.0007`      |
| Call QA — 3 questions, 3-minute call      | `$0.0150`      |
| Call QA — 3 questions, 15-minute call     | `$0.0750`      |
| Deep audit — 10 questions, long document  | `$0.5000`      |
| Extract — one turn                        | `$0.0004`      |

Keep rubrics tight when the content is long — that last row is the same model,
just asked ten questions about twenty thousand characters.

Context is **8K tokens per question**, not per request. An over-long piece of
content is rejected with a `400` before anything is charged.

## Agent-friendly JSON

Every command takes the global `--json` flag for scripting:

```bash theme={null}
60db judge:evaluate --rubric <id> --state-file call.txt --json \
  | jq '.data.answers.tone.choice'
```

## Handling errors

| Exit condition               | Meaning                                                            |
| ---------------------------- | ------------------------------------------------------------------ |
| `400`                        | Your rubric or schema is invalid — the message says which question |
| `402` `INSUFFICIENT_CREDITS` | Top up; the shortfall is printed                                   |
| `429`                        | The judge's queue is full — retry shortly                          |
| `503`                        | Judge service unavailable                                          |

All of these refund automatically. Nothing is charged for a run that failed.

## Related

* [JavaScript SDK](/sdks/javascript) — the same operations as `client.judge.*`
* [Python SDK](/sdks/python) — the same operations as `client.judge.*`
