> ## Documentation Index
> Fetch the complete documentation index at: https://docs.60db.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Judge

> Score content against your own rubric, and classify conversational turns — a classifier, not a generator

## Overview

60db's **Judge** answers questions *you* define about a piece of content, and returns a probability distribution for every answer. It never writes prose.

That is the whole point. Ask a language model "is this caller angry?" and you get a paragraph you then have to parse. Ask Judge and you get `0.86` on `dismissive` — a number you can sort by, threshold on, and compare across a hundred thousand runs.

<CardGroup cols={2}>
  <Card title="Your Labels, Not Ours" icon="tags">
    You supply the options and the descriptions; the model only picks between them
  </Card>

  <Card title="Comparable Output" icon="chart-simple">
    Every answer is a distribution over a fixed answer space, so runs sort and diff
  </Card>

  <Card title="Up to 32 Questions" icon="list-check">
    Each asked independently against the same content, in one request
  </Card>

  <Card title="Saved Rubrics" icon="clipboard-check">
    Store a question set once, then run it by `rubric_id` forever
  </Card>

  <Card title="Confidence & Review Queue" icon="user-check">
    Below `0.85`, route it to a person — filter history on `needs_review`
  </Card>

  <Card title="Turn Extraction" icon="wand-magic-sparkles">
    Intent, operation and entity spans from a single conversational turn
  </Card>
</CardGroup>

<Note>
  **Your existing 60db API key already works.** Judge wraps the 60db Jev model behind the same platform credential as TTS, STT and Memory — there is no separate Judge key to issue. Keys created from now on carry a `judge` scope; keys issued earlier are accepted on their `slm` scope.
</Note>

## Core concepts

### Questions, not prompts

A Judge request has two halves: the **`state`** (the content — a transcript, a ticket, a document, any JSON structure), and the **`questions`** asked about it. Each question is keyed by a name of your choosing, which is how you read the answer back. The key is never sent to the model.

Questions are asked **independently**. Every one sees the same `state` and its own criteria, and nothing else — they are not a conversation, and question 3 cannot see what question 1 answered.

<Tip>
  The **descriptions are what the model reads** — they are the thing to tune, not a prompt. A label with a vague description produces a vague answer.
</Tip>

### The three question types

| Type | Question it answers | You supply | You get back |
| - | - | - | - |
| **`choice`** | Which one of these? | Map of option name → description, max 255 | `choice`, `probabilities` over every option, `confidence` |
| **`score`** | Where on this ladder? | Ordered array, **lowest first**, max 10 levels | `score` (weighted position, e.g. `1.87`), `legend`, `probabilities`, `confidence` |
| **`noul`** | True or false? | Optional `{ "true": "...", "false": "..." }` | `noul` — a raw probability of true |

<Warning>
  Always include an `unknown` option on a `choice`. Without an escape hatch the model must pick one of your real labels even on off-topic content.
</Warning>

Two results are easy to misread:

* **`score` is a weighted position, not an index.** `1.87` on a four-rung ladder means the mass sits between rungs 1 and 2 — it is not "rung 1" rounded.
* **`noul` has deliberately no `confidence` field.** The number *is* the probability of true; it is not a calibrated correctness estimate.

### Confidence and escalation

Every run reports `min_confidence` — the lowest confidence across its answers (`null` when every question was `noul`). **Below `0.85`, a human should look.** That is the service's own escalation threshold, and it is the filter the review queue is built on.

`escalated_questions` tells you how many answers the service already refined with its second-stage model because confidence was low.

## Evaluating content

<Tabs>
  <Tab title="JavaScript">
    ```javascript theme={null}
    import { SixtyDBClient } from '60db-service';

    const client = new SixtyDBClient('your-api-key');

    const run = await client.judge.evaluate({
      state: callTranscript,
      questions: {
        tone: {
          type: 'choice',
          instructions: 'How did the agent come across?',
          criteria: {
            professional: 'Calm, courteous, takes ownership',
            dismissive: 'Brushes the caller off',
            unknown: 'Not enough of the call to tell',   // always add an escape hatch
          },
        },
        satisfaction: { type: 'score', criteria: ['Angry', 'Neutral', 'Happy'] },
        resolved: { type: 'noul', instructions: 'Was it solved?' },
      },
      label: 'support-call-qa',   // groups runs in history
    });

    run.answers.tone.choice           // 'dismissive'
    run.answers.tone.probabilities    // { dismissive: 0.86, professional: 0.11, unknown: 0.03 }
    run.answers.tone.confidence       // 0.79
    run.answers.satisfaction.score    // 1.87 — weighted position, not an index
    run.answers.resolved.noul         // 0.28 — probability of true, NOT a confidence
    run.min_confidence                // under 0.85 → send it to a person
    run.credits_charged               // 0.00000161
    ```
  </Tab>

  <Tab title="Python">
    ```python theme={null}
    from sixtydb import SixtyDBClient

    client = SixtyDBClient("your-api-key")

    run = client.judge.evaluate(
        state=call_transcript,
        questions={
            "tone": {
                "type": "choice",
                "instructions": "How did the agent come across?",
                "criteria": {
                    "professional": "Calm, courteous, takes ownership",
                    "dismissive": "Brushes the caller off",
                    "unknown": "Not enough of the call to tell",   # always add an escape hatch
                },
            },
            "satisfaction": {"type": "score", "criteria": ["Angry", "Neutral", "Happy"]},
            "resolved": {"type": "noul", "instructions": "Was it solved?"},
        },
        label="support-call-qa",
    )

    run["answers"]["tone"]["choice"]          # 'dismissive'
    run["answers"]["tone"]["probabilities"]   # {'dismissive': 0.86, 'professional': 0.11, ...}
    run["answers"]["satisfaction"]["score"]   # 1.87 — weighted, not an index
    run["answers"]["resolved"]["noul"]        # 0.28 — probability of true, not a confidence
    run["min_confidence"]                     # under 0.85 → send to a human
    run["credits_charged"]
    ```
  </Tab>

  <Tab title="cURL">
    ```bash theme={null}
    curl -X POST https://api.60db.ai/judge/evaluate \
      -H "Authorization: Bearer your-api-key" \
      -H "Content-Type: application/json" \
      -d '{
        "state": "Agent: I can'\''t refund that, it'\''s outside the window.\nCaller: This is the third time I'\''ve called about this.",
        "questions": {
          "tone": {
            "type": "choice",
            "instructions": "How did the agent come across?",
            "criteria": {
              "professional": "Calm, courteous, takes ownership",
              "dismissive": "Brushes the caller off",
              "unknown": "Not enough of the call to tell"
            }
          },
          "satisfaction": {
            "type": "score",
            "criteria": ["Angry", "Unhappy", "Neutral", "Satisfied"]
          },
          "resolved": { "type": "noul", "instructions": "Was the problem solved?" }
        },
        "label": "support-call-qa"
      }'
    ```
  </Tab>

  <Tab title="CLI">
    ```bash theme={null}
    # rubric.json holds the questions map
    60db judge:evaluate --rubric-file rubric.json --state-file call.txt

    60db judge:models   # authoritative model list, plus a starter rubric to copy
    ```
  </Tab>
</Tabs>

Alongside `answers`, every run returns a **`summary`** — one headline value per question (the chosen option, the score, or the probability). That is the shape to store on your own record when you do not need the full distribution.

## Saved rubrics

Save a question set once and run it by id from then on. Runs of a saved rubric file themselves under it, so listing by `rubric_id` returns the whole series — which is what makes a rubric a trend line rather than a one-off.

<Tabs>
  <Tab title="JavaScript">
    ```javascript theme={null}
    const rubric = await client.judge.rubrics.create({
      name: 'Call QA',                 // unique per workspace
      description: 'Tone, satisfaction and outcome on inbound calls',
      questions,
      shared: true,                    // workspace-wide; owner/admin only
    });

    await client.judge.evaluate({ state: transcript, rubric_id: rubric.id });

    await client.judge.rubrics.list();
    await client.judge.rubrics.get(id);
    await client.judge.rubrics.update(id, { description: 'updated' });
    await client.judge.rubrics.delete(id);   // past runs stay readable
    ```
  </Tab>

  <Tab title="Python">
    ```python theme={null}
    rubric = client.judge.rubrics.create(
        name="Call QA",                 # unique per workspace
        description="Tone, satisfaction and outcome on inbound calls",
        questions=questions,
        shared=True,                    # workspace-wide; owner/admin only
    )

    client.judge.evaluate(state=transcript, rubric_id=rubric["id"])

    client.judge.rubrics.list()
    client.judge.rubrics.get(rubric_id)
    client.judge.rubrics.update(rubric_id, description="updated")
    client.judge.rubrics.delete(rubric_id)   # past runs stay readable
    ```
  </Tab>

  <Tab title="CLI">
    ```bash theme={null}
    60db judge:create-rubric --name "Call QA" --file rubric.json
    60db judge:evaluate --rubric 8b0da171-... --state-file call.txt

    60db judge:rubrics                    # list
    60db judge:delete-rubric --id <id>    # delete (past runs stay readable)
    ```
  </Tab>
</Tabs>

Pass **either** `questions` or `rubric_id`, never both. Saved rubrics are re-validated on every run — a rubric is not trusted just because it was valid when stored.

## Extracting from a turn

`extract` is the understanding layer between speech-to-text and a reply: what did the speaker just ask for, and what values did they give me. It labels one turn with an **intent** and an **operation**, and pulls **entity spans** out with their exact positions.

<Tabs>
  <Tab title="JavaScript">
    ```javascript theme={null}
    const turn = await client.judge.extract({
      text: 'move my 10am appointment to friday',
      schema: {
        intents:    { booking: 'Wants to arrange or change a booking' },
        operations: { reschedule: 'Move an existing booking' },
        entities:   { date: 'A calendar date', time: 'A clock time' },
        // responsePaths: { confirm: 'Read the change back' }   ← selects the v2 model
      },
    });

    turn.intent.label       // 'booking'
    turn.operation.label    // 'reschedule'
    turn.entities           // [{ label: 'date', text: 'friday', start: 28, end: 34, confidence: 0.8 }]
    ```
  </Tab>

  <Tab title="Python">
    ```python theme={null}
    turn = client.judge.extract(
        text="move my 10am appointment to friday",
        schema={
            "intents": {"booking": "Wants to arrange or change a booking"},
            "operations": {"reschedule": "Move an existing booking"},
            "entities": {"date": "A calendar date", "time": "A clock time"},
            # "responsePaths": {"confirm": "Read the change back"}   ← selects the v2 model
        },
    )

    turn["intent"]["label"]      # 'booking'
    turn["entities"]             # [{'label': 'date', 'text': 'friday', 'start': 28, 'end': 34, ...}]
    ```
  </Tab>

  <Tab title="CLI">
    ```bash theme={null}
    60db judge:extract --text "move my 10am appointment to friday" --schema-file schema.json
    ```
  </Tab>
</Tabs>

<Warning>
  `start`/`end` are **Unicode code point** offsets, not UTF-16 indices. Slice with `Array.from(text).slice(start, end).join('')` in JavaScript, or `"".join(list(text)[start:end])` in Python — the naive `text.slice()` / `text[start:end]` lands in the wrong place once the turn contains an emoji.
</Warning>

Supplying `schema.responsePaths` **selects the v2 model**, which also returns a `response_path` and tracks conversation state. `text` is capped at **2,048 Unicode code points**; the required `unknown` label is added for you if you leave it out of a classification map.

## The review queue

Rather than re-reading everything the model was already sure about, list only what it wasn't:

<Tabs>
  <Tab title="JavaScript">
    ```javascript theme={null}
    const { runs, pagination } = await client.judge.runs.list({
      needs_review: true,     // only runs whose min_confidence fell below 0.85
      limit: 20,
    });

    await client.judge.runs.get(id);      // one run, with what was sent and returned
    await client.judge.runs.delete(id);
    ```
  </Tab>

  <Tab title="Python">
    ```python theme={null}
    page = client.judge.runs.list(needs_review=True, limit=20)
    page["runs"]         # only runs the judge was under 85% sure about
    page["pagination"]   # {'total': ..., 'limit': ..., 'has_more': ...}

    client.judge.runs.get(run_id)      # one run, with what was sent and returned
    client.judge.runs.delete(run_id)
    ```
  </Tab>

  <Tab title="CLI">
    ```bash theme={null}
    60db judge:runs --needs-review
    60db judge:run --id <id>          # one run in full, with what was sent and returned
    60db judge:delete-run --id <id>   # remove it from history
    ```
  </Tab>
</Tabs>

History also filters by `kind` (`evaluate` or `extract`), by `rubric_id` for one rubric's whole series, and by your own `confidence_below` threshold. Pass `save: false` on a run to bill it but keep it out of history entirely.

## Models

Fetch the model list; don't hardcode it. `jev-latest` is an alias that resolves server-side to a versioned id, and that versioned id is what the evaluate response reports back as `model`. Responses are cached for 5 minutes, so polling to populate a picker is free.

```javascript theme={null}
await client.judge.models();                // authoritative list — don't hardcode
await client.judge.usage('current_month');  // spend and run counts
await client.judge.health();                // upstream readiness (owner/admin)
```

## Pricing

Judge is pay-as-you-go on **input tokens only** — $0.010 per 1M. Output is structurally `$0\`, because a classifier emits no text. It bills to the same workspace wallet as every other 60db service.

The one thing that will surprise you:

```
input_tokens ≈ questions × content + every question's own text
```

The service builds **one row per question**, and every row carries the shared `state` again. Ten questions about a 20,000-character document encode that document ten times — so **keep rubrics tight when the content is long**. Context is **8K tokens per question**, not per request; over-long content is rejected with a `400` before anything is charged.

| Scenario | Tokens | Per 1,000 runs |
| - | - | - |
| Ticket triage — 2 questions, short ticket | \~70 | **\$0.0007** |
| Call QA — 3 questions, 3-minute call | \~1,500 | **\$0.0150** |
| Call QA — 3 questions, 15-minute call | \~7,500 | **\$0.0750** |
| Deep audit — 10 questions, long document | \~50,000 | **\$0.5000** |
| Extract — one conversational turn | \~36 | **\$0.0004** |

<Info>
  **A failed run costs nothing.** The wallet is debited before the call and refunded automatically if anything goes wrong — a busy queue (`429`), an unavailable service (`503`), a rejected rubric (`400`). Both entries appear in the ledger, netting to zero. See [Judge pricing](/api-reference/judge/pricing) for the full reconciliation story.
</Info>

<Tip>
  Scoring every call rather than a 2% sample is the point of pricing like this. Ten thousand three-minute calls a day costs roughly **\$4.50/month**.
</Tip>

## Errors

| Status | Meaning |
| - | - |
| `400` | Invalid rubric — the message names the offending question. Also returned over 32 KiB or the 8K-token context. |
| `402` | Insufficient credits; `details.shortfall` says how much is missing. |
| `403` | Your role cannot run the judge, an API key lacks the `judge` scope, or `shared: true` needs owner/admin. |
| `404` | Unknown `rubric_id`, or it belongs to another user. |
| `409` | A rubric with that name already exists in this workspace. |
| `429` | The inference queue is full — back off and retry. |
| `502` / `503` | The judge service is misconfigured or unavailable. |

## API Reference

<CardGroup cols={3}>
  <Card title="Evaluate" icon="scale-balanced" href="/api-reference/judge/evaluate">
    Score content against a rubric
  </Card>

  <Card title="Extract" icon="wand-magic-sparkles" href="/api-reference/judge/extract">
    Intent, operation and entity spans
  </Card>

  <Card title="Rubrics" icon="clipboard-list" href="/api-reference/judge/list-rubrics">
    Save, list, update and delete
  </Card>

  <Card title="Runs" icon="clock-rotate-left" href="/api-reference/judge/list-runs">
    History and the review queue
  </Card>

  <Card title="Models" icon="microchip" href="/api-reference/judge/models">
    Model names this deployment serves
  </Card>

  <Card title="Usage" icon="chart-line" href="/api-reference/judge/usage">
    Spend and run counts
  </Card>

  <Card title="Pricing" icon="tag" href="/api-reference/judge/pricing">
    Rates, token maths and refunds
  </Card>

  <Card title="CLI Commands" icon="terminal" href="/cli-reference/commands/judge">
    `60db judge:*` from the terminal
  </Card>

  <Card title="MCP Tools" icon="plug" href="/mcp-server/judge">
    Judge from Claude and other agents
  </Card>
</CardGroup>
