> ## Documentation Index
> Fetch the complete documentation index at: https://docs.60db.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Judge Tools

> MCP tools for rubric evaluation and turn extraction — score content against questions you define, and parse conversational turns into intent and entities

# Judge Tools

Judge asks a pinned classifier a question about content you supply. It is a
**classifier, not a generator**: it never writes prose, and every answer is a
probability distribution over an answer space *you* defined.

That matters for an agent. A language model asked "is this caller angry?"
returns a paragraph you have to parse and cannot average. Judge returns
`0.86` on `dismissive` — a number you can threshold, sort and chart across
thousands of runs.

<Info>
  **No extra credentials.** Judge uses the same 60db API key the rest of the MCP
  server already authenticates with — `SIXTYDB_API_KEY` as configured, nothing
  new to add.

  The two inference tools are billed per input token. A failed run is refunded
  automatically — nothing is charged for an answer you did not get.
</Info>

## Tools

| Tool                          | Billed | Purpose                            |
| ----------------------------- | ------ | ---------------------------------- |
| `sixtydb_judge_evaluate`      | ✓      | Score content against a rubric     |
| `sixtydb_judge_extract`       | ✓      | Classify a turn, pull entity spans |
| `sixtydb_judge_list_models`   | —      | Model names this deployment serves |
| `sixtydb_judge_list_rubrics`  | —      | Saved rubrics you can run          |
| `sixtydb_judge_create_rubric` | —      | Save a rubric for reuse            |
| `sixtydb_judge_delete_rubric` | —      | Delete a saved rubric              |
| `sixtydb_judge_list_runs`     | —      | History, and the review queue      |
| `sixtydb_judge_get_run`       | —      | One run in full                    |
| `sixtydb_judge_delete_run`    | —      | Remove a run from history          |
| `sixtydb_judge_usage`         | —      | Spend and run counts               |
| `sixtydb_judge_health`        | —      | Upstream readiness (owner/admin)   |

## The three question types

A rubric is a map of *answer key* → *question*. The shape of `criteria`
changes with the type:

| Type     | Asks                   | Returns                                                       |
| -------- | ---------------------- | ------------------------------------------------------------- |
| `choice` | Pick one of my options | The winner, plus a probability for **every** option           |
| `score`  | Rate on my ladder      | A probability-weighted position — `1.87`, **not** an index    |
| `noul`   | True or false          | A raw probability of true. **No confidence field**, by design |

<Warning>
  Always include an `unknown` option on a `choice`. Without an escape hatch the
  model must pick one of your real labels even on off-topic content.
</Warning>

The **descriptions are what the model reads**. A label with a vague
description produces a vague answer — that is the part to tune, not the
prompt.

## sixtydb\_judge\_evaluate

```json theme={null}
{
  "state": "Agent: I can't refund that. Caller: This is the third time I've called.",
  "questions": {
    "tone": {
      "type": "choice",
      "instructions": "How did the agent come across?",
      "criteria": {
        "professional": "Calm, courteous, takes ownership",
        "dismissive": "Brushes the caller off or hides behind policy",
        "unknown": "Not enough of the call to tell"
      }
    },
    "satisfaction": { "type": "score", "criteria": ["Angry", "Unhappy", "Neutral", "Satisfied"] },
    "resolved": { "type": "noul", "instructions": "Was the problem solved?" }
  }
}
```

**Parameters:**

* `state` (required) — the content every question is asked about. Text or any JSON. Each question sees this and nothing else; they are independent, not a conversation.
* `questions` — map of answer key → question, max 32. **Or** `rubric_id` to run a saved one. Never both.
* `model`, `label`, `save` (`false` bills without storing history)

**Returns** one answer per question with its full distribution, the lowest
confidence across the set, token usage and the charge.

## sixtydb\_judge\_extract

The understanding layer for a voice or chat agent — what did they just ask
for, and what values did they give me.

```json theme={null}
{
  "text": "can you move my 10am appointment to friday",
  "schema": {
    "intents":    { "booking": "Wants to arrange or change a booking" },
    "operations": { "reschedule": "Move an existing booking to a new time" },
    "entities":   { "date": "A calendar date", "time": "A clock time" }
  }
}
```

Adding a `responsePaths` map selects the **v2** model, which also returns how
the agent should reply. The mandatory `unknown` label is added server-side if
you leave it out.

<Warning>
  Entity `start`/`end` are **Unicode code point** offsets, not UTF-16 indices.
  Slice with `Array.from(text).slice(start, end)` — a plain `text.slice()`
  lands in the wrong place as soon as the turn contains an emoji.
</Warning>

## The review queue

The single most valuable tool for an agent workflow:

```json theme={null}
{ "needs_review": true, "limit": 20 }
```

`sixtydb_judge_list_runs` with `needs_review` returns only runs where the
judge's lowest confidence fell below **0.85** — the upstream's own escalation
line. That is the queue a person should work through, instead of re-reading
everything the model was already confident about.

Visibility: you see your own runs; workspace owners and admins see everyone's.
A run stores the content it judged, so history is not workspace-public by
default.

## Saving rubrics

```json theme={null}
{ "name": "Call QA", "questions": { "...": "..." }, "shared": true }
```

Saving matters for consistency — everyone scoring against the same rubric
produces comparable numbers; everyone writing their own questions does not.

`shared: true` publishes workspace-wide and is **owner/admin only**. Runs of a
saved rubric file themselves under it, so `list_runs` with `rubric_id` returns
the whole series.

## Cost

Billed per **input token**. The upstream builds one row per question, each
carrying the shared content again, so a 10-question rubric encodes the
transcript **ten times**.

| Example                                   | Per 1,000 runs |
| ----------------------------------------- | -------------- |
| Ticket triage — 2 questions, short ticket | `$0.0007`      |
| Call QA — 3 questions, 3-minute call      | `$0.0150`      |
| Call QA — 3 questions, 15-minute call     | `$0.0750`      |
| Deep audit — 10 questions, long document  | `$0.5000`      |
| Extract — one turn                        | `$0.0004`      |

Context is **8K tokens per question**, not per request. Over-long content is
rejected with a `400` before anything is charged.

## Error handling

| Code          | Meaning                                                             |
| ------------- | ------------------------------------------------------------------- |
| `400`         | Invalid rubric or schema — the message names the offending question |
| `402`         | Insufficient credits; the shortfall is reported                     |
| `403`         | `shared: true` needs owner/admin; health is admin-only              |
| `404`         | Unknown rubric or run, or it belongs to someone else                |
| `429` / `503` | The judge is busy or unavailable — retry                            |

All of these refund automatically.

## When to reach for it

* **Scoring at volume** — every call, not a 2% sample
* **Routing** — send tickets to a team, rank a queue by urgency
* **Grading generated output** against the source it cited
* **Turn parsing** inside a live voice agent
* **Escalation** — anything under 0.85 goes to a human, automatically

<Warning>
  Judge gives consistent, comparable opinions — **not truth**. The model's own
  spec says outputs are *"not calibrated correctness guarantees."* Right tool for
  ranking, routing and flagging; wrong tool as the final word on a consequential
  decision without a person in the loop.
</Warning>

## Related

* [Judge CLI commands](/cli-reference/commands/judge)
* [JavaScript SDK](/sdks/javascript)
* [Python SDK](/sdks/python)
