Skip to main content

Judge Tools

Judge asks a pinned classifier a question about content you supply. It is a classifier, not a generator: it never writes prose, and every answer is a probability distribution over an answer space you defined. That matters for an agent. A language model asked “is this caller angry?” returns a paragraph you have to parse and cannot average. Judge returns 0.86 on dismissive — a number you can threshold, sort and chart across thousands of runs.
No extra credentials. Judge uses the same 60db API key the rest of the MCP server already authenticates with — SIXTYDB_API_KEY as configured, nothing new to add.The two inference tools are billed per input token. A failed run is refunded automatically — nothing is charged for an answer you did not get.

Tools

The three question types

A rubric is a map of answer key → question. The shape of criteria changes with the type:
Always include an unknown option on a choice. Without an escape hatch the model must pick one of your real labels even on off-topic content.
The descriptions are what the model reads. A label with a vague description produces a vague answer — that is the part to tune, not the prompt.

sixtydb_judge_evaluate

Parameters:
  • state (required) — the content every question is asked about. Text or any JSON. Each question sees this and nothing else; they are independent, not a conversation.
  • questions — map of answer key → question, max 32. Or rubric_id to run a saved one. Never both.
  • model, label, save (false bills without storing history)
Returns one answer per question with its full distribution, the lowest confidence across the set, token usage and the charge.

sixtydb_judge_extract

The understanding layer for a voice or chat agent — what did they just ask for, and what values did they give me.
Adding a responsePaths map selects the v2 model, which also returns how the agent should reply. The mandatory unknown label is added server-side if you leave it out.
Entity start/end are Unicode code point offsets, not UTF-16 indices. Slice with Array.from(text).slice(start, end) — a plain text.slice() lands in the wrong place as soon as the turn contains an emoji.

The review queue

The single most valuable tool for an agent workflow:
sixtydb_judge_list_runs with needs_review returns only runs where the judge’s lowest confidence fell below 0.85 — the upstream’s own escalation line. That is the queue a person should work through, instead of re-reading everything the model was already confident about. Visibility: you see your own runs; workspace owners and admins see everyone’s. A run stores the content it judged, so history is not workspace-public by default.

Saving rubrics

Saving matters for consistency — everyone scoring against the same rubric produces comparable numbers; everyone writing their own questions does not. shared: true publishes workspace-wide and is owner/admin only. Runs of a saved rubric file themselves under it, so list_runs with rubric_id returns the whole series.

Cost

Billed per input token. The upstream builds one row per question, each carrying the shared content again, so a 10-question rubric encodes the transcript ten times. Context is 8K tokens per question, not per request. Over-long content is rejected with a 400 before anything is charged.

Error handling

All of these refund automatically.

When to reach for it

  • Scoring at volume — every call, not a 2% sample
  • Routing — send tickets to a team, rank a queue by urgency
  • Grading generated output against the source it cited
  • Turn parsing inside a live voice agent
  • Escalation — anything under 0.85 goes to a human, automatically
Judge gives consistent, comparable opinions — not truth. The model’s own spec says outputs are “not calibrated correctness guarantees.” Right tool for ranking, routing and flagging; wrong tool as the final word on a consequential decision without a person in the loop.