Judge Tools
Judge asks a pinned classifier a question about content you supply. It is a classifier, not a generator: it never writes prose, and every answer is a probability distribution over an answer space you defined. That matters for an agent. A language model asked “is this caller angry?” returns a paragraph you have to parse and cannot average. Judge returns0.86 on dismissive — a number you can threshold, sort and chart across
thousands of runs.
No extra credentials. Judge uses the same 60db API key the rest of the MCP
server already authenticates with —
SIXTYDB_API_KEY as configured, nothing
new to add.The two inference tools are billed per input token. A failed run is refunded
automatically — nothing is charged for an answer you did not get.Tools
The three question types
A rubric is a map of answer key → question. The shape ofcriteria
changes with the type:
The descriptions are what the model reads. A label with a vague
description produces a vague answer — that is the part to tune, not the
prompt.
sixtydb_judge_evaluate
state(required) — the content every question is asked about. Text or any JSON. Each question sees this and nothing else; they are independent, not a conversation.questions— map of answer key → question, max 32. Orrubric_idto run a saved one. Never both.model,label,save(falsebills without storing history)
sixtydb_judge_extract
The understanding layer for a voice or chat agent — what did they just ask for, and what values did they give me.responsePaths map selects the v2 model, which also returns how
the agent should reply. The mandatory unknown label is added server-side if
you leave it out.
The review queue
The single most valuable tool for an agent workflow:sixtydb_judge_list_runs with needs_review returns only runs where the
judge’s lowest confidence fell below 0.85 — the upstream’s own escalation
line. That is the queue a person should work through, instead of re-reading
everything the model was already confident about.
Visibility: you see your own runs; workspace owners and admins see everyone’s.
A run stores the content it judged, so history is not workspace-public by
default.
Saving rubrics
shared: true publishes workspace-wide and is owner/admin only. Runs of a
saved rubric file themselves under it, so list_runs with rubric_id returns
the whole series.
Cost
Billed per input token. The upstream builds one row per question, each carrying the shared content again, so a 10-question rubric encodes the transcript ten times.
Context is 8K tokens per question, not per request. Over-long content is
rejected with a
400 before anything is charged.
Error handling
All of these refund automatically.
When to reach for it
- Scoring at volume — every call, not a 2% sample
- Routing — send tickets to a team, rank a queue by urgency
- Grading generated output against the source it cited
- Turn parsing inside a live voice agent
- Escalation — anything under 0.85 goes to a human, automatically