Overview
60db’s Judge answers questions you define about a piece of content, and returns a probability distribution for every answer. It never writes prose. That is the whole point. Ask a language model “is this caller angry?” and you get a paragraph you then have to parse. Ask Judge and you get0.86 on dismissive — a number you can sort by, threshold on, and compare across a hundred thousand runs.
Your Labels, Not Ours
You supply the options and the descriptions; the model only picks between them
Comparable Output
Every answer is a distribution over a fixed answer space, so runs sort and diff
Up to 32 Questions
Each asked independently against the same content, in one request
Saved Rubrics
Store a question set once, then run it by
rubric_id foreverConfidence & Review Queue
Below
0.85, route it to a person — filter history on needs_reviewTurn Extraction
Intent, operation and entity spans from a single conversational turn
Your existing 60db API key already works. Judge wraps the 60db Jev model behind the same platform credential as TTS, STT and Memory — there is no separate Judge key to issue. Keys created from now on carry a
judge scope; keys issued earlier are accepted on their slm scope.Core concepts
Questions, not prompts
A Judge request has two halves: thestate (the content — a transcript, a ticket, a document, any JSON structure), and the questions asked about it. Each question is keyed by a name of your choosing, which is how you read the answer back. The key is never sent to the model.
Questions are asked independently. Every one sees the same state and its own criteria, and nothing else — they are not a conversation, and question 3 cannot see what question 1 answered.
The three question types
Two results are easy to misread:
scoreis a weighted position, not an index.1.87on a four-rung ladder means the mass sits between rungs 1 and 2 — it is not “rung 1” rounded.noulhas deliberately noconfidencefield. The number is the probability of true; it is not a calibrated correctness estimate.
Confidence and escalation
Every run reportsmin_confidence — the lowest confidence across its answers (null when every question was noul). Below 0.85, a human should look. That is the service’s own escalation threshold, and it is the filter the review queue is built on.
escalated_questions tells you how many answers the service already refined with its second-stage model because confidence was low.
Evaluating content
- JavaScript
- Python
- cURL
- CLI
answers, every run returns a summary — one headline value per question (the chosen option, the score, or the probability). That is the shape to store on your own record when you do not need the full distribution.
Saved rubrics
Save a question set once and run it by id from then on. Runs of a saved rubric file themselves under it, so listing byrubric_id returns the whole series — which is what makes a rubric a trend line rather than a one-off.
- JavaScript
- Python
- CLI
questions or rubric_id, never both. Saved rubrics are re-validated on every run — a rubric is not trusted just because it was valid when stored.
Extracting from a turn
extract is the understanding layer between speech-to-text and a reply: what did the speaker just ask for, and what values did they give me. It labels one turn with an intent and an operation, and pulls entity spans out with their exact positions.
- JavaScript
- Python
- CLI
schema.responsePaths selects the v2 model, which also returns a response_path and tracks conversation state. text is capped at 2,048 Unicode code points; the required unknown label is added for you if you leave it out of a classification map.
The review queue
Rather than re-reading everything the model was already sure about, list only what it wasn’t:- JavaScript
- Python
- CLI
kind (evaluate or extract), by rubric_id for one rubric’s whole series, and by your own confidence_below threshold. Pass save: false on a run to bill it but keep it out of history entirely.
Models
Fetch the model list; don’t hardcode it.jev-latest is an alias that resolves server-side to a versioned id, and that versioned id is what the evaluate response reports back as model. Responses are cached for 5 minutes, so polling to populate a picker is free.
Pricing
Judge is pay-as-you-go on input tokens only — 0`, because a classifier emits no text. It bills to the same workspace wallet as every other 60db service. The one thing that will surprise you:state again. Ten questions about a 20,000-character document encode that document ten times — so keep rubrics tight when the content is long. Context is 8K tokens per question, not per request; over-long content is rejected with a 400 before anything is charged.
A failed run costs nothing. The wallet is debited before the call and refunded automatically if anything goes wrong — a busy queue (
429), an unavailable service (503), a rejected rubric (400). Both entries appear in the ledger, netting to zero. See Judge pricing for the full reconciliation story.Errors
API Reference
Evaluate
Score content against a rubric
Extract
Intent, operation and entity spans
Rubrics
Save, list, update and delete
Runs
History and the review queue
Models
Model names this deployment serves
Usage
Spend and run counts
Pricing
Rates, token maths and refunds
CLI Commands
60db judge:* from the terminalMCP Tools
Judge from Claude and other agents