Skip to main content
POST
Ask up to 32 questions about one piece of content and get a probability distribution for each. Judge is a classifier, not a generator. It never writes prose: every answer is a distribution over an answer space you defined, which is what makes results comparable across thousands of runs. A language model asked “is this caller angry?” returns a paragraph; this returns 0.86 on dismissive. Each question is asked independently — every one sees the same state and its own criteria, and nothing else. They are not a conversation.
Billed per input token. A failed run is refunded automatically — you are never charged for an answer you did not get. See Judge pricing.
Your existing 60db API key already works. Judge wraps the 60db Jev model behind the same platform credential as TTS, STT and Memory — you do not need a separate key, and you do not need to reissue the one you have. Keys created from now on carry a judge scope; keys issued earlier are accepted on their slm scope.

Request

Headers

string
required
Bearer token — your standard 60db API key (sk_…), or a user JWT. There is no separate Judge credential.
string
required
application/json

Body

string | object | array
required
The content every question is asked about — a transcript, a ticket, a document, or any JSON structure. Top-level numbers, booleans and null are rejected.
object
Map of answer key → question, 1–32 entries. The key is your handle on the answer; it is never sent to the model.Supply this or rubric_id, never both.Each question is one of three shapes:choice — pick one of your options.
criteria is a map of option name → description, max 255 options.score — rate on a ladder.
An ordered array, lowest first, max 10 levels.noul — true or false.
criteria is optional: { "true": "...", "false": "..." }.
string
Run a saved rubric by id instead of inline questions. See List rubrics.
string
default:"jev-latest"
Model name. Fetch the valid list from List models rather than hardcoding it.
boolean
default:"true"
false bills the run but keeps it out of history.
string
Tag for grouping runs in history. Max 120 characters.

Response

object
One answer per question, under the same key you supplied.choice — choice (the winning option), probabilities (every option, summing to 1), confidence.score — score (a probability-weighted position on your ladder, e.g. 1.87 — not an index), legend (level index → your description), probabilities, confidence.noul — noul, a raw probability of true. There is deliberately no confidence field: the upstream is explicit that this is not a calibrated correctness estimate.
object
One headline value per question — the chosen option, the score, or the probability. What the history list renders.
number | null
Lowest confidence across the answers, or null when every question was noul. Below 0.85 a human should look — that is the upstream’s own escalation threshold.
integer
How many answers the service refined with its second-stage model because confidence was low.
object
input_tokens and output_tokens. Output is always 0 — a classifier emits no text.
number
What this run cost, in USD.
string | null
The stored run’s id, or null when save: false.

Notes

Always include an unknown option on a choice. Without an escape hatch the model must pick one of your real labels even on off-topic content.
The descriptions are what the model reads — they are the thing to tune, not a prompt. A label with a vague description produces a vague answer. Context is 8K tokens per question, not per request. The shared state is re-encoded for every question, so a long document with many questions is fine only while each individual row fits. Over-long content is rejected with a 400 before anything is charged.

Errors

Example