Skip to main content

Judge

Judge asks a pinned classifier a question about content you supply. It never generates text: every answer is a probability distribution over an answer space you defined, which is what makes the results averageable, sortable and safe to feed straight into other code.
Uses your existing 60db credentials. Judge wraps the 60db Jev model behind the same API key as TTS, STT and Memory — there is no separate Judge key, and keys you already issued keep working. 60db config and X60DB_API_KEY apply unchanged.The two inference commands are billed per input token. A failed run is refunded automatically — you are never charged for an answer you did not get. Every billed command prints the charge and remaining balance in a footer.

Commands

The three question types

A rubric is a JSON map of answer key → question. Each question is one of:
rubric.json
Always include an unknown option on a choice. Without an escape hatch the model has to pick one of your real labels even on off-topic content.
The descriptions are what the model reads — they are the part to tune. A label with a vague description produces vague answers.

Evaluate content

Options:
  • -s, --state <text> — The content to judge
  • -f, --state-file <path> — Read the content from a file instead
  • -r, --rubric <id> — Run a saved rubric by id
  • --rubric-file <path> — Run an inline rubric from a JSON file
  • -m, --model <name> — Model to use (see judge:models)
  • -l, --label <text> — Tag the run for later filtering
  • --no-save — Bill the run but keep it out of history
Pass either --rubric or --rubric-file, never both. Output:
The distribution is the answer. A choice with 0.86 on the winner and one with 0.34 read identically if you only print the label.

Extract from a turn

Classify one conversational turn and pull the values out of it:
schema.json
Options:
  • -t, --text <text> — The turn (required, max 2,048 characters)
  • --schema-file <path> — JSON file with the label schema (required)
  • --profile <name> — generic (default) or medical
  • --budget <ms> — Total request budget, 1–30000
  • --no-save — Bill the run but keep it out of history
Adding a responsePaths map to the schema selects the v2 model, which also returns how the agent should reply. Output:
Entity start/end are Unicode code point offsets, not byte or UTF-16 indices. An emoji earlier in the turn shifts every naive slice after it.

Save and re-run a rubric

Build the rubric once, then run it by id from then on:
Options:
  • -n, --name <name> — Unique per workspace
  • -f, --file <path> — JSON file containing the questions map
  • -d, --description <text> — What the rubric is for
  • -m, --model <name> — Default model for this rubric
  • --shared — Publish to the whole workspace (owner/admin only)
Every run of a saved rubric files itself under it, so 60db judge:runs --rubric <id> gives you the whole series.

The review queue

The single most useful command. Anything the judge was under 85% sure about — the upstream’s own escalation line:
Options:
  • --needs-review — Only low-confidence runs
  • --kind <kind> — evaluate or extract
  • -r, --rubric <id> — Only runs of one rubric
  • --limit <n> / --offset <n> — Paging (max 100 per page)
History is private to whoever made the run. Workspace owners and admins see everyone’s — a run stores the content it judged.

Models and spend

Do not hardcode model names — judge:models is the source of truth and changes when the deployment does.

Pricing

Billed on input tokens, the model’s own unit. The upstream builds one row per question, each carrying the shared content again, so a 10-question rubric encodes the transcript ten times. Question count and content length both land in the token total on their own. Keep rubrics tight when the content is long — that last row is the same model, just asked ten questions about twenty thousand characters. Context is 8K tokens per question, not per request. An over-long piece of content is rejected with a 400 before anything is charged.

Agent-friendly JSON

Every command takes the global --json flag for scripting:

Handling errors

All of these refund automatically. Nothing is charged for a run that failed.