Small decisions. Interesting possibilities.Submit contentSubmit

Jev Calibrate

A CLI for testing authored Jev questions against labeled examples, selecting thresholds, measuring calibration, and checking changes against held-out data.

Added to Jevfast

How it uses Jev

The user defines Noul, Choice, and Score questions. The tool runs those questions against known examples, compares typed answers with labels, and reports metrics and suggested Noul thresholds.

What Jev decides

Representative source-derived requestExample answers · not a recorded Jev response · Source ↗
Question 1 · refund requested
YOUR APP
INSTRUCTION

Does the customer ask for money already paid to be returned?

STATE

One labeled support-ticket example and the project’s configured typed questions; the question set is user-authored and varies by calibration project.

JEV · NOUL
YesNo
Question 2 · owner
YOUR APP
INSTRUCTION

Which team should handle this message?

STATE

One labeled support-ticket example and the project’s configured typed questions; the question set is user-authored and varies by calibration project.

JEV · CHOICE
  1. billing
  2. technical
  3. account
  4. other
Question 3 · frustration
YOUR APP
INSTRUCTION

How frustrated is the customer under the example rubric?

STATE

One labeled support-ticket example and the project’s configured typed questions; the question set is user-authored and varies by calibration project.

JEV · SCORE
02

App workflow

  1. Define questions and labels

    The user writes typed questions and labeled examples in a project directory.

  2. Validate project

    The CLI checks schema and label compatibility before calling Jev.

  3. Run examples

    Jev evaluates eligible labeled examples and the tool records each answer.

  4. Measure and tune

    The CLI reports calibration metrics and threshold suggestions, then can compare revisions on a held-out set.

Why it is interesting

It treats a Jev prompt as an evaluable instrument: examples, held-out checks, threshold selection, and regression comparison help distinguish a useful gate from a ranker or unusable question.