Jev Calibrate
A CLI for testing authored Jev questions against labeled examples, selecting thresholds, measuring calibration, and checking changes against held-out data.
A CLI for testing authored Jev questions against labeled examples, selecting thresholds, measuring calibration, and checking changes against held-out data.
The user defines Noul, Choice, and Score questions. The tool runs those questions against known examples, compares typed answers with labels, and reports metrics and suggested Noul thresholds.
Does the customer ask for money already paid to be returned?
One labeled support-ticket example and the project’s configured typed questions; the question set is user-authored and varies by calibration project.
Which team should handle this message?
One labeled support-ticket example and the project’s configured typed questions; the question set is user-authored and varies by calibration project.
How frustrated is the customer under the example rubric?
One labeled support-ticket example and the project’s configured typed questions; the question set is user-authored and varies by calibration project.
The user writes typed questions and labeled examples in a project directory.
The CLI checks schema and label compatibility before calling Jev.
Jev evaluates eligible labeled examples and the tool records each answer.
The CLI reports calibration metrics and threshold suggestions, then can compare revisions on a held-out set.
It treats a Jev prompt as an evaluable instrument: examples, held-out checks, threshold selection, and regression comparison help distinguish a useful gate from a ranker or unusable question.