ExperimentsNew
Jev as a Judge
An evaluation harness comparing Jev with LLM judges on the same recorded agent runs.
An evaluation harness comparing Jev with LLM judges on the same recorded agent runs.
Jev answers typed yes-or-no, score and choice questions about recorded agent behavior, while the harness measures its judgments against LLM evaluators.
It tests where a fast calibrated decision model can replace an LLM judge and records accuracy, cost and latency on identical agent traces.
This project calls Jev on demand, in response to user input or events.