Small decisions. Interesting possibilities.Submit a project / videoSubmit

How it uses Jev

Jev answers typed yes-or-no, score and choice questions about recorded agent behavior, while the harness measures its judgments against LLM evaluators.

Why it's interesting

It tests where a fast calibrated decision model can replace an LLM judge and records accuracy, cost and latency on identical agent traces.

How often does Jev run?

This project calls Jev on demand, in response to user input or events.