LangWatch Instant Evals
Runs typed Jev judgments over selected traces and spans in a user-triggered Instant Eval job.
Runs typed Jev judgments over selected traces and spans in a user-triggered Instant Eval job.
An Instant Eval job selects traces or spans and passes their text to Jev as state. Configured classifier questions produce Boolean, integer Score, or Category Choice judgments, and the run service stores those verdicts with the evaluated traces.
Judge whether the configured criterion holds for the supplied trace text.
Text extracted from a selected trace or span, plus configured evaluator questions.
A user or API requests an evaluation job over selected traces or spans.
The run service loads a bounded page and extracts text for the configured evaluator.
Build typed classifier questions and send them with the trace text.
Validate and persist results against the evaluated traces or spans.
LangWatch connects typed evaluation criteria to observability data: a user can define an evaluator once, run it over selected traces, and inspect stored verdicts beside the original activity.