SREGym
An SRE agent evaluation framework with optional Jev decision support for reviewing diagnostic plans and submitted evidence.
An SRE agent evaluation framework with optional Jev decision support for reviewing diagnostic plans and submitted evidence.
When enabled for an operator-launched SRE run, Jev first ranks 3 to 5 proposed read-only diagnostic tests with one Choice and per-test Scores. A separate jev_submit path evaluates diagnosis-stage or verification-stage evidence with phase-specific Noul questions. The Conductor remains authoritative for accepting submissions; Jev does not execute commands or repairs.
Which proposed read-only test best distinguishes competing causes of the current application failure? Prefer a fresh, interpretable observation of failing behavior over rereading historical warnings or confirming healthy unrelated components. Select a test, not the most persuasive written diagnosis. A repair is not a diagnostic test.
Planning call: direct namespace snapshot, agent observations, and 3 to 5 candidate read-only diagnostic tests. Candidate test records contain hypothesis, command, supports_if, and rejects_if. The tool computes one Choice and a Score for each candidate.
Rate the diagnostic value of candidate test 1, including whether its two predicted outcomes distinguish competing explanations. Judge a read-only observation, not a repair or submission.
Planning call: direct namespace snapshot, agent observations, and 3 to 5 candidate read-only diagnostic tests. Candidate test records contain hypothesis, command, supports_if, and rejects_if. The tool computes one Choice and a Score for each candidate.
Does the evidence establish the agent's proposed causal mechanism for an active or repeatable application failure? An unusual setting, historical error, or correlation alone is insufficient. If there is no proposed mechanism, answer no.
Submission review at phase diagnose: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis in agent_hypothesis, and proposed_action (empty unless supplied).
Do the supplied observations demonstrate a recent or repeatable failed application operation relevant to the proposed diagnosis? Historical startup errors, healthy-workload restart counts, unusual settings, and hypothetical risks alone are insufficient. Judge the actual evidence, not the agent's assertion that an outage exists.
Submission review at phase diagnose: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis in agent_hypothesis, and proposed_action (empty unless supplied).
Do the supplied before-and-after observations support that the applied repair addressed the demonstrated cause of the original application failure? Evaluate the historical failure together with the current state. A healthy state after repair does not contradict a previously demonstrated failure, but fixing an unrelated anomaly or testing an unaffected path is insufficient.
Submission review at phase verify: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis and applied action, including supplied before-and-after functional evidence.
Does the proposed or applied repair correct the established mechanism while preserving application behavior during ordinary restarts, placement changes, and requests? A workaround that only avoids the currently failing path or relies on accidental placement is insufficient. Repeated successes under unchanged conditions alone do not prove this.
Submission review at phase verify: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis and applied action, including supplied before-and-after functional evidence.
Do the fresh functional tests exercise the actual failing behavior and support the claimed result? Healthy pods or tests of unaffected paths are insufficient. If no functional result is supplied, answer no.
Submission review at phase verify: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis and applied action, including supplied before-and-after functional evidence.
The user selects a benchmark problem and agent configuration.
The agent interacts with the problem environment and collects evidence.
When enabled, Jev reviews diagnostic tests or a submission using the supplied evidence.
The agent or operator uses the review within the benchmark workflow; Jev does not execute repairs.
It places typed review decisions inside an SRE benchmark workflow and keeps Jev as an optional support path rather than the agent that runs repairs.