Small decisions. Interesting possibilities.Submit contentSubmit

SREGym

An SRE agent evaluation framework with optional Jev decision support for reviewing diagnostic plans and submitted evidence.

Added to Jevfast

How it uses Jev

When enabled for an operator-launched SRE run, Jev first ranks 3 to 5 proposed read-only diagnostic tests with one Choice and per-test Scores. A separate jev_submit path evaluates diagnosis-stage or verification-stage evidence with phase-specific Noul questions. The Conductor remains authoritative for accepting submissions; Jev does not execute commands or repairs.

What Jev decides

Representative Jev planning request; minimum 3 dynamic candidate testsExample answers · not a recorded Jev response · Source ↗
Question 1 · next test
YOUR APP
INSTRUCTION

Which proposed read-only test best distinguishes competing causes of the current application failure? Prefer a fresh, interpretable observation of failing behavior over rereading historical warnings or confirming healthy unrelated components. Select a test, not the most persuasive written diagnosis. A repair is not a diagnostic test.

STATE

Planning call: direct namespace snapshot, agent observations, and 3 to 5 candidate read-only diagnostic tests. Candidate test records contain hypothesis, command, supports_if, and rejects_if. The tool computes one Choice and a Score for each candidate.

JEV · CHOICE
  1. test_1
  2. test_2
  3. test_3
  4. revise_tests
Question 2 · value 1
YOUR APP
INSTRUCTION

Rate the diagnostic value of candidate test 1, including whether its two predicted outcomes distinguish competing explanations. Judge a read-only observation, not a repair or submission.

STATE

Planning call: direct namespace snapshot, agent observations, and 3 to 5 candidate read-only diagnostic tests. Candidate test records contain hypothesis, command, supports_if, and rejects_if. The tool computes one Choice and a Score for each candidate.

JEV · SCORE
04
Diagnosis-stage review requestExample answers · not a recorded Jev response · Source ↗
Question 1 · causal support
YOUR APP
INSTRUCTION

Does the evidence establish the agent's proposed causal mechanism for an active or repeatable application failure? An unusual setting, historical error, or correlation alone is insufficient. If there is no proposed mechanism, answer no.

STATE

Submission review at phase diagnose: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis in agent_hypothesis, and proposed_action (empty unless supplied).

JEV · NOUL
YesNo
Question 2 · active failure
YOUR APP
INSTRUCTION

Do the supplied observations demonstrate a recent or repeatable failed application operation relevant to the proposed diagnosis? Historical startup errors, healthy-workload restart counts, unusual settings, and hypothetical risks alone are insufficient. Judge the actual evidence, not the agent's assertion that an outage exists.

STATE

Submission review at phase diagnose: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis in agent_hypothesis, and proposed_action (empty unless supplied).

JEV · NOUL
YesNo
Verification-stage review requestExample answers · not a recorded Jev response · Source ↗
Question 1 · causal support
YOUR APP
INSTRUCTION

Do the supplied before-and-after observations support that the applied repair addressed the demonstrated cause of the original application failure? Evaluate the historical failure together with the current state. A healthy state after repair does not contradict a previously demonstrated failure, but fixing an unrelated anomaly or testing an unaffected path is insufficient.

STATE

Submission review at phase verify: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis and applied action, including supplied before-and-after functional evidence.

JEV · NOUL
YesNo
Question 2 · durable repair
YOUR APP
INSTRUCTION

Does the proposed or applied repair correct the established mechanism while preserving application behavior during ordinary restarts, placement changes, and requests? A workaround that only avoids the currently failing path or relies on accidental placement is insufficient. Repeated successes under unchanged conditions alone do not prove this.

STATE

Submission review at phase verify: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis and applied action, including supplied before-and-after functional evidence.

JEV · NOUL
YesNo
Question 3 · functional evidence
YOUR APP
INSTRUCTION

Do the fresh functional tests exercise the actual failing behavior and support the claimed result? Healthy pods or tests of unaffected paths are insufficient. If no functional result is supplied, answer no.

STATE

Submission review at phase verify: direct fresh compact Kubernetes namespace snapshot, agent observations, proposed causal diagnosis and applied action, including supplied before-and-after functional evidence.

JEV · NOUL
YesNo

App workflow

  1. Choose an SRE problem

    The user selects a benchmark problem and agent configuration.

  2. Run diagnostic work

    The agent interacts with the problem environment and collects evidence.

  3. Request Jev review

    When enabled, Jev reviews diagnostic tests or a submission using the supplied evidence.

  4. Continue or submit

    The agent or operator uses the review within the benchmark workflow; Jev does not execute repairs.

Why it is interesting

It places typed review decisions inside an SRE benchmark workflow and keeps Jev as an optional support path rather than the agent that runs repairs.