Small decisions. Interesting possibilities.Submit contentSubmit
Jev Arena GIF preview

Jev Arena

A benchmark harness that compares Jev and DeepSeek on comment-labeling tasks and records per-item outcomes, usage, cost, and latency.

Added to Jevfast

How it uses Jev

For each comment in a batch, Jev answers a fixed family of typed questions about relevance, sentiment, intent, aspects, and emotions, plus a sentiment score. The host maps answers into labels and derives evidence quotes mechanically from source text.

What Jev decides

Representative source-derived requestExample answers · not a recorded Jev response · Source ↗
Question 1 · c0 relevant
YOUR APP
INSTRUCTION

Is the representative comment about or evaluating an AI model or product?

STATE

A batch of source comments/reviews. Each comment gets the same set of classification questions; IDs include the index of that comment in the batch.

JEV · NOUL
YesNo
Question 2 · c0 sentiment
YOUR APP
INSTRUCTION

Which overall attitude does the comment express?

STATE

A batch of source comments/reviews. Each comment gets the same set of classification questions; IDs include the index of that comment in the batch.

JEV · CHOICE
  1. positive
  2. negative
  3. neutral
  4. mixed
Question 3 · c0 intent
YOUR APP
INSTRUCTION

What is the comment’s main intent?

STATE

A batch of source comments/reviews. Each comment gets the same set of classification questions; IDs include the index of that comment in the batch.

JEV · CHOICE
  1. praise
  2. complaint
  3. question
  4. suggestion
  5. correction
  6. agreement
  7. disagreement
  8. joke
  9. information
  10. other
Question 4 · c0 score
YOUR APP
INSTRUCTION

Which source-defined sentiment tier best fits the comment?

STATE

A batch of source comments/reviews. Each comment gets the same set of classification questions; IDs include the index of that comment in the batch.

JEV · SCORE
04

App workflow

  1. Load evaluation set

    The operator selects a dataset and run configuration.

  2. Build batches

    The runner groups comments and generates per-comment question IDs.

  3. Call backends

    Jev and the comparison backend label each batch; oversized or failed batches follow the documented split and retry handling.

  4. Record and report

    The runner stores labels and events, then builds comparison, coverage, usage, cost, and latency reports.

Why it is interesting

The arena evaluates structured decisions at batch scale and preserves failure, cost, latency, and labeling evidence alongside comparative outcomes.