awesome jevDIRECTORYSponsorStar on GitHub

jev-agent-failure-benchmark

TokenTrim

A benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.

GitHub repository ↗Search and filter
License
Apache-2.0
GitHub Stars
1
Source reviewed
2026-09-19

Where Jev makes a decision

Builds candidate sets from traces and submits three choice questions.

What this project offers

Provides evaluation scripts and author results; some baselines generate answers while Jev selects candidates.

Review scope

Some Jev benchmark axes use constrained choices while paper baselines generate freely; not every metric is a like-for-like comparison.

Sources and implementation

Related projects