jev-agent-failure-benchmark
TokenTrim
A benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.
- License
- Apache-2.0
- GitHub Stars
- 1
- Source reviewed
- 2026-09-19
Where Jev makes a decision
Builds candidate sets from traces and submits three choice questions.
What this project offers
Provides evaluation scripts and author results; some baselines generate answers while Jev selects candidates.
Review scope
Some Jev benchmark axes use constrained choices while paper baselines generate freely; not every metric is a like-for-like comparison.
Sources and implementation
Related projects
jev-review
NiazMorshed2007 · Evaluation & Observability
A local MCP code-quality reviewer returning structured scores to coding Agents.
jev-benchmarks
AbdelStark · Evaluation & Observability
A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.
jev-playground
mizchi · Evaluation & Observability
A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.
jev-lm
y0usaf · Evaluation & Observability
A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.