jev-benchmark
wondertwins
Benchmarks Jev on chess moves and identifying which game NPC a player addresses.
- License
- MIT
- GitHub Stars
- 2
- Source reviewed
- 2026-09-19
Where Jev makes a decision
Selects legal chess moves or judges whether an utterance addresses each NPC.
What this project offers
Publishes labeled data, raw requests and responses, and evaluation code.
Review scope
Chess strength, F1 and latency come from the author’s specific tasks and settings, not general game-playing capability. README and integration source reviewed at a fixed commit; not independently run or benchmarked by this site.
Sources and implementation
Related projects
jev-review
NiazMorshed2007 · Evaluation & Observability
A local MCP code-quality reviewer returning structured scores to coding Agents.
jev-benchmarks
AbdelStark · Evaluation & Observability
A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.
jev-playground
mizchi · Evaluation & Observability
A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.
jev-lm
y0usaf · Evaluation & Observability
A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.