jev-benchmarks
AbdelStark
A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.
- License
- Apache-2.0
- GitHub Stars
- 9
- Source reviewed
- 2026-09-19
Where Jev makes a decision
Runs the same labeled text tasks through both backends and records probabilities, latency and failures.
What this project offers
Helps examine task-specific accuracy and whether confidence scores support chosen thresholds.
Review scope
The author reports a 300-example pilot. Hosted Jev and local GLiNER timings are not hardware-normalized; this site did not rerun it.
Sources and implementation
Related projects
jev-review
NiazMorshed2007 · Evaluation & Observability
A local MCP code-quality reviewer returning structured scores to coding Agents.
jev-playground
mizchi · Evaluation & Observability
A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.
jev-lm
y0usaf · Evaluation & Observability
A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.
jev-chat
adhyaay-karnwal · Evaluation & Observability
A research chat decoder that repeatedly asks Jev to choose words or phrases and assembles them in code.