typesafe-ai-benchmark
iammrduncan
Compares Jev with other structured-output models on shared application tasks, recording errors, latency, Tokens, and estimated cost.
- License
- MIT
- GitHub Stars
- 32
- Source reviewed
- 2026-09-19
Where Jev makes a decision
Maps the same tasks to Jev Choice/Noul questions and normalizes answers to a shared result format.
What this project offers
Preserves comparison methods and results for inspecting model differences.
Review scope
The current README describes limited synthetic workloads and separately captured results. The demo video is not measurement data, and results are not a ranking for all tasks. This site has not reproduced them.
Sources and implementation
Related projects
jev-review
NiazMorshed2007 · Evaluation & Observability
A local MCP code-quality reviewer returning structured scores to coding Agents.
jev-benchmarks
AbdelStark · Evaluation & Observability
A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.
jev-playground
mizchi · Evaluation & Observability
A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.
jev-lm
y0usaf · Evaluation & Observability
A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.