Evaluation & Observability
29 · Public sources and integration logic reviewed.
jev-review
NiazMorshed2007 · Evaluation & Observability
A local MCP code-quality reviewer returning structured scores to coding Agents.
jev-benchmarks
AbdelStark · Evaluation & Observability
A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.
jev-playground
mizchi · Evaluation & Observability
A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.
jev-lm
y0usaf · Evaluation & Observability
A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.
jev-chat
adhyaay-karnwal · Evaluation & Observability
A research chat decoder that repeatedly asks Jev to choose words or phrases and assembles them in code.
jev-agent-failure-benchmark
TokenTrim · Evaluation & Observability
A benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.
ask-jev
omni- · Evaluation & Observability
A Windows PowerShell tool for auditing recorded Codex execution evidence with :jev.
jev-rerank-bench
anessbelbati · Evaluation & Observability
Compares Jev, dedicated rerankers and chat models on the same retrieved passages.
jev-demos
Bud-ro · Evaluation & Observability
Maze experiments comparing Jev single-step choices with multi-step lookahead.
foreman-jev
Shifty-Eye-Games · Evaluation & Observability
An experimental Jev supervisor for Codex workers with programmer-selected acceptance commands.
latitude-llm
latitude-dev · Evaluation & Observability
Latitude includes an optional Jev preclassifier for conversation checks and their selection records.
taskuary
ldbumble · Evaluation & Observability
An optional Jev judgment module in Taskuary for checking user-defined conditions on task state.
jev-gomoku
XieChengYuan · Evaluation & Observability
A nine-board, 15×15 Gomoku workbench comparing how two Jev players respond to different input representations.
jevcal
abhixhek · Evaluation & Observability
A toolkit for evaluating Jev probabilities on labeled data, selecting confidence thresholds and checking model drift.
jev-benchmark
wondertwins · Evaluation & Observability
Benchmarks Jev on chess moves and identifying which game NPC a player addresses.
jev-frontend-qa
Nainish-Rai · Evaluation & Observability
Frontend QA that uses Jev to choose browser actions and checks contracts through DOM, HTTP and database evidence.
hermes-jev-north-star
poponline63 · Evaluation & Observability
A Hermes goal-checking skill that saves requirements, creates a run prompt and checks completion evidence.
jev-synergy-screening
PistachioAIHQ · Evaluation & Observability
A Jev title-and-abstract screening experiment compared with Cohen Abstract Triage labels for an ADHD review.
jev-exploration
SamuelSacco · Evaluation & Observability
A research repository tracking Jev claims and limitations, with calibration experiments and runnable examples.
jev-flash-review
TheBous · Evaluation & Observability
An MCP review engine that evaluates Agent-supplied diffs against explicit rules.
supercov
supercorp-ai · Evaluation & Observability
A quality and coverage CLI for coding agents: Jev assesses source properties while local coverage highlights testing targets.
jev-pref
doeixd · Evaluation & Observability
Turns AGENTS.md preferences into rules checked by Jev against hunks, staged files or PRs.
typesafe-playground
kavehmz · Evaluation & Observability
Interactive Jev experiments for support-routing previews and 3D driving simulations.
goodwatch-monorepo
alp82 · Evaluation & Observability
A film-and-TV attribute-scoring experiment inside GoodWatch comparing Jev question designs and batch sizes.
jev-calibration-audit
jujumilk3 · Evaluation & Observability
Audits Jev calibration, option-wording effects and Korean judgments through public APIs and datasets.
typesafe-ai-benchmark
iammrduncan · Evaluation & Observability
Compares Jev with other structured-output models on shared application tasks, recording errors, latency, Tokens, and estimated cost.
Canny
qkal · Evaluation & Observability
Keeps an execution ledger for Claude Code and Codex CLI to check for passing validation after edits.
jev-behavior-study
RINNECODER · Evaluation & Observability
An independent Jev 1.13.0 behavior study recording successes and failures across question framing, input conditions and games.
jev-eval
4esv · Evaluation & Observability
Compares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.