Jev-Persian-Benchmark
ArmanJR
Benchmarks for **Jev** on **480 authored general Persian questions** and a **24-excerpt classical Persian poetry pilot** (48 main questions plus 48 controls). Related general questions are batched; poetry questions run individually. Raw responses are saved and answers are scored locally, without a runtime model judge. The general benchmark's historical Laya comparison is retained below.
- License
- Not declared
- GitHub Stars
- 3
- Source reviewed
- —
Where Jev makes a decision
Jev returns a structured decision for the local program; consult the source for the exact decision policy.
What this project offers
Adds structured choices or scores to the workflow; performance and cost benefits have not been independently verified.
Review scope
Based on repository metadata and README with rule-based classification; pending human review, with no independent runtime or performance verification.
Sources and implementation
Related projects
jev
dannote · SDK & Decision Frameworks
Integrates Jev as an asynchronous Elixir/OTP process whose replies are handled with GenServer pattern matching.
zod-jev
jomatsu · SDK & Decision Frameworks
Adds semantic rules to Zod validation, such as checking whether text matches a description or contains personal information.
daf-jev
docxology · SDK & Decision Frameworks
A Python toolkit for Jev requests, batch evaluation, calibration and MCP access.
jev-starter
hamakyo · SDK & Decision Frameworks
TypeScript patterns for thresholds, fallbacks, human review and evaluation on top of the TypeSafe SDK.