Jev-Persian-Benchmark
ArmanJR
Benchmarks for **Jev** on **480 authored general Persian questions** and a **24-excerpt classical Persian poetry pilot** (48 main questions plus 48 controls). Related general questions are batched; poetry questions run individually. Raw responses are saved and answers are scored locally, without a runtime model judge. The general benchmark's historical Laya comparison is retained below.
- ライセンス
- 記載なし
- GitHub Stars
- 3
- ソース確認日
- —
Jev が判断する箇所
Jev returns a structured decision for the local program; consult the source for the exact decision policy.
このプロジェクトの用途
Adds structured choices or scores to the workflow; performance and cost benefits have not been independently verified.
確認の範囲
公開ソースと連携ロジックを確認済み。
ソースと実装
同じカテゴリのプロジェクト
jev
dannote · SDK・判断フレームワーク
Jev を Elixir/OTP の非同期プロセスとして組み込み、GenServer のパターンマッチで応答を処理する。
zod-jev
jomatsu · SDK・判断フレームワーク
説明との一致や個人情報の有無など、意味に基づくルールを Zod の検証に加える。
daf-jev
docxology · SDK・判断フレームワーク
Jev の呼び出し・バッチ評価・校正・MCP 接続をまとめた Python ツールキット。
jev-starter
hamakyo · SDK・判断フレームワーク
TypeSafe SDK に、しきい値、代替経路、人による確認、評価のパターンを加える TypeScript ツール集。