发明 RLHF 的人,这次做了个不会说话的模型:Jev 独立研究报告。52 页 PDF + 50 条中文实测复现包 + 143 条可回溯数据表
Benchmarks & Research for Jev — page 20
474 repositories with documented relationships and source evidence.
A small second eval for shadcn-ui/lint that uses TypeSafe's Jev to judge the linter's own output.
Application of TypeSafe Jev (noul judgment primitive) on the collusion.wiki corpus: agent vs human page authorship, head-to-head vs local Qwen3.8-Flash-Next
Checking JEV's Contradiction detection (Model by TypeSafe.AI)
Local, single-user job-posting evaluator: typed semantic signals plus a deterministic scorer produce explainable 0-100 fit scores, with every claim traced to a line of the posting.
A typed decision-routing prototype for turning medical learning material into deterministic study actions, designed to evaluate TypeSafe Jev.
Ongoing Japanese research deck on Jev and System One models, maintained as Markdown slides.
Evaluating TypeSafe's Jev as a fast monitor and action gate for agent sabotage in SHADE-Arena, compared with Gemini 2.5 Flash/Pro.
Charts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models
Smoke test: route prompts to a work or life database with TypeSafe AI (Jev), 30-case eval
Does Jev predict stock returns from news? It reads the news well; there is no tradeable alpha. Three arms separate reading from recall.
Jev (TypeSafe) vs. Gemini 3.8 Flash vs. GPT-5.6 Luna na anotação estruturada de sentenças do TJSP: qualidade, tempo e custo
Position paper: the Hidden-Markov and fuzzy primitives missing from TypeSafe AI's Jev and System-One decision models. Two lemmas, one principle (Deferred Crispification), one architecture (BSF-S1).
Can Jev pick the winner of a real headline A/B test? 64.5% across 10,984 Upworthy randomized experiments, 74.7% when the difference was decisive.
TypeScript experiments, evaluations, and latency benchmarks for TypeSafe's Jev model
Rust port of TypeSafe system-one-adapter (LLM-backed system_one evaluations)
Evaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call
Durable orchestration, delegation, evidence, evaluation, review, and recovery for coding-agent workflows.