Benchmark TypeSafe Jev against any OpenRouter model on your own data.
Benchmarks & Research for Jev — page 11
474 repositories with documented relationships and source evidence.
Ask Jev by TypeSafe AI whether a number is odd. TypeScript, real token usage, and latency benchmarks.
Reproducible riichi mahjong benchmark for Jev, GPT, Mortal, and hybrid agents using MJAI and RiichiEnv.
Local evaluation workbench for TypeSafe Jev
Independent benchmark data for TypeSafe's Jev (System One model) vs LLMs: accuracy, calibration, cost. Boards + per-decision logs, CC-BY-4.0
Drive a real browser or a real LangGraph agent through one declared user journey, and report what actually happened. A page saying 'success' is never accepted as proof.
Local stand-ins for TypeSafe Jev's System One interface: typed calibrated decisions from small models, with a routing benchmark and a distilled 19 ms encoder
Open source implementation of Poke
Jev test bench — Texas Hold'em edition: live-fire testing of TypeSafe's Jev evaluation model through a React poker game (Vercel AI Gateway). MIT.
Semantic firewall for LLM agents: tool calls gated by TypeSafe Jev (System One decision model via OpenRouter) + deterministic policy. PoC with corpus, stability eval, baseline, results.
Recorded comparison of Jev, Terra, and Opus on 100 synthetic security-triage cases, five passes each, with a static inspectable dashboard
A Jev-inspired decision interface for existing LLMs. Explicit choices, scores, calibration, and review thresholds.
Reproducible benchmark for TypeSafe AI's Jev on agent tool-call risk classification: accuracy, latency, and whether the confidence score is worth routing on.
Measures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets
TypeSafe's Jev model plays Snake live in the browser, with a headless eval and a self-improvement loop
Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.
TypeSafe ranking, classification, retrieval and structured-decision tools for Pi coding agents
I tortured Jev into being a RISC-V CPU.
Jev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels
Evaluation System for DG Fundraising Cause
面向语义决策模型的中文言下之意评测集:100 道伴侣对话 Choice 题,支持本地 ONNX、Jev 官方 API 与内网模型对比
Evaluate cloud spend against a budget and gate CI with Jev (approve, warn, block, or review).
⚡ JEV vs LAYA · Next-Gen AI Tetris Duel Benchmark Platform (TypeSafe Jev vs Local GPU ModernBERT) with React-Bits UI/UX
Reproducible experiment: can TypeSafe's Jev decision model prioritize SCA and SAST findings from context? Frozen benchmark, 50k scale run, frontier-model comparison, raw data and blog.