Can a System One model play arcade games? TypeSafe's Jev plays Tetris, Snake and 2048 — benchmarked against random and heuristic baselines.
Benchmarks & Research for Jev — page 15
474 repositories with documented relationships and source evidence.
Evaluate typed AI decisions on labeled French-language cases.
Measured benchmarks of TypeSafe's Jev decision model vs LLMs: ticket triage, voice-agent decision layer, bulk tagging
Finite-sample guarantees for Jev (TypeSafe's System One). Conformal risk control turns calibrated probabilities into certified routing thresholds; prediction-powered inference audits them. 2,412 decisions on CLINC150 for $0.23 — including the shift and prevalence cases where the guarantee breaks.
Early evaluation of TypeSafe's Jev on clinic call triage: 98.6% accurate when confident, 116k decisions per dollar
An 11-level crash course on Jev, TypeSafe AI's System One decision model — runnable examples against the real API, plus a capstone project with unit tests and evals. Works with any LLM provider via LiteLLM.
An 11-level crash course on Jev, TypeSafe AI's System One decision model — runnable examples against the real API, plus a capstone project with unit tests and evals. Works with any LLM provider via LiteLLM.
Prioritize human labels with active-learning strategies, durable review history and clean holdout evaluation.
A cottage in React Three Fiber whose eleven working parts are judged, as you flip them, by typesafe-ai/jev — an evaluation model reached through the Vercel AI Gateway.
Independent playground for TypeSafe AI Jev decision models: typed decisions, support-ticket routing, reproducible evaluations, and a local browser demo.
MCP server exposing TypeSafe AI's Jev decision model to Claude (stdio, two tools: jev_evaluate, jev_route)
Jev + Multimodal: shared visual decisions, evidence adapters, source audits and reproducible public-image benchmarks.
Reproducible benchmark evaluating TypeSafe AI's Jev (System One paradigm) on Brazil's ENEM 2025 standardized exam. Evaluates typed decision-making, domain-specific accuracy, and RLCD uncertainty calibration against open LLM baselines with an interactive GitHub Pages dashboard.
A playground for TypeSafe AI's Jev (System One model), built around five real production workflows: support triage, RAG relevance gating, agent action firewall, inline moderation, and CI eval judging.
Benchmark harness evaluating TypeSafe's Jev model on six public safety benchmarks against Shieldstral-1.0-3B and top SOTA guard models
Agent skill for TypeSafe AI's Jev decision model: fit assessment, integration recipes, calibration, multi-Jev, benchmarks
Can an at-home model match Jev? TypeSafe Jev vs self-hosted Laya on Moroccan Darija sentiment, with reproducible zero-shot evaluation.
A benchmark comparing TypeSafe's Jev versus local open-weight and API LLMs, as well as making Jev chat :) (ish)
Resume↔JD apply-gate cost eval: TypeSafe Jev vs Claude/Grok/GPT tiers (10 real JDs)
Local-first WebUI for TypeSafe Jev: batch document evaluation, typed questions, formulas, and CSV export. Unofficial.
Jevals (jevals.com): independent benchmark of TypeSafe's Jev vs LLMs
JevKit is an open-source decision infrastructure that helps developers integrate Jev into their applications, define structured decision tasks, evaluate model performance, and build reliable fallback policies.