Auto-marking maths scripts with Jev (TypeSafe System One): 2,054 scripts, 96.6% agreement with human markers
Benchmarks & Research for Jev — page 19
474 repositories with documented relationships and source evidence.
Jev (TypeSafe System One) research: API notes, benchmarks, community experiments, agent-loop patterns
Meridian | AI context-intake and contraction risk & compliance micro-app with live TypeSafe AI vs OpenAI evaluation
Discord bot that evaluates message like chess
Browser and macOS desktop agent for pi: Jev (TypeSafe System One) chooses each action from a structured observation in a bounded, surface-agnostic loop. Isolated Playwright tools, an allow-listed accessibility-tree tool, deterministic selectors, four-tier benchmarks.
Benchmark for Pi skill routing: BM25 vs TypeSafe Jev across roster sizes 50–500
TypeSafe System One (Jev) evaluated on synthetic FHIR clinical scenarios, with Claude as System Two
Eval framework library to evaluate AI hallucinations, correctness, and instruction following, powered by Jev AI.
Independent API-only calibration audit of TypeSafe AI's Jev decision model
A visual control room for typed AI evaluations, structured decision-making, and confidence-aware policy gating using Vercel AI SDK and TypeSafe Jev.
TypeSafe's Jev vs gpt-5.4-mini and gpt-5.6-luna on four public classification sets: cases, per-item answers, scoring, charts
Early-access test of TypeSafe's Jev on 24 Norwegian hearing responses, next to DeepSeek V4.1 Flash
Reproducible synthetic benchmark for temporal loan identity resolution with TypeSafe Jev
Claude Code skill: swap System 2 LLM pipeline components for System 1 TypeSafe Jev decisions via investigation, live three-arm eval, fallback, and an independent judge
Evaluating TypeSafe's Jev on Michael Mauboussin's 50-question decision calibration test
High-concurrency Elixir actor runtime for large-scale semantic graphs, search swarms, and batched model evaluation.
Synthetic smoking-history extraction benchmark comparing TypeSafe Jev and OpenAI structured outputs, with reproducible accuracy, cost, and latency results.
Automation Anywhere custom package wrapping the TypeSafe (Jev) API for typed yes/no, categorical, and scored-scale text evaluation.
A tug-of-war over one shared actuator — you and TypeSafe's Jev push the same CartPole and the forces add. Physics verified against Gymnasium, with fair-baseline benchmarks and token accounting.
Sandbox & benchmark: optimize how you ask Jev (TypeSafe System One) to play falling-block puzzle games, head-to-head against LLMs
Jev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証
A small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.
Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.