JEValuate

Akeel-Majeed/JEValuate
Research

Auto-marking maths scripts with Jev (TypeSafe System One): 2,054 scripts, 96.6% agreement with human markers

TypeScriptDocumented

jevx

umgbhalla/jevx
Research

Jev (TypeSafe System One) research: API notes, benchmarks, community experiments, agent-loop patterns

PythonDocumented

meridian-demo

MoRohn/meridian-demo
Research

Meridian | AI context-intake and contraction risk & compliance micro-app with live TypeSafe AI vs OpenAI evaluation

TypeScriptDocumented

pi-Jev-browser

laihenyi/pi-Jev-browser
Research

Browser and macOS desktop agent for pi: Jev (TypeSafe System One) chooses each action from a structured observation in a bounded, surface-agnostic loop. Isolated Playwright tools, an allow-listed accessibility-tree tool, deterministic selectors, four-tier benchmarks.

TypeScriptDocumented

pi-jev-skill-bench

iamdin/pi-jev-skill-bench
Research

Benchmark for Pi skill routing: BM25 vs TypeSafe Jev across roster sizes 50–500

TypeScriptDocumented

explore-typesafe-ai

si618/explore-typesafe-ai
Research

TypeSafe System One (Jev) evaluated on synthetic FHIR clinical scenarios, with Claude as System Two

PythonDocumented

ground-zero

zavocc/ground-zero
Research

Eval framework library to evaluate AI hallucinations, correctness, and instruction following, powered by Jev AI.

PythonDocumented

jev-calibration-audit

jujumilk3/jev-calibration-audit
Research

Independent API-only calibration audit of TypeSafe AI's Jev decision model

PythonDocumented

jev-control-room

thisisjorge/jev-control-room
Research

A visual control room for typed AI evaluations, structured decision-making, and confidence-aware policy gating using Vercel AI SDK and TypeSafe Jev.

TypeScriptDocumented

jev-eval

onlyoneaman/jev-eval
Research

TypeSafe's Jev vs gpt-5.4-mini and gpt-5.6-luna on four public classification sets: cases, per-item answers, scoring, charts

TypeScriptDocumented

jev-horingssvar-eval

EmilLindfors/jev-horingssvar-eval
Research

Early-access test of TypeSafe's Jev on 24 Norwegian hearing responses, next to DeepSeek V4.1 Flash

PythonDocumented

jev-loan-identity-benchmark

KiishiAD/jev-loan-identity-benchmark
Research

Reproducible synthetic benchmark for temporal loan identity resolution with TypeSafe Jev

PythonDocumented

jev-swap

HomenShum/jev-swap
Research

Claude Code skill: swap System 2 LLM pipeline components for System 1 TypeSafe Jev decisions via investigation, live three-arm eval, fallback, and an independent judge

PythonDocumented

jev-takes-mauboussin

eggmasonvalue/jev-takes-mauboussin
Research

Evaluating TypeSafe's Jev on Michael Mauboussin's 50-question decision calibration test

PythonDocumented

plexus

nshkrdotcom/plexus
Research

High-concurrency Elixir actor runtime for large-scale semantic graphs, search swarms, and batched model evaluation.

ElixirDocumented

smoking-extraction-benchmark

vclic/smoking-extraction-benchmark
Research

Synthetic smoking-history extraction benchmark comparing TypeSafe Jev and OpenAI structured outputs, with reproducible accuracy, cost, and latency results.

PythonDocumented

typesafe-package

AutomationAnywhere/typesafe-package
Research

Automation Anywhere custom package wrapping the TypeSafe (Jev) API for typed yes/no, categorical, and scored-scale text evaluation.

JavaDocumented

cartpole-jev

yslinear/cartpole-jev
Research

A tug-of-war over one shared actuator — you and TypeSafe's Jev push the same CartPole and the forces add. Physics verified against Gymnasium, with fair-baseline benchmarks and token accounting.

JavaScriptDocumented

jev-gamebenchmark

aieo-product/jev-gamebenchmark
Research

Sandbox & benchmark: optimize how you ask Jev (TypeSafe System One) to play falling-block puzzle games, head-to-head against LLMs

PythonDocumented

jev-jp-address

smasato/jev-jp-address
Research

Jev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証

TypeScriptDocumented

jev-pick-and-place-study

tryaksh/jev-pick-and-place-study
Research

A small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.

PythonDocumented

jev-playground

hegargarcia/jev-playground
Research

Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.

TypeScriptDocumented