jev-bias-bench

Fox-Islam/jev-bias-bench
Research

A benchmark for Jev's biases, built from people who differ in one attribute at a time.

PHPDocumented

jev-cookbook

chr-kelly/jev-cookbook
Research

Runnable question sets for TypeSafe's Jev, an eval harness with measured CLINC150 results, and a linter for the request shapes the API silently mis-reads.

PythonDocumented

jev-demo

co1smos/jev-demo
Research

Historical paper-trading simulator for evaluating TypeSafe AI JEV decisions

PythonDocumented

jev-evaluation

willkelly/jev-evaluation
Research

An adversarial evaluation of TypeSafe's jev decision model: nine experiments and 28 predictions fixed before any data was collected. 123,805 requests, $12.69.

PythonDocumented

jev-fastloop

SupremeDreamZ/jev-fastloop
Research

Cheap fail-open semantic edge layer for Jev (TypeSafe System One): decision service + confidence policy + decision logging + blind eval harness + skill router over 1,000+ skills. Stdlib-only Python.

PythonDocumented

jev-graph-walk

timsamart/jev-graph-walk
Research

Context-carrying graph retrieval with Jev: branching walks, reproducible ablations, and a visual replay.

PythonDocumented

jev-guardrails

codebam/jev-guardrails
Research

Jev-backed guardrails for agent tool calls: library, native OpenCode/Hermes/DeepSeek Harness hooks, and a hosted eval-credit service.

TypeScriptDocumented

jev-heart-risk-bench

rubinagentagi-tech/jev-heart-risk-bench
Research

Benchmarking Jev (TypeSafe System One) on 5,000 real CDC survey respondents, with an interactive demo where every profile has a real model answer

HTMLDocumented

jev-injection-bench

ASEVlad/jev-injection-bench
Research

Benchmarking TypeSafe's Jev against four rival prompt-injection detectors on 11,900 labelled prompts

PythonDocumented

jev-kit

FlorianRiquelme/jev-kit
Research

Typed client and benchmark harness for Jev, TypeSafe AI's System One decision model, through the Vercel AI Gateway. Measure accuracy, calibration and cost on your own data before you trust a threshold.

TypeScriptDocumented

jev-preflop-poker

marcbara/jev-preflop-poker
Research

Is Jev (TypeSafe System One model) a decent preflop poker player? Reproducible benchmark vs 4 reference styles

PythonDocumented

jev-rag-benchmark

emretheus/jev-rag-benchmark
Research

Free English RAG benchmark for TypeSafe Jev 1.13 — frozen candidate pools, calibration, paired bootstrap CIs, $0 runs. Compared with OpenJev and NVIDIA cross-encoders.

PythonDocumented

jev-replay-lab

Tomdachs/jev-replay-lab
Research

Local Jev evaluation workbench: datasets, typed questions, threshold simulation and run comparison

TypeScriptDocumented

jev-source-evaluator

stephotee/jev-source-evaluator
Research

Auditing LLM deep-research reports with TypeSafe's Jev: source authority, relevance, and whether cited figures mean what the report says

PythonDocumented

jev-tree-memory

Pizzawookiee/jev-tree-memory
Research

TypeSafe’s Jev AI + an n-ary memory tree = efficient agentic memory routing and retrieval.

PythonDocumented

jev-vs-llm-ticket-router

SarathChandraBellam/jev-vs-llm-ticket-router
Research

Benchmark: TypeSafe Jev vs traditional LLM on support-ticket routing accuracy, latency, and cost

PythonDocumented

jev-workbench

Chunky83/jev-workbench
Research

Jev Workbench: a C++ desktop workspace with Python diagnostics and structured TypeSafe evaluations

PythonDocumented

jev48

edgelabs-ai/jev48
Local alternative

Open, auditable reproduction of TypeSafe Jev: a 2B probabilistic decision model benchmarked across six public suites with weights, code, and receipts.

PythonDocumented

jev_rag

aijnek/jev_rag
Research

RAGのretrievalにTypeSafe APIを使うアイデアの実装

HTMLCode reference

jev_tool_calling_experiment

javierBrenesAI/jev_tool_calling_experiment
Research

Use Jev model from TypeSafe to evaluate if there is an improvement in latency and cost in the agentic loop for tool decisioning.

PythonDocumented

jevdash

Sunwood-ai-labs/jevdash
Research

100% clean-room, copyright-free 2D platformer benchmarking TypeSafe Jev AI in real-time control (60 FPS)

PythonDocumented

jevlergy

daisuke7/jevlergy
Research

Evaluating TypeSafe's System One model Jev through camera-based food allergen estimation (Flutter, iOS/Android)

PythonDocumented

livesignal

sandeeppanem/livesignal
Research

Real-time creator intelligence powered by direct TypeSafe Jev decisions and measurable full-corpus evaluation.

PythonDocumented

playwright-ai-benchmark

cancakmk/playwright-ai-benchmark
Research

Playwright AI Benchmark: Klasik Playwright vs JEV (TypeSafe AI) vs GPT-5.6 Luna vs GLM-5.3

TypeScriptDocumented