jev-eval

4esv/jev-eval
Research

Benchmark TypeSafe Jev against any OpenRouter model on your own data.

PythonDocumented

Jev-is-odd

robipop22/Jev-is-odd
Research

Ask Jev by TypeSafe AI whether a number is odd. TypeScript, real token usage, and latency benchmarks.

JavaScriptDocumented

jev-mahjong-bench

hamakyo/jev-mahjong-bench
Research

Reproducible riichi mahjong benchmark for Jev, GPT, Mortal, and hybrid agents using MJAI and RiichiEnv.

TypeScriptDocumented

jevals

dayhaysoos/jevals
Research

Local evaluation workbench for TypeSafe Jev

TypeScriptDocumented

jevals-data

Jevals/jevals-data
Research

Independent benchmark data for TypeSafe's Jev (System One model) vs LLMs: accuracy, calibration, cost. Boards + per-decision logs, CC-BY-4.0

Documented

journey-evals

gargpratyush/journey-evals
Research

Drive a real browser or a real LangGraph agent through one declared user journey, and report what actually happened. A page saying 'success' is never accepted as proof.

PythonDocumented

local-system-one

djhoomin/local-system-one
Research

Local stand-ins for TypeSafe Jev's System One interface: typed calibrated decisions from small models, with a routing benchmark and a distilled 19 ms encoder

PythonDocumented

poker-jev-test-bench

Xy2002/poker-jev-test-bench
Research

Jev test bench — Texas Hold'em edition: live-fire testing of TypeSafe's Jev evaluation model through a React poker game (Vercel AI Gateway). MIT.

JavaScriptDocumented

semantic-firewall

CeamKrier/semantic-firewall
Research

Semantic firewall for LLM agents: tool calls gated by TypeSafe Jev (System One decision model via OpenRouter) + deterministic policy. PoC with corpus, stability eval, baseline, results.

TypeScriptDocumented

system-one-security-triage

Robertzu43/system-one-security-triage
Research

Recorded comparison of Jev, Terra, and Opus on 100 synthetic security-triage cases, five passes each, with a static inspectable dashboard

TypeScriptDocumented

decisionbridge

grishahq/decisionbridge
Research

A Jev-inspired decision interface for existing LLMs. Explicit choices, scores, calibration, and review thresholds.

HTMLDocumented

jev-benchmark

themsquared/jev-benchmark
Research

Reproducible benchmark for TypeSafe AI's Jev on agent tool-call risk classification: accuracy, latency, and whether the confidence score is worth routing on.

PythonDocumented

jev-secret-detection

teyhouse/jev-secret-detection
Research

Measures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets

PythonDocumented

jev-snake

MoonTory/jev-snake
Research

TypeSafe's Jev model plays Snake live in the browser, with a headless eval and a self-improvement loop

TypeScriptDocumented

padflow-jev-evals

zsavage8/padflow-jev-evals
Research

Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.

PythonDocumented

pi-jev-tools

Jabbslad/pi-jev-tools
Research

TypeSafe ranking, classification, retrieval and structured-decision tools for Pi coding agents

TypeScriptDocumented

RISC-jeV

i2cjak/RISC-jeV
Research

I tortured Jev into being a RISC-V CPU.

PythonCode reference

jev-synergy-screening

PistachioAIHQ/jev-synergy-screening
Research

Jev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels

PythonDocumented

Olave

stalinzbb/Olave
Research

Evaluation System for DG Fundraising Cause

TypeScriptDocumented

jev-benchmark

haxudev/jev-benchmark
Research

面向语义决策模型的中文言下之意评测集:100 道伴侣对话 Choice 题,支持本地 ONNX、Jev 官方 API 与内网模型对比

Jupyter NotebookDocumented

jev-cloud-cost-guardian

JevForge/jev-cloud-cost-guardian
Research

Evaluate cloud spend against a budget and gate CI with Jev (approve, warn, block, or review).

TypeScriptDocumented

jev-laya-tetris

HarryReidx/jev-laya-tetris
Research

⚡ JEV vs LAYA · Next-Gen AI Tetris Duel Benchmark Platform (TypeSafe Jev vs Local GPU ModernBERT) with React-Bits UI/UX

JavaScriptDocumented

jev-security-prioritization

san3ncrypt3d/jev-security-prioritization
Research

Reproducible experiment: can TypeSafe's Jev decision model prioritize SCA and SAST findings from context? Frozen benchmark, 50k scale run, frontier-model comparison, raw data and blog.

PythonDocumented