jev-arcade

CankatSarac/jev-arcade
Research

Can a System One model play arcade games? TypeSafe's Jev plays Tetris, Snake and 2048 — benchmarked against random and heuristic baselines.

PythonDocumented

jev-banc-francais

gbesse/jev-banc-francais
Research

Evaluate typed AI decisions on labeled French-language cases.

JavaScriptDocumented

jev-benchmarks

GaNotchVFX/jev-benchmarks
Research

Measured benchmarks of TypeSafe's Jev decision model vs LLMs: ticket triage, voice-agent decision layer, bulk tagging

PythonDocumented

jev-certify

nikkoxgonzales/jev-certify
Research

Finite-sample guarantees for Jev (TypeSafe's System One). Conformal risk control turns calibrated probabilities into certified routing thresholds; prediction-powered inference audits them. 2,412 decisions on CLINC150 for $0.23 — including the shift and prevalence cases where the guarantee breaks.

PythonDocumented

jev-clinic-triage-eval

Autometrixai/jev-clinic-triage-eval
Research

Early evaluation of TypeSafe's Jev on clinic call triage: 98.6% accurate when confident, 116k decisions per dollar

PythonDocumented

jev-crash-course

nadeemcite/jev-crash-course
Research

An 11-level crash course on Jev, TypeSafe AI's System One decision model — runnable examples against the real API, plus a capstone project with unit tests and evals. Works with any LLM provider via LiteLLM.

PythonDocumented

jev-crash-course

nadyth/jev-crash-course
Research

An 11-level crash course on Jev, TypeSafe AI's System One decision model — runnable examples against the real API, plus a capstone project with unit tests and evals. Works with any LLM provider via LiteLLM.

PythonDocumented

jev-label

gbesse/jev-label
Research

Prioritize human labels with active-learning strategies, durable review history and clean holdout evaluation.

PythonDocumented

jev-lamplight-house

nicksonthc/jev-lamplight-house
Research

A cottage in React Three Fiber whose eleven working parts are judged, as you flip them, by typesafe-ai/jev — an evaluation model reached through the Vercel AI Gateway.

TypeScriptDocumented

Jev-LLM-Playground

STiFLeR7/Jev-LLM-Playground
Research

Independent playground for TypeSafe AI Jev decision models: typed decisions, support-ticket routing, reproducible evaluations, and a local browser demo.

JavaScriptDocumented

jev-mcp

ThePFMind/jev-mcp
Research

MCP server exposing TypeSafe AI's Jev decision model to Claude (stdio, two tools: jev_evaluate, jev_route)

PythonDocumented

jev-multimodal

Alpha-Harper-Franklin/jev-multimodal
Research

Jev + Multimodal: shared visual decisions, evidence adapters, source audits and reproducible public-image benchmarks.

PythonDocumented

jev-no-enem

patryckalves/jev-no-enem
Research

Reproducible benchmark evaluating TypeSafe AI's Jev (System One paradigm) on Brazil's ENEM 2025 standardized exam. Evaluates typed decision-making, domain-specific accuracy, and RLCD uncertainty calibration against open LLM baselines with an interactive GitHub Pages dashboard.

PythonDocumented

jev-playground

AviroopPaul/jev-playground
Research

A playground for TypeSafe AI's Jev (System One model), built around five real production workflows: support triage, RAG relevance gating, agent action firewall, inline moderation, and CI eval judging.

JavaScriptDocumented

jev-safety-benchmark

mbburabak/jev-safety-benchmark
Research

Benchmark harness evaluating TypeSafe's Jev model on six public safety benchmarks against Shieldstral-1.0-3B and top SOTA guard models

PythonDocumented

jev-skill

abhisheksharma001/jev-skill
Research

Agent skill for TypeSafe AI's Jev decision model: fit assessment, integration recipes, calibration, multi-Jev, benchmarks

PythonDocumented

jev-vs-laya

mouadse/jev-vs-laya
Research

Can an at-home model match Jev? TypeSafe Jev vs self-hosted Laya on Moroccan Darija sentiment, with reproducible zero-shot evaluation.

PythonDocumented

jev-vs-llm

slobodaapl/jev-vs-llm
Research

A benchmark comparing TypeSafe's Jev versus local open-weight and API LLMs, as well as making Jev chat :) (ish)

PythonDocumented

jev-vs-llm-resume-jd-eval

prakash5284/jev-vs-llm-resume-jd-eval
Research

Resume↔JD apply-gate cost eval: TypeSafe Jev vs Claude/Grok/GPT tiers (10 real JDs)

PythonDocumented

jev-webui

chcknnbn/jev-webui
Research

Local-first WebUI for TypeSafe Jev: batch document evaluation, typed questions, formulas, and CSV export. Unofficial.

TypeScriptDocumented

Jevals

Jevals/Jevals
Research

Jevals (jevals.com): independent benchmark of TypeSafe's Jev vs LLMs

Documented

JevKit

Ernosto0/JevKit
Research

JevKit is an open-source decision infrastructure that helps developers integrate Jev into their applications, define structured decision tasks, evaluate model performance, and build reliable fallback policies.

PythonDocumented