jev-system-one-playground

adyoi/jev-system-one-playground
Research

Windows Forms playground for Jev System One — visualize and evaluate TypeSafe primitives (Choice, Score, Noul) with a 3-pane editor, preset builders, and simulated API evaluation.

C#Documented

jev-vs-cosine

Madheshvivekanandan/jev-vs-cosine
Research

Do you need TypeSafe's Jev, or is cosine similarity enough? A $0, reproducible benchmark on Banking77 intent routing.

PythonDocumented

JevExperiment

Bernardbyy/JevExperiment
Research

Jev vs LLMs: benchmarking a decision model against small LLMs on accuracy, latency and cost.

PythonDocumented

jevkit

JasmineAIGC/jevkit
Research

Policy + evaluation framework for Jev / Kev (TypeSafe System One): typed probabilities in, auditable actions out — calibration, budget-compiled thresholds, permutation & drift checks. Zero-dependency core, kev-native.

PythonDocumented

korean-decision-benchmark

jkf87/korean-decision-benchmark
Research

Reproducible Colab benchmark of SemIf, Decider, Laya and TypeSafe Jev on Korean hate speech; pinned models, isolated environments, MIT code.

PythonDocumented

adaptive-rag-router-typesafe-jev

wimaniac/adaptive-rag-router-typesafe-jev
Research

Adaptive RAG Router using TypeSafe Jev, with budget constrained evaluation against rule based and LLM routers.

PythonDocumented

ai-decision-engine

sergio-lim/ai-decision-engine
Research

Give your agents typed AI decisions instead of parsing LLM text. Jev (TypeSafe) classifier with confidence-gated routing + regex baseline benchmark.

PythonDocumented

Biased-Decisions

AnthusAI/Biased-Decisions
Research

A benchmark harness for bias in fast decision models: causal counterfactuals on real text, with control floors, reported in the decision's own currency

PythonDocumented

date-with-jev

rishi-raj-jain/date-with-jev
Research

Anonymous, share-card-first dating-chat evaluator: upload a screenshot, ask Jev for the read, and get the tea. Built on Next.js, Neon (Lakebase Postgres + Object Storage), and Vercel.

TypeScriptDocumented

jev-audit

neozhu/jev-audit
Research

AI-powered contract comparison with Jev atomic evaluations—spot substantive changes, filter OCR noise, and generate reviewable audit reports.

TypeScriptDocumented

jev-bench

PavelRavvich/jev-bench
Research

Independent benchmark of TypeSafe Jev (System One) vs cheap and frontier LLMs: accuracy, calibration, latency, cost

PythonDocumented

jev-benchmark

ejs-5/jev-benchmark
Research

An independent benchmark of TypeSafe's Jev on 868 real decisions from a public repo, with mechanical (non-model-generated) labels.

PythonDocumented

jev-benchmark-demo

Correa-Gui/jev-benchmark-demo
Research

Jev (TypeSafe AI) vs LLMs: benchmark + demo web ao vivo comparando latencia, custo e concordancia

PythonDocumented

jev-dj-router

panchicore/jev-dj-router
Research

Learning TypeSafe AI's Jev evaluation model fast by building a /play request router for a DJ chatbot

PythonDocumented

jev-eval-harness

laszloblum/jev-eval-harness
Research

LLM-as-a-judge evaluation harness: a confidence-gated cascade (fast typed judge → Claude) that scores LLM outputs and gates CI on a regression diff. TypeScript, real recorded transcripts, real benchmark numbers.

TypeScriptDocumented

jev-eval-lab

sureshbujji/jev-eval-lab
Research

Eval harness for TypeSafe Jev decision models: golden-dataset evaluation, calibration metrics, and schema-guarantee tests for typed decisions (Choice / Score / Noul)

PythonDocumented

jev-evaluation

leecoder/jev-evaluation
Research

Standalone TypeSafe Jev evaluations on 2026 Korean CSAT and SAT Practice Test #11

PythonDocumented

jev-fanout-bench

blowxian/jev-fanout-bench
Research

Measured: asking TypeSafe Jev N questions in one call bills the state once. 2,976 real requests, raw data, exact billing check.

PythonDocumented

jev-lab

eyesofish/jev-lab
Local alternative

Local playground for decision models in the shape of TypeSafe's Jev, running the open openjev replication (Qwen3.5 as an NLI cross-encoder). CPU-only, no API key.

PythonDocumented

jev-replay

poisson-labs/jev-replay
Local alternative

Measures whether TypeSafe's Jev decision model flips actions on identical replays near shipped thresholds across 12,705 calls. Replication data, harness, and independent verification.

PythonDocumented

jev-router-lab

Suryals/jev-router-lab
Research

Shadow-mode eval of TypeSafe Jev as an AIOps triage router vs Qwen3.8-27B and Claude Sonnet 5

PythonDocumented

jev-skillbench

MohibShaikh/jev-skillbench
Research

Benchmark of TypeSafe's Jev as a malicious agent-skill detector on MalSkillBench, with a verify-and-escalate cascade

PythonDocumented

JevApi

JawzoD3TH/JevApi
Research

JevApi — .NET tooling for the TypeSafe AI Jev model, which evaluates typed questions (noul/choice/score) against a state and returns structured answers with probabilities and confidence.

C#Documented

LAYA-RLCD

shyamsridhar123/LAYA-RLCD
Research

🎮 Play. Learn. Respawn. Turn LAYA & ModernBERT Decoder into tactical AI pilots in Cinder Station, an original Doom-inspired game. Explore RLCD + CE, TypeSafe Jev, interactive lessons & transparent benchmarks. Train locally or in Colab. Fork it. Train it. Play it.

Jupyter NotebookDocumented