Windows Forms playground for Jev System One — visualize and evaluate TypeSafe primitives (Choice, Score, Noul) with a 3-pane editor, preset builders, and simulated API evaluation.
Benchmarks & Research for Jev — page 12
474 repositories with documented relationships and source evidence.
Do you need TypeSafe's Jev, or is cosine similarity enough? A $0, reproducible benchmark on Banking77 intent routing.
Jev vs LLMs: benchmarking a decision model against small LLMs on accuracy, latency and cost.
Policy + evaluation framework for Jev / Kev (TypeSafe System One): typed probabilities in, auditable actions out — calibration, budget-compiled thresholds, permutation & drift checks. Zero-dependency core, kev-native.
Reproducible Colab benchmark of SemIf, Decider, Laya and TypeSafe Jev on Korean hate speech; pinned models, isolated environments, MIT code.
Adaptive RAG Router using TypeSafe Jev, with budget constrained evaluation against rule based and LLM routers.
Give your agents typed AI decisions instead of parsing LLM text. Jev (TypeSafe) classifier with confidence-gated routing + regex baseline benchmark.
A benchmark harness for bias in fast decision models: causal counterfactuals on real text, with control floors, reported in the decision's own currency
Anonymous, share-card-first dating-chat evaluator: upload a screenshot, ask Jev for the read, and get the tea. Built on Next.js, Neon (Lakebase Postgres + Object Storage), and Vercel.
AI-powered contract comparison with Jev atomic evaluations—spot substantive changes, filter OCR noise, and generate reviewable audit reports.
Independent benchmark of TypeSafe Jev (System One) vs cheap and frontier LLMs: accuracy, calibration, latency, cost
An independent benchmark of TypeSafe's Jev on 868 real decisions from a public repo, with mechanical (non-model-generated) labels.
Jev (TypeSafe AI) vs LLMs: benchmark + demo web ao vivo comparando latencia, custo e concordancia
Learning TypeSafe AI's Jev evaluation model fast by building a /play request router for a DJ chatbot
LLM-as-a-judge evaluation harness: a confidence-gated cascade (fast typed judge → Claude) that scores LLM outputs and gates CI on a regression diff. TypeScript, real recorded transcripts, real benchmark numbers.
Eval harness for TypeSafe Jev decision models: golden-dataset evaluation, calibration metrics, and schema-guarantee tests for typed decisions (Choice / Score / Noul)
Standalone TypeSafe Jev evaluations on 2026 Korean CSAT and SAT Practice Test #11
Measured: asking TypeSafe Jev N questions in one call bills the state once. 2,976 real requests, raw data, exact billing check.
Local playground for decision models in the shape of TypeSafe's Jev, running the open openjev replication (Qwen3.5 as an NLI cross-encoder). CPU-only, no API key.
Measures whether TypeSafe's Jev decision model flips actions on identical replays near shipped thresholds across 12,705 calls. Replication data, harness, and independent verification.
Shadow-mode eval of TypeSafe Jev as an AIOps triage router vs Qwen3.8-27B and Claude Sonnet 5
Benchmark of TypeSafe's Jev as a malicious agent-skill detector on MalSkillBench, with a verify-and-escalate cascade
JevApi — .NET tooling for the TypeSafe AI Jev model, which evaluates typed questions (noul/choice/score) against a state and returns structured answers with probabilities and confidence.
🎮 Play. Learn. Respawn. Turn LAYA & ModernBERT Decoder into tactical AI pilots in Cinder Station, an original Doom-inspired game. Explore RLCD + CE, TypeSafe Jev, interactive lessons & transparent benchmarks. Train locally or in Colab. Fork it. Train it. Play it.