A benchmark for Jev's biases, built from people who differ in one attribute at a time.
Benchmarks & Research for Jev — page 17
474 repositories with documented relationships and source evidence.
Runnable question sets for TypeSafe's Jev, an eval harness with measured CLINC150 results, and a linter for the request shapes the API silently mis-reads.
Historical paper-trading simulator for evaluating TypeSafe AI JEV decisions
An adversarial evaluation of TypeSafe's jev decision model: nine experiments and 28 predictions fixed before any data was collected. 123,805 requests, $12.69.
Cheap fail-open semantic edge layer for Jev (TypeSafe System One): decision service + confidence policy + decision logging + blind eval harness + skill router over 1,000+ skills. Stdlib-only Python.
Context-carrying graph retrieval with Jev: branching walks, reproducible ablations, and a visual replay.
Jev-backed guardrails for agent tool calls: library, native OpenCode/Hermes/DeepSeek Harness hooks, and a hosted eval-credit service.
Benchmarking Jev (TypeSafe System One) on 5,000 real CDC survey respondents, with an interactive demo where every profile has a real model answer
Benchmarking TypeSafe's Jev against four rival prompt-injection detectors on 11,900 labelled prompts
Typed client and benchmark harness for Jev, TypeSafe AI's System One decision model, through the Vercel AI Gateway. Measure accuracy, calibration and cost on your own data before you trust a threshold.
Is Jev (TypeSafe System One model) a decent preflop poker player? Reproducible benchmark vs 4 reference styles
Free English RAG benchmark for TypeSafe Jev 1.13 — frozen candidate pools, calibration, paired bootstrap CIs, $0 runs. Compared with OpenJev and NVIDIA cross-encoders.
Local Jev evaluation workbench: datasets, typed questions, threshold simulation and run comparison
Auditing LLM deep-research reports with TypeSafe's Jev: source authority, relevance, and whether cited figures mean what the report says
TypeSafe’s Jev AI + an n-ary memory tree = efficient agentic memory routing and retrieval.
Benchmark: TypeSafe Jev vs traditional LLM on support-ticket routing accuracy, latency, and cost
Jev Workbench: a C++ desktop workspace with Python diagnostics and structured TypeSafe evaluations
Open, auditable reproduction of TypeSafe Jev: a 2B probabilistic decision model benchmarked across six public suites with weights, code, and receipts.
RAGのretrievalにTypeSafe APIを使うアイデアの実装
Use Jev model from TypeSafe to evaluate if there is an improvement in latency and cost in the agentic loop for tool decisioning.
100% clean-room, copyright-free 2D platformer benchmarking TypeSafe Jev AI in real-time control (60 FPS)
Evaluating TypeSafe's System One model Jev through camera-based food allergen estimation (Flutter, iOS/Android)
Real-time creator intelligence powered by direct TypeSafe Jev decisions and measurable full-corpus evaluation.
Playwright AI Benchmark: Klasik Playwright vs JEV (TypeSafe AI) vs GPT-5.6 Luna vs GLM-5.3