Pluggable, observable RAG MCP server — hybrid retrieval (BM25+dense+RRF), BEIR-benchmarked rerankers (BGE / TypeSafe Jev / LLM), Streamlit dashboard + Langfuse, multimodal ingestion
Benchmarks & Research for Jev — page 13
474 repositories with documented relationships and source evidence.
Ultra-fast, sub-100ms API & webhook guardrail and triage gateway powered by TypeSafe AI System One (Jev). Parallel 7-dimension speculative evaluation, deterministic policy router, zero-dependency SQLite audit trail, and automated outbound dispatch.
S1Rank: Can a System-One decision model (TypeSafe Jev) rerank? Benchmark, raw responses, and paper.
Strands agents that make small local LLMs (Qwen3 0.6B, MiniCPM5 2B, Qwen3.5 4B) answer like Jev, a System 1 model: typed yes/no, choice and score answers with probabilities. Benchmarked against Jev on the jevals suite, via logprob readout and written-out JSON probabilities.
Small, fast, honest decision models. Open alternative to Jev: zero-shot, fit on your labels in seconds, calibrated with a coverage guarantee, Jev wire-compatible.
Deterministic evaluation harness for LLM agents, with a calibrated judge (Jev) that can act as a CI gate.
results from iterative benchmark runs
Jev is actually not a traditional LLM, it doesn’t generate text. It’s what the TypeSafe AI team calls a System One model: 📖 System One models are a class of AI models built to make fast, structured decisions that software can use directly. A System One model evaluates a state and returns typed answers and probabilities.
Open, self-hostable, faster System One decision model — API-compatible alternative to TypeSafe's Jev
Evidence-aware semiconductor alert triage research demo with TypeSafe Jev, policy guards, and reproducible evaluation
JEV evaluation (FM-JEV-01): replay-only, advisory-only evaluation of Foreman-style supervision with TypeSafe Jev. Measures how much safety comes from the model versus a deterministic evidence gate. No worker authority.
This repository contains a GitHub issue classifier built on Jev, TypeSafe AI's System One model. It labels every new issue with typed values and calibrated confidence in milliseconds, labelling what it is sure about and escalating what it is not. Three guardrail layers guard every write, and a frozen eval suite gates each deploy.
Alert triage with Jev (TypeSafe System One): proof of concept and benchmark against a rules-only baseline
Clean-room CLI for TypeSafe System One / Jev Decision Contracts, search, and benchmarks.
typesafe.ai Jev(jev-1.13.0) benchmarked against the 2026 수능 국어영역 — JSON passages/questions + scored results
Production evaluation and gateway tooling for Jev and System One-compatible decision models
Decision Arena: TypeSafe's Jev vs open-source Laya playing highway-env, Snake and Blackjack with zero training, plus benchmarks and a Claude Code watchdog
1,000-decision behavioral benchmark of Jev across 25 deterministic reasoning families.
A hybrid CLI and SPA dashboard leveraging the TypeSafe System One (Jev) API to perform quantitative, rubric-based evaluations of markdown documentation.
Stance test of TypeSafe Jev vs Claude on Taiwan sovereignty questions, in Traditional Chinese, Simplified Chinese and English
在自己的数据上评测 TypeSafe Jev 的准确率、概率校准和可用阈值
975 PLC tag Expressions in Under a Quarter of a Second
Evaluate Jev (TypeSafe System One decision model) on GAOKAO-Bench objective questions - accuracy, confidence calibration and latency over 1497 Chinese college-entrance-exam multiple-choice questions
Typed, confidence-aware Jev routing for LangGraph.js with reproducible benchmarks