Evaluation and calibration toolkit for TypeSafe Jev decisions
Benchmarks & Research for Jev — page 16
474 repositories with documented relationships and source evidence.
Jevtrieval is a retrieval-augmented generation demo. It retrieves documents from Qdrant, scores their relevance with Jev, and generates an answer from the selected documents.
Real-time last-mile delivery exception triage: Jev (TypeSafe AI System One) evaluates each event, deterministic guardrails decide. FastAPI + Streamlit.
LLM movie recommender using Jev (TypeSafe AI) as a calibrated reranker, with an offline eval against an LLM judge
Local Nevada statute search and a reproducible evaluation of Jev-assisted ranking against keyword search.
An independent, reproducible benchmark of TypeSafe's Jev against open, CPU-only alternatives — 10,000 decisions, all raw results published.
Reproducible long-workflow benchmarks for Pi Sieve, with Astra and Luna comparisons.
Comprehensive poker game theory knowledge base, instruction-tuning dataset, and TypeSafe AI (Jev) evaluation benchmark for Limit & No-Limit Texas Hold'em.
Compare question formulations with held-out evaluation, bounded experiments and DecisionPack export.
Candidate-experience feedback triage showcase using TypeSafe Jev, with synthetic data, human review, and measured evaluation.
Monitor semantic conditions with exact-input caching, hysteresis and source/evaluator plugins.
Model Context Protocol (MCP) server & OpenCode skill for San Francisco Early Learning For All (ELFA) preschool search, voucher calculations, and TypeSafe Jev System One evaluation
Benchmark riproducibile di TypeSafe Jev via API diretta/OpenRouter contro un modello locale MLX/Qwen su BANKING77.
Hybrid System 1 (TypeSafe Jev decisions) + System 2 (Claude Fable) demo with confidence-gated routing and cost/accuracy benchmarks.
A small RAG web app that visualizes every step of the retrieval-augmented generation pipeline
Large typed decision model evaluation benchmark and novel calibration standard (`calibration.json`) for any typed decision inference system, covering 275 distinct use-cases and 27,598 individual decisions for Jev-like open typed decision models.
Modular Python integration layer for TypeSafe Jev, with a typed evaluation client, CLI, HTTP API, Docker deployment, confidence policies, and reusable Choice, Score, and Noul primitives.
Federated knowledge retrieval engine for AI agents. Routes queries across specialized backends (FTS5, ChromaDB, graph) via a registry-first orchestrator with TypeSafe/Jev AI routing and evidence re-ranking. Read-only by default, security-gated by design.
TypeSafe Jev (System One) as DeepSeek Harness agent tools, with the token accounting and benchmark to prove what they cost
Jev (TypeSafe AI) evaluation plugin for Hermes Agent — structured decision/classification/routing tool via Vercel AI Gateway
Independent, pre-registered evaluation of TypeSafe AI's Jev. 5,721 calls, 21 experiments, 50 predictions registered before collection. Raw data included.
Real-time ad ranking and creative review on a typed evaluation model — one call prices a whole auction, no trained CTR model, no logged clicks.
Code when exact. Jev when useful. Agent when uncertain. A lightweight routing layer for computer-use agents, with real workflow benchmarks.