Text/PDF - File categorization and sorting with Typesafe AI Jev or local calibrated decision model
Benchmarks & Research for Jev — page 8
474 repositories with documented relationships and source evidence.
DSH plugin: register TypeSafe Jev (System One decision model) as an agent tool — jev_decide returns calibrated probabilities (noul/choice/score) for routing/triage/guardrail judgments, no text generation. 把 TypeSafe Jev 决策模型注册为 DSH agent 工具
Chess moves, evaluations, persona opponents, and game classification with TypeSafe AI System One
Chain-of-thought and self-refinement for TypeSafe's Jev: feed its typed answers back as state and ask again. Benchmarks vs TypeSafe's own cookbook numbers.
Fast, cheap judgment for AI coding agents: semantic search, focused reads and list picking in ~2s. CLI + MCP server on TypeSafe Jev. Benchmarked on SWE-bench.
FUn little experiment with Typesafe AI Jev Model playing chess against stockfish :)
Pre-registered independent eval of TypeSafe Jev against a nano-class LLM, a frontier LLM, and a supervised encoder (Banking77 + CLINC150 zero-shot)
Rubric-based eval harness cheap enough to run on every PR, powered by typesafe-ai/jev
I kept watching coding agents burn context on decisions that aren't hard - triage 400 tickets, tag 600 files, route to one of six teams. jev-mode moves those verdicts to a typed-judgment model. I A/B'd it: 78% fewer tokens, 16x less work-attributable input, accuracy 96.1% vs 93.7%. Python, no deps, MIT.
Train lightweight language backbones for typed decisions and candidate probabilities. CE/Brier training, evaluation, and an offline end-to-end demo.
Fast and cheap agent evals. jev as judge.
Agente de trading para o mercado BTC Up/Down de 5 minutos da Polymarket: modelo em código, Jev (TypeSafe System One) como portão, ordens maker, calibração e shadows em paper
TypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers
TypeSafe System One models (Jev) for RubyLLM: typed judgments, evaluations and reranking.
Open auto mode for AI agents — a calibrated tool-call firewall powered by TypeSafe Jev. Ships as a Claude Code hook
Autonomous System-One Triage Engine & Benchmark powered by TypeSafe AI (Jev). 75ms inference, $0 output tokens, and RLCD epistemic safety gates.
Measure when to use Jev and other models on your data, then route accordingly.
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
An eval-first MCP server for TypeSafe's Jev, a System One model that returns typed judgments (noul, choice, score) with probabilities instead of generated text.
Reproducible Jev Ultrafast research-browser eval harness + field note (QC’d cases, suite runner, report generator). Not investment advice.
Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena
Zero-shot spam filtering with TypeSafe Jev Noul questions, compared with TF-IDF baselines
Typed, policy-driven decision workflows on top of TypeSafe AI Jev: confidence routing, fallbacks, evaluation, and RAG patterns for TypeScript apps.
An observable raw-character chat experiment powered entirely by TypeSafe Jev Choice