MCP server for TypeSafe Jev (System One): one typed evaluate tool over streamable HTTP
Benchmarks & Research for Jev — page 14
474 repositories with documented relationships and source evidence.
TypeSafe's Jev and four fast LLMs added to Stanford MedHELM's MedHallu results: harness, preregistered run plans and every run file.
Review public code and documents with Jev in Codex and Claude Code: ranking, batch screening, claim checks, spend limits, and evaluation tools.
Weekly builds on Jev (TypeSafe AI), benchmarked honestly enough to publish. Week 01: inbound lead triage, 90% routing at 366ms and $0.04 per thousand leads.
Probabilistic soft-constraint evaluation for volunteer scheduling using TypeSafe Jev, with deterministic application policy.
Claude Code skill: integrate TypeSafe's Jev (System One model) into JS/TS projects — primitives, SDK, patterns, runnable example, triggering evals.
Multi-agent customer support API where the LLMs write the text and a typed decision model makes the decisions. Jev (TypeSafe System One) handles routing, response evaluation and retry/escalation; FastAPI + Ollama. Includes a 60-ticket benchmark against LLM free-text and structured-JSON baselines.
Empirical benchmarks, task telemetry, and engineering thesis for Jev (TypeSafe AI) System One in autonomous agentic organizations
JEVChess is an elegant, minimal chess platform designed around the chess.com aesthetic, It delivers rapid tactical responses using TypeSafe AI's **Jev** System-One model and heuristic fallback evaluation, without requiring any external database.
Behavioural anti-cheat research bench for Minecraft (Paper): mining telemetry, typed LLM questions, and an offline benchmark against classic X-Ray heuristics. Shadow mode only, never bans.
Head-to-head evaluation of Laya (open weights) and TypeSafe Jev on email intent classification
Benchmark de triagem de revisões sistemáticas: JEV (TypeSafe) × DeepSeek V4 Flash × Qwen3-Reranker, contra o julgamento documentado de autores reais (SYNERGY+ v3.0)
JEV-inspired typed runtime decisions over long-horizon PSG: reusable overnight encoding, sparse retrieval, and Choice/Noul/Score probabilities.
Benchmark for typed System One decision models (Jev, Laya, and whatever comes next): calibration against a noise floor, framing sensitivity, selective prediction, ordinal fidelity, interference, robustness. pip install sys1bench.
Local alternative to Jev: typed decisions (choice/score/yes-no) with real probability distributions, from any llama.cpp-supported GGUF model. Easy CLI for humans and AI agents. No generation, no fine-tuning, no compiler.
Benchmarking decision models (TypeSafe Jev, Convai Laya) on the 2026 Turkish university entrance exam — 593 questions, full controls, uncontaminated test set
Jev-powered decision workflows for Codex image generation. Explicit constraints, bounded repairs, and honest evaluation.
Hosted Jev policy workspaces with tenant isolation, finite evaluations and revocable aggregate reports
Score written content on house style, editorial quality, and audience fit. Regex linter for mechanical rules; semantic dimensions via TypeSafe's Jev model.
Stranger danger for AI agents: detect prompt injection in content your agent fetches, before it reads it. Runs on TypeSafe's Jev System One model by default; local and keyword detectors included. Benchmarked, failures published.
Independent 100-question English benchmark of TypeSafe's Jev decision model (jev-1.13.0)
A runnable decision-first agent evaluation example using Jev and TypeSafe AI.