Discriminative Monte Carlo Tree Search using System One and Harnesses
Benchmarks & Research for Jev — page 5
474 repositories with documented relationships and source evidence.
A verification layer for AI evaluations. Checks the instrument, not just the score: data, scorer, runs, numbers, claims, and itself.
Tests, calibration audits and failure-mode studies of Jev (TypeSafe System One): jaggedness, consistency, injection, abstention.
Benchmark TypeSafe JEV against LLMs, fine-tuned BERT, Laya and zero-shot NLI on text classification: accuracy, calibration, latency, throughput, cost
An instrumented 2048 web lab where every move is a Jev (TypeSafe AI System One) Choice, with no heuristic fallback | 用 Jev 决策模型驱动每一步的 2048 网页实验台,概率、置信度、延迟与成本全部摊开可见,且刻意不做启发式兜底
Jev, TypeSafe's System One model, plays chess against any OpenRouter LLM, Stockfish and you. One-page web app with live moves, Jev's move probabilities, saved games and win rates.
Claude Code plugin that scores review findings, debug hypotheses and design options with TypeSafe's Jev — calibrated probabilities instead of one more opinion.
OpenJev: an independent Jev-inspired System One decision API based on TypeSafe.ai concepts. Choice, score and noul primitives, local mock server, Python and TypeScript SDKs. Real inference planned; not affiliated with TypeSafe AI.
Make room for useful evidence. Inspectable context selection for RAG, with Jev reranking and open benchmark studies.
A show-and-tell capability study for Jev, TypeSafe's System One decision model.
Jev (TypeSafe) vs Claude Haiku 4.5 on 2 000 phishing emails: accuracy, calibration, latency, cost. Reproducible benchmark.
LegalForecast-MTD benchmark alpha and official evaluation workflows
Self-hosted Jev decision workbench and Feishu bot: automatic choices, probabilities, and experimental word/character writing.
First independent head-to-head benchmark of System One decision models (Laya vs Jev) on byte-identical inputs
An OpenAI-compatible proxy that sits between your LLM and your users. It evaluates each sliding window of tokens while the response is still streaming and cuts the stream before a violating token can reach the screen.
Jev 1.13 reward-model evaluation across 8 benchmark tracks, with an interactive report and 54-row SOTA comparison
Abandoned Qwen3.5-0.8B LoRA fine-tuning experiments for Jev-like typed judgments, with datasets, adapters, evaluations, and a full retrospective.
Jev-compatible System 开源Jev
Community-maintained Go SDK for the TypeSafe AI System One evaluation API, with typed questions, fluent builders, retries, and examples.
Catch breaking API behavior hidden in OpenAPI prose with deterministic checks and TypeSafe JEV System One semantic review.
Open-source BYOK arena for Jev and other AI judges. Find failures, compare quality, cost, and latency.
Experimental Jev permission gate for Claude Code via OpenRouter, with reproducible latency and cost benchmarks
Calibrated alignment verifier for LLM responses and agent plans — powered by Jev
Audit your git diff against YAML coding-standards packs using TypeSafe's Jev model, from a CLI or your AI agent's command/skill.