TypeSafe Jev shadow-evaluation plugin for Rspamd with optional GPT provider comparison
Benchmarks & Research for Jev — page 18
474 repositories with documented relationships and source evidence.
Demo de un interceptor de fraude simulado que compara en paralelo un modelo Sistema 1 (Jev, TypeSafe AI) con un LLM (Gemini) sobre transacciones sintéticas: velocidad, costo por llamada y decisiones. Proyecto personal de experimentación, no es un benchmark.
A small CLI that uses the TypeSafe API (Jev model) to evaluate emails
A 60-game benchmark of TypeSafe's Jev evaluation model playing Battleship. The model matches plain code; it does not beat it.
A learning scaffold for TypeSafe AI's System One models: eval harness plus a measured, plain-language comparison of the Jev decision model vs an LLM stand-in on 24 real operational decisions. All numbers reproducible from committed run files.
Runtime-agnostic algorithms built on TypeSafe's Jev structured-evaluation model
A traffic city where every car is driven by TypeSafe's Jev model, benchmarked against a rule-based driver.
Benchmarking TypeSafe Jev against 15 chat-model configurations on support-ticket classification: latency, tokens, cost, accuracy.
A 3D driving simulator where Jev, TypeSafe's decision model, chooses what the car does. Code eye or Gemini camera eye, code reflexes, an evaluation harness.
Architecture for running many semantic decisions through one validated path, with a labelled 111-case benchmark comparing TypeSafe Jev against a Claude baseline, and a dashboard for inspecting any single decision. Experimental, not production.
Classification benchmark: Jev vs Claude Sonnet 5 vs Claude Opus 5 on AG News
Benchmarking TypeSafe Jev against general-purpose LLMs on support-ticket routing, with a focus on latency, accuracy, and confidence.
Standalone Go CLI and MCP server for TypeSafe Jev: typed judgments, JSONL evaluation, and resumable batches.
Evaluate Jev (TypeSafe AI) on Japanese customer-inquiry data: 150-item dataset + comparison script (Jev / LLM / rule-based)
Experiments with TypeSafe's Jev: triage benchmark vs LLM, and a source study of jev-ultrafast vs Cline's jev-browser
Small apps for evaluating TypeSafe AI's System One model (Jev). Unofficial.
Triage de leads de WhatsApp/CRM con Jev (TypeSafe AI): ruteo por confianza, plantilla n8n y benchmark en español. Sin dependencias. No afiliado.
Measured cost, latency and raw output from the live Jev API (TypeSafe AI System One model) across 8 use cases — reproducible
Does ORDER BY over a Jev probability put rows in a defensible order? Independent ranking, calibration and invariant measurements of TypeSafe AI's Jev: passes six pre-registered gates on 360 labeled rows, fails four of six on graded product relevance.
CLI for evaluating Microsoft Outlook email with TypeSafe AI’s Jev model.
Test harness + benchmark for TypeSafe's Jev decision model (noul/choice/score) via OpenRouter's Decisions API
Reproducible Tetris decision benchmark comparing TypeSafe Jev with Claude Haiku
jeval: open-source evaluations for AI outputs and agents, judged by Jev