typed_evals

TrustifAI/typed_evals
Research

Fast, typed, calibrated evaluations for LLM and agent outputs, powered by Jev — with simple, framework-agnostic Python APIs

PythonDocumented

jev.nu

cablehead/jev.nu
Research

Nushell module for the TypeSafe System One API: typed decisions with calibrated probabilities

NushellDocumented

jev-dspy-lab

jmanhype/jev-dspy-lab
Research

Reproducible calibration and selective-risk benchmarks for Jev/TypeSafe decisions in DSPy workflows

PythonDocumented

trade-jev

justinhe16/trade-jev
Research

Backtest Jev (TypeSafe) as a BUY/SELL/HOLD trader on NQ L10 order-book data

PythonDocumented

JEV-Paper-Radar

Eliot5566/JEV-Paper-Radar
Research

Let Jev read every new arXiv paper each morning and surface the few you should read. Plain-English interests, calibrated probabilities, ~$0.06/day, fork and go.

PythonDocumented

Kev

arjun988/Kev
Research

Open-source System One decision engine. Typed choice / score / noul with calibrated probabilities. Self-host with Ollama or any OpenAI-compatible model. Apache-2.0.

TypeScriptDocumented

laya-browser-agent

ChenneyZhuang/laya-browser-agent
Local alternative

Local, open-source Jev alternative: browser agent decisions with Laya (System One model) on your own machine. No cloud, no API key. Playwright/CDP, MCP-friendly.

PythonDocumented

luce

scienthoon/luce
Research

Luce: a recipe for calibrated decision models — a sentence about your task in, a small model that answers typed questions with honest probabilities out (init → synth → train → eval → serve)

PythonDocumented

nanojev-arena

caijinchun/nanojev-arena
Local alternative

NanoJev Snake Arena: 1v4 human-vs-AI battleship + 100-agent swarm simulator. Local demo of Jev System-One model (open-source mini replica).

HTMLDocumented

poorjev

rupeshpoojary9/poorjev
Local alternative

Open-source, local Jev alternative: a System One decision layer with provably calibrated confidence (ECE 0.170→0.071). Typed decisions, runs offline, no API key, no waitlist.

PythonDocumented

pi-jev-context

Nyarlathoteppppp/pi-jev-context
Research

Model performance first. Token savings second. A Pi extension with freshness-aware read dedupe, Jev log filtering, and searchable verbatim recall. Keeps existing message history intact.

TypeScriptDocumented

jev-rerank-bench

anessbelbati/jev-rerank-bench
Research

Can a decision model beat dedicated rerankers? TypeSafe Jev vs Cohere Rerank 4 vs ZeroEntropy zerank-2 vs a chat-model baseline: 14 datasets, every raw API response, bootstrap ranges on every gap.

PythonDocumented

learn-jev-end-to-end

harshithsunku/learn-jev-end-to-end
Jev application

Learn Jev end to end: a free hands-on course. Build 13 AI agent use cases with a fast brain (Jev) and a slow brain (LLM). One OpenRouter key.

Jupyter NotebookDocumented

jev-dimabsa

ZhangYiqun018/jev-dimabsa
Research

TypeSafe Jev baseline for DimABSA (SemEval-2026 Task 3) subtask 1: zero-shot and 3-shot valence-arousal regression

PythonDocumented

jev-architect

karanb192/jev-architect
Research

Find, design, and evaluate TypeSafe Jev decision loops.

HTMLDocumented

Foq

yohanargentina-oss/Foq
Local alternative

⚡ Foq — the FREE, local, open-source alternative to Jev. Typed System 1 decisions in ~25 ms — no waitlist, no cloud, no per-token cost. foq.fr

PythonDocumented

jev-ood-calibration

scienthoon/jev-ood-calibration
Research

Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06.

PythonDocumented

jev-search

larguesa/jev-search
Research

Experimental semantic line search with TypeSafe Jev via OpenRouter. Python CLI with no runtime dependencies.

PythonDocumented

jev-search-rerank-eval

zhuyansen/jev-search-rerank-eval
Research

Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.

PythonDocumented

ruling

bradAGI/ruling
Research

Typed, calibrated decisions from a local model. No text generated.

PythonDocumented

daf-jev

docxology/daf-jev
Research

daf-jev: composable Python toolkit for TypeSafe's Jev (System One) decision API — question builders, confidence gates, evaluator, calibration, CLI, MCP server, agent skill

PythonDocumented

jev-korean-benchmark

mahlernim/jev-korean-benchmark
Research

Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence

PythonDocumented

jev-benchmark

wondertwins/jev-benchmark
Research

Benchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs

PythonDocumented

jev-lm

y0usaf/jev-lm
Research

A word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval

TypeScriptDocumented