jev-robotics-eval

lose4578/jev-robotics-eval
Research

Evaluate JEV robot control in MetaWorld and RoboTwin: discrete actions, hierarchical control, privilege-level ablations, and replay analysis.

PythonDocumented

jev-zig-cli

SupratimSircar05/jev-zig-cli
Research

Unofficial Jev-powered Zig terminal agent with deterministic policy gates and encrypted audit trails.

ZigDocumented

DecisionEngine

swisnl/DecisionEngine
Research

Typed, probabilistic decisions (Choice / Score / Noul) for System One tasks, powered by TypeSafe Jev, with OpenAI and Anthropic as alternative engines

PHPDocumented

firefox-jev-mcp

kitoutou999/firefox-jev-mcp
Research

Serveur MCP pour piloter Firefox avec Claude : Jev (TypeSafe) choisit les clics, Claude planifie. Benchmark inclus.

JavaScriptDocumented

inventio

minhquan23102000/inventio
Research

Local retrieval without embeddings. Every answer comes back as the original passage with its coordinates (path:start-end), so a person or an agent can open the exact lines. No Embedding, Zero Cost, Fully Local.

PythonDocumented

jev-accessibility-tree

vandenbogart/jev-accessibility-tree
Research

An experiment: rebuilding a web page's accessibility tree in under a second with TypeSafe's Jev. Evals + a Chrome extension.

TypeScriptDocumented

jev-amazon-keyword-checker

nexscope-ai/jev-amazon-keyword-checker
Research

Find Amazon keyword ideas with Nexscope and evaluate their relevance to listing text with Jev.

PythonDocumented

jev-tmmluplus-eval

lianghsun/jev-tmmluplus-eval
Research

Evaluate TypeSafe AI's Jev (System One Model) on TMMLU+ v1.1 — four-way choice via the API's own response schema, 100% parse rate by construction

PythonDocumented

ollaya

ollaya-dev/ollaya
Local alternative

Run open decision models locally — pull, run and serve Laya and other open Jev alternatives behind a TypeSafe-compatible API.

RustDocumented

reflexbench

brida-ai/reflexbench
Local alternative

ReflexBench — open benchmark and evaluation harness for System One models and typed decision engines

PythonDocumented

system-one-router

mmornati/system-one-router
Research

A fast System One decision model (Jev or local Laya) picks which LLM answers each prompt: OpenAI-compatible Go gateway + decision benchmark

GoDocumented

toolgate-experiment

diorrego/toolgate-experiment
Research

Reproducible benchmarks of direct MCP and remote tool selection with Go, Rust, TypeScript, and Jev. Compares latency, token usage, estimated API cost, and query correctness.

TypeScriptDocumented

can-jev-play

carlaiau/can-jev-play
Research

Can Jev infer whether a bet is worthwhile from its payout table, and do recent wins or losses sway that choice

TypeScriptDocumented

flagship-jev

karishnu/flagship-jev
Research

Semantic feature flag evaluation for Cloudflare Workers using Jev and Flagship

TypeScriptDocumented

jev-filter

apixly-ai/jev-filter
Research

Semantic filtering for AI tool results. Native npm CLI, caller-controlled context, batching, and reproducible benchmarks.

PythonDocumented

jev-laya-benchmark

EnesDemir143/jev-laya-benchmark
Research

Local benchmark comparing TypeSafe Jev and Laya-MLX for structured issue classification

PythonDocumented

jev-vs-spacy

oluies/jev-vs-spacy
Research

spaCy vs TypeSafe Jev on English spam, support routing and Swedish routing, evaluated in Braintrust

PythonDocumented

nanojev

sshh12/nanojev
Research

A hypothetical reconstruction of Jev, TypeSafe's closed "System One" decision model, in the spirit of nanoGPT: the smallest working version of what black-box probing suggests. It reads a state and answers typed questions (yes/no, pick one, ordered score) with probabilities instead of generated text.

PythonDocumented

openjev-multimodal

jev-skills/openjev-multimodal
Research

Local multimodal decisions on your Mac. Jev-compatible typed probabilities with Qwen, llama.cpp and Metal.

PythonDocumented

rerank-bench-jev

denser-org/rerank-bench-jev
Research

Production reranker benchmark: TypeSafe Jev vs Qwen3-Reranker-0.6B on BEIR SciFact and NFCorpus — nDCG@10, cost/query and p50/p95 latency. Quality is indistinguishable; Jev is ~2x faster, Qwen ~4x cheaper. From the team at denser.ai

PythonDocumented

system-one

mpuig/system-one
Research

An open-source System One decision model — Kahneman's term for the fast, automatic judgment faculty. Calibrated choice/score/noul probabilities in one forward pass on small fine-tuned open models. Local on Apple Silicon (MLX), Jev-compatible API, audited experiment log.

PythonDocumented

yc-jev-bench

PPRAMANIK62/yc-jev-bench
Research

Can Jev replace an LLM reranker? A benchmark of TypeSafe's Jev vs Claude Haiku and BGE on natural-language search over all 6,245 YC companies, with a live search demo.

TypeScriptDocumented

dsh-jev-verify

xienda/dsh-jev-verify
Research

Jev (TypeSafe System One) decision tools + live verification benchmark for DeepSeek Harness: jev_decision (choice/score/noul) and jev_verify, honest by design.

JavaScriptDocumented