hermes-jev-context-engine

jasonjeske/hermes-jev-context-engine
Research

Experimental selective context compaction for Hermes Agent using TypeSafe Jev. Native plugin, local archives, bounded scoring and honest evaluation.

PythonDocumented

jcol

keltokhy/jcol
Research

Apply natural-language codebooks to tables: a CLI and Python API with resumable annotation, exports, and label evaluation.

PythonDocumented

jev-inner-speech-bci

Dililianxice/jev-inner-speech-bci
Research

A reproducible benchmark connecting Jev semantic priors with intracortical inner-speech BCI decoding.

PythonDocumented

jev-screening-benchmark

Saeedabdf/jev-screening-benchmark
Research

Benchmark of TypeSafe Jev (System One decision model) for systematic review title/abstract screening vs GLM/Sonnet on Cohen_2006 gold standard

TypeScriptDocumented

jev-sim

dashbi1/jev-sim
Research

Jev-compatible /v1/systemone server reading typed decisions from LLM logits, benchmarked against TypeSafe's Jev on the same items via JevBench

PythonDocumented

decision-bench

Hanno-Labs/decision-bench
Research

Open benchmark runtime for document-grounded decision models

PythonCode reference

jev-bench

TheWayWithin/jev-bench
Research

Does the cited source actually say it? A 42-claim benchmark: Jev (TypeSafe System One) against GPT-5.4, Claude Sonnet 5 and Gemini 3.1 Pro.

PythonDocumented

jev-score

a-Fig/jev-score
Research

Local-first document evaluation workspaces powered by Jev

JavaScriptDocumented

jev-the-spire

alexmeckes/jev-the-spire
Research

Watch TypeSafe Jev play Slay the Spire 2, with a local decision dashboard and offline evaluations.

JavaScriptDocumented

jev-vs-open-decision-models

elcronos/jev-vs-open-decision-models
Research

Zero-shot benchmark of TypeSafe Jev 1.13 (decision model) vs open-weight non-generative models PrismNLI-0.4B and Laya: frozen protocol, raw predictions, calibration, latency, report

PythonDocumented

mcp-server-jev

MattiooFR/mcp-server-jev
Research

Typed AI decisions for Codex, Claude and any MCP client, powered by TypeSafe Jev. Classify, score and evaluate with one generic tool.

JavaScriptDocumented

paper-radar-jev

LYchoon/paper-radar-jev
Research

An automated research paper radar that fetches the latest papers from arXiv, evaluates their relevance to a configurable research profile using TypeSafe AI, and ranks them by relevance score. Designed for personalized, daily literature discovery across different research domains.

PythonDocumented

what-is-jev

g0runmezadam/what-is-jev
Research

Independent, source-linked research on TypeSafe AI's Jev (System One), with 947 rubric-scored public repositories, recurring patterns, datasets, and bilingual documentation.

PythonDocumented

cc-mod-jev

JayDoubleu/cc-mod-jev
Research

Claude Code mod: Jev-scored context pruning through OpenRouter. Verbatim compaction, an optional gate on oversized tool outputs, tests and evals.

TypeScriptDocumented

decide

alsoleg89/decide
Research

Bulk decisions for AI agents. Jev classifies files and logs; your agent reviews exceptions. Reproducible cost and quality benchmarks.

PythonDocumented

jev-1.13-mini-benchmark

alperenerol/jev-1.13-mini-benchmark
Research

Mini benchmark of TypeSafe's jev-1.13 structured decision model (OpenRouter Decisions API) on labeled support-triage: noul/choice/score, consistency, cost, lessons learned

PythonDocumented

jev-bbq-experiment

simonmesmith/jev-bbq-experiment
Research

Reproducible evaluation of TypeSafe Jev on all 58,492 BBQ questions: accuracy, stereotype bias, uncertainty, cost and latency.

RDocumented

jev-lab

llt22/jev-lab
Research

Hands-on research lab for TypeSafe's Jev (System One model): reproducible benchmarks of Noul/Choice/Score primitives, confidence gating, fan-out latency, agent control — plus a living audit of the Jev ecosystem.

PythonDocumented

jev-labs

copyleftdev/jev-labs
Research

Never confidently wrong: a TLA+-verified consensus kernel around TypeSafe's Jev, run through 1,680 chaos-tested pharmacy decisions with zero wrong verdicts. Film, code, and every captured call.

PythonDocumented

jev-test

clduab11/jev-test
Research

Pre-registered benchmark: can a 2B local model (Gemma 4 E2B) answer web questions without making things up when a decision model (TypeSafe Jev) makes every call? SearXNG for search, MemPalace for verbatim memory, seven arms including open local judges. Spec and thresholds fixed before any run.

PythonDocumented

jevaluate

ElshinQ/jevaluate
Research

Jevaluate: evaluate before you trust. Field notes, runnable scripts and an agent skill for TypeSafe Jev: gated evals, a browser loop, a product walk with DeepSeek vision, a UI text judge and a first-click tree test. Co-authored with Claude Fable 5.1.

JavaScriptDocumented

jevcheck

sathariels/jevcheck
Research

Behavioral contracts for TypeSafe Jev — pin production expectations, eval model upgrades, catch flips and confidence regressions.

PythonDocumented

kelpie

seahsky/kelpie
Research

Delegation policy for Claude Code, cut down to what its own benchmark supports: two pinned roles, and a skill that argues against delegating by default.

JavaScriptDocumented

jev-carryforward

Dharundp6/jev-carryforward
Research

What your last session knew, scored against what this one is doing. MCP server: a per-project ledger written as things happen, recalled per task with TypeSafe's Jev evaluation model via Vercel AI Gateway.

TypeScriptDocumented