Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
jev-medhallu-benchmark
TypeSafe's Jev and four fast LLMs added to Stanford MedHELM's MedHallu results: harness, preregistered run plans and every run file.
Project facts
- Relationship to Jev
- Research
- Evidence
- Documented
- Language
- Python
- License
- MIT
- Origin
- Original repository
- Repository status
- Not archived
- Created
- 2026-09-22
- GitHub stars
- 0
- Evidence checked
- 2026-09-24T08:40:10.508Z
- Metadata checked
- 2026-09-24T08:40:10.508Z
- Check status
- current
Stars measure the whole repository, including work unrelated to Jev.
Evidence and scope
Documented records the linked documentation or source. JevHunt has not independently run or benchmarked this project.
uv run score.py runs/medhallu/test-v2 --threshold 0.65 --ref jev-1.13 --xlsx
Evidence commit: 8a7f2f88eb258fd0bc859ed4aeca3b80643e5d94
Discovered through: github-search.
How we review →