Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
jev-ood-calibration
Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06.
Project facts
- Relationship to Jev
- Research
- Evidence
- Documented
- Language
- Python
- License
- MIT
- Origin
- Original repository
- Repository status
- Not archived
- Created
- 2026-09-19
- GitHub stars
- 6
- Evidence checked
- 2026-09-24T08:40:10.508Z
- Metadata checked
- 2026-09-24T08:40:10.508Z
- Check status
- current
Stars measure the whole repository, including work unrelated to Jev.
Evidence and scope
Documented records the linked documentation or source. JevHunt has not independently run or benchmarked this project.
All calls through Vercel AI Gateway (model id `typesafe-ai/jev`; the `typesafe-ai/jev-latest` id in the AI SDK docs returns "Model not found" on the Gateway), AI SDK 7.0.107 `experimental_evaluate`, `zeroDataRetention: true`, 2026-09-19. 3,721 public-benchmark items + 900 synthetic items, 0 failed calls. The Gateway does not expose a model version; every response resolves to canonical slug `typesafe-ai/jev` and declares `rounding: {probabilityDecimals: 2, scoreDecimals: 2}`, which is where the 0
Evidence commit: 914d87ab16517f14e833cff63d73b51235b14bd6
Discovered through: awesome-jev, typesafe-field-guide.
How we review →