Jev: TypeSafe's first System One Model types decisions and prices them two orders of magnitude lower
typesafe-jev
TypeSafe's first System One Model: outputs are type-safe structured values under three question primitives (Noul/Choice/Score) rather than strings, trained with RLCD for calibrated decision probabilities. The official 13-question parallel experiment makes one synthesized call ~12.2x cheaper and ~10.0x faster than 13 separate calls; four-workflow mean accuracy 67.8% at /tmp/run_ingest2.sh.0004 and 0.4s per instance - opus 5 / sol tier accuracy at almost two orders of magnitude less cost; input $0.042/M tokens, output free. Ships an OpenJev local reproduction path and an eight-item caveats list.
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- 四工作流平均准确率(官方评测)
- Vendor Claim · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeOur read: Jev carves "decisions" out of generation - type-safe outputs plus calibrated probabilities - a rare shape innovation at the base-model layer, and its ~2-order-of-magnitude cost gap makes it the default tier for high-frequency small decisions (routing, classification, next-action). The gap is equally clear: 67.8% four-workflow mean versus 73-74% for opus 5 / sol means long-horizon work still belongs to sequential flagships, and the official eval's reference labels come from competitor high-thinking tiers, so read it with a methodology discount. Confidence C (vendor claim): an OpenJev local reproduction path exists, but we have not finished our own harness comparison.
Jev is TypeSafe's first commercial System One Model: instead of free-form strings it returns type-safe structured values under three question primitives (Noul / Choice / Score), so type errors are impossible at the protocol layer. Its training objective is not RLHF or RLVR but RLCD - strictly proper scoring rules that train decision probabilities into calibrated confidences callers can route on directly.
Its context accounting differs from sequential models: the whole state is read once and every question is answered in parallel. The official parallelism handbook runs 13 questions (8 Noul, 2 Choice, 3 Score) over a pinned GDPR Wikipedia full text: one synthesized call is about 12.2x cheaper and 10.0x faster than 13 separate calls, with identical answers; the quickstart page cites 11.5x and 9.6x for the same experiment, same order of magnitude. State plus all questions totals 64k tokens; state plus the longest single question, 32k.
The official eval decomposes tasks into questions plus code rules, with reference labels taken as the mean answer of the two strongest models at the time (GPT-6 Astra and Fable 5.1, both at high thinking). Across four workflows Jev averages 67.8% accuracy at $0.0004 and 0.4s per instance; the opus 5 workflow baseline is 73.1% / $0.1761 / 37.8s and sol is 74.1% / $0.0836 / 23.3s. Accuracy lands in the sonnet 5 / terra tier while cost sits almost two orders of magnitude away from everyone. The homepage's "193.6x faster, 444.6x cheaper" comes from this eval, and TypeSafe concedes it is the optimistic end of the real gain.
The bill is its hardest edge: $0.042 per million input tokens ($42 per billion), output free - 238x cheaper than Claude Fable 5.1's input price. The Doom demo at 10 queries per second costs about $7 per hour. Limits are 250k tokens/s and 1200 requests/min, explicitly dynamic during the early stage. Caveats matter too: on an open parallel constrained-decoding model (same calibrated-probability mechanism) we ran 30 gold-labeled decision cases and saw mean confidence 0.82 against actual accuracy 0.70 - calibrated confidences from this family cannot be consumed as accuracies.