Model · hosted decision model
Jev accuracy on multi-hop reasoning
Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.
Jev accuracy by proof depth
Each point is about 300 problems that need that many inference steps.
- Jev
- Jev
Open world: true, false or unknown
| Proof depth | Items | Jev | Best constant |
|---|---|---|---|
| depth 0 | 300 | 97.7% 96.0–99.3 | 33.3% |
| depth 1 | 302 | 87.7% 83.8–91.4 | 33.4% |
| depth 2 | 303 | 84.8% 80.2–88.4 | 33.3% |
| depth 3 | 303 | 79.9% 75.2–84.5 | 33.3% |
| depth 4 | 303 | 71.9% 66.7–77.2 | 33.3% |
| depth 5 | 289 | 81.0% 76.5–85.5 | 34.9% |
Closed world: true or false
| Proof depth | Items | Jev | Best constant |
|---|---|---|---|
| depth 0 | 300 | 99.3% 98.3–100.0 | 50.0% |
| depth 1 | 300 | 93.7% 91.0–96.3 | 50.0% |
| depth 2 | 300 | 92.0% 88.7–94.7 | 50.0% |
| depth 3 | 300 | 85.0% 81.3–89.3 | 50.0% |
| depth 4 | 300 | 76.3% 71.7–81.0 | 50.0% |
| depth 5 | 300 | 89.3% 85.7–92.3 | 50.0% |
True, false and unknown: where Jev errs
Accuracy alone hides which answers a model gets wrong. Rows are the correct answer; columns are what Jev said.
Open world: true, false or unknown
| Correct answer | Said true | Said false | Said unknown | Recall |
|---|---|---|---|---|
| true | 537 | 5 | 63 | 88.8% |
| false | 10 | 534 | 61 | 88.3% |
| unknown | 124 | 28 | 438 | 74.2% |
Jev answered false 31.5%, true 37.3%, unknown 31.2% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.
Closed world: true or false
| Correct answer | Said true | Said false | Recall |
|---|---|---|---|
| true | 807 | 93 | 89.7% |
| false | 100 | 800 | 88.9% |
Jev answered false 49.6%, true 50.4% of the time; the correct answers are false 50.0%, true 50.0%.
Jev across every breakdown
For each property of the problems, Jev's weakest and strongest value on the open-world task (slices of at least 50 items). Only depth was controlled; the rest are descriptive.
| Breakdown | Weakest · strongest |
|---|---|
| Proof depth | depth 4 71.9% · depth 0 97.7% (25.7 points) |
| Gold answer | unknown 74.2% · true 88.8% (14.5 points) |
| Theory kind | attribute 82.8% · relation 85.3% (2.5 points) |
| Negation in the theory | without negation 80.3% · with negation 87.5% (7.2 points) |
| Negated statement | negated statement 78.6% · plain statement 89.1% (10.5 points) |
| Question strategy | inv-rconc 50.2% · random 98.0% (47.8 points) |
| Paraphrased rules | paraphrased (NatLang) 78.1% · templated 86.1% (8.0 points) |
| Theory's deepest proof | 4 67.9% · 0 90.1% (22.2 points) |
| Theory length | 80-109 79.8% · 0-49 93.1% (13.3 points) |
| Number of rules | 8+ 80.0% · 0-2 92.2% (12.2 points) |
| Number of facts | 8-12 80.3% · 0-3 90.4% (10.1 points) |
| Proof size | 0-1 80.4% · 2 98.1% (17.8 points) |
Jev against every other model
Paired differences: both models answered exactly the same items, so the gap has its own 95% interval. Positive numbers mean Jev was more accurate. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.
Jev minus GPT-6 Luna, open world: +19.8 points overall (+17.2 to +22.4)
Jev minus Kev-9B, open world: +25.3 points overall (+22.7 to +27.8)
Jev minus Kev-4B, open world: +30.3 points overall (+27.4 to +32.9)
Jev minus Kev-0.8B, open world: +30.3 points overall (+27.7 to +32.7)
Jev minus Laya, open world: +41.4 points overall (+38.8 to +44.0)
Jev minus GPT-6 Luna, closed world: +24.3 points overall (+21.9 to +26.6)
Jev minus Kev-9B, closed world: +25.5 points overall (+23.0 to +27.8)
Jev minus Kev-4B, closed world: +30.5 points overall (+27.9 to +33.1)
Jev minus Kev-0.8B, closed world: +33.5 points overall (+30.7 to +36.3)
Jev minus Laya, closed world: +33.7 points overall (+30.7 to +36.5)
Each rival in full: Jev vs GPT-6 Luna · Kev vs Jev · Jev vs Laya
Consistency and calibration
Asked every open world question a second time, Jev gave the same answer on 97.2% of 1,800 items (50 changed; Gwet's AC1 0.958, interval 0.948 to 0.969). Accuracy was 83.8% the first time and 84.2% the second.
Asked every closed world question a second time, Jev gave the same answer on 97.9% of 1,800 items (38 changed; Gwet's AC1 0.958, interval 0.944 to 0.971). Accuracy was 89.3% the first time and 89.7% the second.
Calibration on the open world task: Jev's probability for its own answer averaged 87.8% against 83.8% accuracy; expected calibration error 4.1 points.
Calibration on the closed world task: Jev's probability for its own answer averaged 92.1% against 89.3% accuracy; expected calibration error 2.9 points.
Size, hardware and speed
- Parameters
- undisclosed
- Architecture
- undisclosed
- Weights
- not available (hosted)
- Where it ran
- TypeSafe API (typesafe-sdk); version recorded per row (jev-1.13.0)
- Maker's stated hardware
- not applicable (hosted)
- Price
- $42 per billion input tokens; output free
- Time per decision, open world
- median 189 ms, 90th percentile 241 ms, over the network
- Time per decision, closed world
- median 170 ms, 90th percentile 207 ms, over the network
Hosted: timed over the network against the vendor's servers, so its latency is not like for like with a model on a laptop. Speed, size and memory for every model.
What we predicted about Jev
- 2 of 5 predictions about Jev held (1 partly held, 2 failed). See How we measured.