Model · hosted decision model

Jev accuracy on multi-hop reasoning

Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.

Open world83.8%82.2 to 85.5 (95%) on 1,800 problems; place 1 of 6
Closed world89.3%87.8 to 90.7 (95%) on 1,800 problems; place 1 of 6
Depth 0 to depth 5−16.7points on the open-world task, 97.7% to 81.0%

Jev accuracy by proof depth

Each point is about 300 problems that need that many inference steps.

Open world: true, false or unknown
  • Jev
0%25%50%75%100%012345proof depth (inference steps)chance 33%
Closed world: true or false
  • Jev
0%25%50%75%100%012345proof depth (inference steps)chance 50%

Open world: true, false or unknown

Jev by depth, Open world
Proof depthItemsJevBest constant
depth 030097.7% 96.0–99.333.3%
depth 130287.7% 83.8–91.433.4%
depth 230384.8% 80.2–88.433.3%
depth 330379.9% 75.2–84.533.3%
depth 430371.9% 66.7–77.233.3%
depth 528981.0% 76.5–85.534.9%

Closed world: true or false

Jev by depth, Closed world
Proof depthItemsJevBest constant
depth 030099.3% 98.3–100.050.0%
depth 130093.7% 91.0–96.350.0%
depth 230092.0% 88.7–94.750.0%
depth 330085.0% 81.3–89.350.0%
depth 430076.3% 71.7–81.050.0%
depth 530089.3% 85.7–92.350.0%

True, false and unknown: where Jev errs

Accuracy alone hides which answers a model gets wrong. Rows are the correct answer; columns are what Jev said.

Open world: true, false or unknown

Correct answer by what the model said
Correct answerSaid trueSaid falseSaid unknownRecall
true53756388.8%
false105346188.3%
unknown1242843874.2%

Jev answered false 31.5%, true 37.3%, unknown 31.2% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.

Closed world: true or false

Correct answer by what the model said
Correct answerSaid trueSaid falseRecall
true8079389.7%
false10080088.9%

Jev answered false 49.6%, true 50.4% of the time; the correct answers are false 50.0%, true 50.0%.

Jev across every breakdown

For each property of the problems, Jev's weakest and strongest value on the open-world task (slices of at least 50 items). Only depth was controlled; the rest are descriptive.

BreakdownWeakest · strongest
Proof depthdepth 4 71.9% · depth 0 97.7% (25.7 points)
Gold answerunknown 74.2% · true 88.8% (14.5 points)
Theory kindattribute 82.8% · relation 85.3% (2.5 points)
Negation in the theorywithout negation 80.3% · with negation 87.5% (7.2 points)
Negated statementnegated statement 78.6% · plain statement 89.1% (10.5 points)
Question strategyinv-rconc 50.2% · random 98.0% (47.8 points)
Paraphrased rulesparaphrased (NatLang) 78.1% · templated 86.1% (8.0 points)
Theory's deepest proof4 67.9% · 0 90.1% (22.2 points)
Theory length80-109 79.8% · 0-49 93.1% (13.3 points)
Number of rules8+ 80.0% · 0-2 92.2% (12.2 points)
Number of facts8-12 80.3% · 0-3 90.4% (10.1 points)
Proof size0-1 80.4% · 2 98.1% (17.8 points)

Jev against every other model

Paired differences: both models answered exactly the same items, so the gap has its own 95% interval. Positive numbers mean Jev was more accurate. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.

Jev minus GPT-6 Luna, open world: +19.8 points overall (+17.2 to +22.4)

depth 0300 items
+4.3
depth 1302 items
+9.9
depth 2303 items
+24.4
depth 3303 items
+18.2
depth 4303 items
+27.4
depth 5289 items
+34.9

Jev minus Kev-9B, open world: +25.3 points overall (+22.7 to +27.8)

depth 0300 items
+8.3
depth 1302 items
+13.9
depth 2303 items
+26.4
depth 3303 items
+30.4
depth 4303 items
+33.3
depth 5289 items
+39.8

Jev minus Kev-4B, open world: +30.3 points overall (+27.4 to +32.9)

depth 0300 items
+10.0
depth 1302 items
+24.8
depth 2303 items
+32.7
depth 3303 items
+35.3
depth 4303 items
+34.7
depth 5289 items
+44.6

Jev minus Kev-0.8B, open world: +30.3 points overall (+27.7 to +32.7)

depth 0300 items
+27.3
depth 1302 items
+33.8
depth 2303 items
+38.0
depth 3303 items
+31.0
depth 4303 items
+21.5
depth 5289 items
+30.1

Jev minus Laya, open world: +41.4 points overall (+38.8 to +44.0)

depth 0300 items
+35.0
depth 1302 items
+45.7
depth 2303 items
+43.6
depth 3303 items
+41.6
depth 4303 items
+36.0
depth 5289 items
+46.7

Jev minus GPT-6 Luna, closed world: +24.3 points overall (+21.9 to +26.6)

depth 0300 items
+1.3
depth 1300 items
+17.0
depth 2300 items
+26.7
depth 3300 items
+28.0
depth 4300 items
+28.0
depth 5300 items
+44.7

Jev minus Kev-9B, closed world: +25.5 points overall (+23.0 to +27.8)

depth 0300 items
+9.7
depth 1300 items
+16.7
depth 2300 items
+28.0
depth 3300 items
+26.3
depth 4300 items
+26.0
depth 5300 items
+46.3

Jev minus Kev-4B, closed world: +30.5 points overall (+27.9 to +33.1)

depth 0300 items
+13.7
depth 1300 items
+25.3
depth 2300 items
+37.7
depth 3300 items
+29.0
depth 4300 items
+31.3
depth 5300 items
+46.0

Jev minus Kev-0.8B, closed world: +33.5 points overall (+30.7 to +36.3)

depth 0300 items
+31.3
depth 1300 items
+34.0
depth 2300 items
+40.7
depth 3300 items
+32.0
depth 4300 items
+25.3
depth 5300 items
+37.7

Jev minus Laya, closed world: +33.7 points overall (+30.7 to +36.5)

depth 0300 items
+27.3
depth 1300 items
+42.7
depth 2300 items
+37.3
depth 3300 items
+33.7
depth 4300 items
+17.7
depth 5300 items
+43.3

Consistency and calibration

Asked every open world question a second time, Jev gave the same answer on 97.2% of 1,800 items (50 changed; Gwet's AC1 0.958, interval 0.948 to 0.969). Accuracy was 83.8% the first time and 84.2% the second.

Asked every closed world question a second time, Jev gave the same answer on 97.9% of 1,800 items (38 changed; Gwet's AC1 0.958, interval 0.944 to 0.971). Accuracy was 89.3% the first time and 89.7% the second.

Calibration on the open world task: Jev's probability for its own answer averaged 87.8% against 83.8% accuracy; expected calibration error 4.1 points.

Calibration on the closed world task: Jev's probability for its own answer averaged 92.1% against 89.3% accuracy; expected calibration error 2.9 points.

Size, hardware and speed

Parameters
undisclosed
Architecture
undisclosed
Weights
not available (hosted)
Where it ran
TypeSafe API (typesafe-sdk); version recorded per row (jev-1.13.0)
Maker's stated hardware
not applicable (hosted)
Price
$42 per billion input tokens; output free
Time per decision, open world
median 189 ms, 90th percentile 241 ms, over the network
Time per decision, closed world
median 170 ms, 90th percentile 207 ms, over the network

Hosted: timed over the network against the vendor's servers, so its latency is not like for like with a model on a laptop. Speed, size and memory for every model.

What we predicted about Jev