Breakdown · the controlled difficulty axis

Accuracy by proof depth: where multi-hop reasoning breaks

How many rule applications the shortest proof of the answer needs. Depth 0 is a fact stated in the text; depth 5 needs five chained inferences. This is the benchmark's difficulty axis, and the sample was built to hold it fixed: about 300 items at each depth.

As proofs get deeper, every model we tested except Jev falls to coin-flip accuracy by five chained inferences: the open decision models Kev and Laya, and GPT-6 Luna used as a one-shot classifier. Jev still answers 81.0–89.3% correctly at that depth.

25.7On the open-world task, Jev's accuracy ranges from 71.9% (depth 4) to 97.7% (depth 0), a spread of 25.7 points. Jev is the most accurate model overall; every model is below.

Proof depth, open world task

A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (33%); the grey tick is the score of always giving the slice's most common answer.

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
depth 0300 items
depth 1302 items
depth 2303 items
depth 3303 items
depth 4303 items
depth 5289 items
Proof depth, Open world
Proof depthItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constant
depth 030097.7% 96.0–99.393.3% 90.3–96.089.3% 85.7–92.787.7% 84.0–91.370.3% 65.3–75.762.7% 57.0–68.333.3%
depth 130287.7% 83.8–91.477.8% 73.2–82.873.8% 68.9–78.862.9% 57.3–68.254.0% 48.3–59.642.1% 36.8–47.733.4%
depth 230384.8% 80.2–88.460.4% 54.5–66.058.4% 53.1–63.752.1% 46.5–57.846.9% 41.3–52.841.3% 36.0–46.933.3%
depth 330379.9% 75.2–84.561.7% 56.8–67.749.5% 44.2–55.444.6% 39.3–49.848.8% 43.6–54.538.3% 33.0–44.633.3%
depth 430371.9% 66.7–77.244.6% 38.6–49.538.6% 33.3–43.637.3% 31.7–42.250.5% 44.9–56.136.0% 31.0–41.633.3%
depth 528981.0% 76.5–85.546.0% 40.1–51.641.2% 35.3–46.736.3% 31.1–41.950.9% 45.3–56.734.3% 29.1–39.834.9%

Proof depth, closed world task

Anything that cannot be derived is false, so there is no unknown answer. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (50%); the grey tick is the score of always giving the slice's most common answer.

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
depth 0300 items
depth 1300 items
depth 2300 items
depth 3300 items
depth 4300 items
depth 5300 items
Proof depth, Closed world
Proof depthItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constant
depth 030099.3% 98.3–100.098.0% 96.3–99.389.7% 86.0–93.085.7% 81.7–89.768.0% 62.7–73.372.0% 66.7–76.750.0%
depth 130093.7% 91.0–96.376.7% 72.3–81.077.0% 72.7–81.768.3% 63.0–73.759.7% 54.0–65.751.0% 45.0–56.050.0%
depth 230092.0% 88.7–94.765.3% 60.0–70.764.0% 59.0–69.354.3% 49.0–59.751.3% 45.7–57.054.7% 48.7–59.750.0%
depth 330085.0% 81.3–89.357.0% 51.7–63.058.7% 53.0–64.356.0% 50.7–61.353.0% 47.7–58.751.3% 45.7–56.750.0%
depth 430076.3% 71.7–81.048.3% 43.0–53.750.3% 44.7–55.745.0% 39.0–50.351.0% 45.7–57.058.7% 53.0–64.350.0%
depth 530089.3% 85.7–92.344.7% 38.7–50.343.0% 37.3–48.743.3% 38.3–49.051.7% 46.0–57.346.0% 40.3–51.050.0%

What a deeper proof looks like

The shortest templated item in the open-world sample at each depth whose answer is true. Each needs one more rule application than the last.

Depth 0 · answer: true

Gary is not white. If something is cold and big then it is not young.
Statement: Gary is not white.

Depth 1 · answer: true

The lion is nice. If something is nice then it is not blue.
Statement: The lion is not blue.

Depth 2 · answer: true

The lion is kind. Blue things are red. All kind things are blue.
Statement: The lion is red.

Depth 3 · answer: true

The squirrel is kind. All kind people are big. Green, big people are red. Big, kind people are green.
Statement: The squirrel is red.

Depth 4 · answer: true

Bob is big. Bob is cold. Bob is kind. Bob is round. Bob is smart. Dave is cold. Erin is big. Erin is green. Fiona is big. Fiona is smart. Big, green things are round. If something is cold and blue then it is smart. Smart, round things are kind. Round, big things are cold. Cold things are blue.
Statement: Erin is smart.

Depth 5 · answer: true

Bob is big. Bob is cold. Bob is kind. Bob is round. Bob is smart. Dave is cold. Erin is big. Erin is green. Fiona is big. Fiona is smart. Big, green things are round. If something is cold and blue then it is smart. Smart, round things are kind. Round, big things are cold. Cold things are blue.
Statement: Erin is kind.

The gap to Jev grows with depth

Paired differences on the same items: Jev minus each model, by depth. Positive means Jev was more accurate.

Jev minus GPT-6 Luna, open world

depth 0300 items
+4.3
depth 1302 items
+9.9
depth 2303 items
+24.4
depth 3303 items
+18.2
depth 4303 items
+27.4
depth 5289 items
+34.9

Jev minus GPT-6 Luna, closed world

depth 0300 items
+1.3
depth 1300 items
+17.0
depth 2300 items
+26.7
depth 3300 items
+28.0
depth 4300 items
+28.0
depth 5300 items
+44.7

Jev minus Kev-9B, open world

depth 0300 items
+8.3
depth 1302 items
+13.9
depth 2303 items
+26.4
depth 3303 items
+30.4
depth 4303 items
+33.3
depth 5289 items
+39.8

Jev minus Kev-9B, closed world

depth 0300 items
+9.7
depth 1300 items
+16.7
depth 2300 items
+28.0
depth 3300 items
+26.3
depth 4300 items
+26.0
depth 5300 items
+46.3

Jev minus Kev-4B, open world

depth 0300 items
+10.0
depth 1302 items
+24.8
depth 2303 items
+32.7
depth 3303 items
+35.3
depth 4303 items
+34.7
depth 5289 items
+44.6

Jev minus Kev-4B, closed world

depth 0300 items
+13.7
depth 1300 items
+25.3
depth 2300 items
+37.7
depth 3300 items
+29.0
depth 4300 items
+31.3
depth 5300 items
+46.0

Jev minus Kev-0.8B, open world

depth 0300 items
+27.3
depth 1302 items
+33.8
depth 2303 items
+38.0
depth 3303 items
+31.0
depth 4303 items
+21.5
depth 5289 items
+30.1

Jev minus Kev-0.8B, closed world

depth 0300 items
+31.3
depth 1300 items
+34.0
depth 2300 items
+40.7
depth 3300 items
+32.0
depth 4300 items
+25.3
depth 5300 items
+37.7

Jev minus Laya, open world

depth 0300 items
+35.0
depth 1302 items
+45.7
depth 2303 items
+43.6
depth 3303 items
+41.6
depth 4303 items
+36.0
depth 5289 items
+46.7

Jev minus Laya, closed world

depth 0300 items
+27.3
depth 1300 items
+42.7
depth 2300 items
+37.3
depth 3300 items
+33.7
depth 4300 items
+17.7
depth 5300 items
+43.3

The unknown answer at each depth

Recall on each correct answer, by depth, on the open-world task: the share of items with that answer each model got right. Counted from the saved answers of every fully scored model.

ModelCorrect answerdepth 0depth 1depth 2depth 3depth 4depth 5
Jevtrue98%89%86%86%80%93%
false99%92%87%80%79%92%
unknown96%82%81%73%56%54%
GPT-6 Lunatrue96%77%56%52%33%36%
false85%61%46%52%27%35%
unknown99%95%79%80%74%71%
Kev-9Btrue97%78%61%53%30%36%
false94%73%44%40%19%23%
unknown77%70%70%55%67%69%
Kev-4Btrue96%64%44%35%20%22%
false82%48%32%27%16%12%
unknown85%77%81%72%76%82%
Kev-0.8Btrue91%61%48%48%50%52%
false92%88%87%89%92%89%
unknown28%12%6%10%10%5%
Layatrue89%37%38%33%29%17%
false97%88%84%81%79%81%
unknown2%1%2%1%0%0%

About 100 items per cell, so each number is uncertain by roughly ±10 points.