Breakdown · the controlled difficulty axis
Accuracy by proof depth: where multi-hop reasoning breaks
How many rule applications the shortest proof of the answer needs. Depth 0 is a fact stated in the text; depth 5 needs five chained inferences. This is the benchmark's difficulty axis, and the sample was built to hold it fixed: about 300 items at each depth.
As proofs get deeper, every model we tested except Jev falls to coin-flip accuracy by five chained inferences: the open decision models Kev and Laya, and GPT-6 Luna used as a one-shot classifier. Jev still answers 81.0–89.3% correctly at that depth.
25.7On the open-world task, Jev's accuracy ranges from 71.9% (depth 4) to 97.7% (depth 0), a spread of 25.7 points. Jev is the most accurate model overall; every model is below.
Proof depth, open world task
A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (33%); the grey tick is the score of always giving the slice's most common answer.
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Proof depth | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant |
|---|---|---|---|---|---|---|---|---|
| depth 0 | 300 | 97.7% 96.0–99.3 | 93.3% 90.3–96.0 | 89.3% 85.7–92.7 | 87.7% 84.0–91.3 | 70.3% 65.3–75.7 | 62.7% 57.0–68.3 | 33.3% |
| depth 1 | 302 | 87.7% 83.8–91.4 | 77.8% 73.2–82.8 | 73.8% 68.9–78.8 | 62.9% 57.3–68.2 | 54.0% 48.3–59.6 | 42.1% 36.8–47.7 | 33.4% |
| depth 2 | 303 | 84.8% 80.2–88.4 | 60.4% 54.5–66.0 | 58.4% 53.1–63.7 | 52.1% 46.5–57.8 | 46.9% 41.3–52.8 | 41.3% 36.0–46.9 | 33.3% |
| depth 3 | 303 | 79.9% 75.2–84.5 | 61.7% 56.8–67.7 | 49.5% 44.2–55.4 | 44.6% 39.3–49.8 | 48.8% 43.6–54.5 | 38.3% 33.0–44.6 | 33.3% |
| depth 4 | 303 | 71.9% 66.7–77.2 | 44.6% 38.6–49.5 | 38.6% 33.3–43.6 | 37.3% 31.7–42.2 | 50.5% 44.9–56.1 | 36.0% 31.0–41.6 | 33.3% |
| depth 5 | 289 | 81.0% 76.5–85.5 | 46.0% 40.1–51.6 | 41.2% 35.3–46.7 | 36.3% 31.1–41.9 | 50.9% 45.3–56.7 | 34.3% 29.1–39.8 | 34.9% |
Proof depth, closed world task
Anything that cannot be derived is false, so there is no unknown answer. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (50%); the grey tick is the score of always giving the slice's most common answer.
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Proof depth | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant |
|---|---|---|---|---|---|---|---|---|
| depth 0 | 300 | 99.3% 98.3–100.0 | 98.0% 96.3–99.3 | 89.7% 86.0–93.0 | 85.7% 81.7–89.7 | 68.0% 62.7–73.3 | 72.0% 66.7–76.7 | 50.0% |
| depth 1 | 300 | 93.7% 91.0–96.3 | 76.7% 72.3–81.0 | 77.0% 72.7–81.7 | 68.3% 63.0–73.7 | 59.7% 54.0–65.7 | 51.0% 45.0–56.0 | 50.0% |
| depth 2 | 300 | 92.0% 88.7–94.7 | 65.3% 60.0–70.7 | 64.0% 59.0–69.3 | 54.3% 49.0–59.7 | 51.3% 45.7–57.0 | 54.7% 48.7–59.7 | 50.0% |
| depth 3 | 300 | 85.0% 81.3–89.3 | 57.0% 51.7–63.0 | 58.7% 53.0–64.3 | 56.0% 50.7–61.3 | 53.0% 47.7–58.7 | 51.3% 45.7–56.7 | 50.0% |
| depth 4 | 300 | 76.3% 71.7–81.0 | 48.3% 43.0–53.7 | 50.3% 44.7–55.7 | 45.0% 39.0–50.3 | 51.0% 45.7–57.0 | 58.7% 53.0–64.3 | 50.0% |
| depth 5 | 300 | 89.3% 85.7–92.3 | 44.7% 38.7–50.3 | 43.0% 37.3–48.7 | 43.3% 38.3–49.0 | 51.7% 46.0–57.3 | 46.0% 40.3–51.0 | 50.0% |
What a deeper proof looks like
The shortest templated item in the open-world sample at each depth whose answer is true. Each needs one more rule application than the last.
Gary is not white. If something is cold and big then it is not young.
Statement: Gary is not white.
The lion is nice. If something is nice then it is not blue.
Statement: The lion is not blue.
The lion is kind. Blue things are red. All kind things are blue.
Statement: The lion is red.
The squirrel is kind. All kind people are big. Green, big people are red. Big, kind people are green.
Statement: The squirrel is red.
Bob is big. Bob is cold. Bob is kind. Bob is round. Bob is smart. Dave is cold. Erin is big. Erin is green. Fiona is big. Fiona is smart. Big, green things are round. If something is cold and blue then it is smart. Smart, round things are kind. Round, big things are cold. Cold things are blue.
Statement: Erin is smart.
Bob is big. Bob is cold. Bob is kind. Bob is round. Bob is smart. Dave is cold. Erin is big. Erin is green. Fiona is big. Fiona is smart. Big, green things are round. If something is cold and blue then it is smart. Smart, round things are kind. Round, big things are cold. Cold things are blue.
Statement: Erin is kind.
The gap to Jev grows with depth
Paired differences on the same items: Jev minus each model, by depth. Positive means Jev was more accurate.
Jev minus GPT-6 Luna, open world
Jev minus GPT-6 Luna, closed world
Jev minus Kev-9B, open world
Jev minus Kev-9B, closed world
Jev minus Kev-4B, open world
Jev minus Kev-4B, closed world
Jev minus Kev-0.8B, open world
Jev minus Kev-0.8B, closed world
Jev minus Laya, open world
Jev minus Laya, closed world
The unknown answer at each depth
Recall on each correct answer, by depth, on the open-world task: the share of items with that answer each model got right. Counted from the saved answers of every fully scored model.
| Model | Correct answer | depth 0 | depth 1 | depth 2 | depth 3 | depth 4 | depth 5 |
|---|---|---|---|---|---|---|---|
| Jev | true | 98% | 89% | 86% | 86% | 80% | 93% |
| false | 99% | 92% | 87% | 80% | 79% | 92% | |
| unknown | 96% | 82% | 81% | 73% | 56% | 54% | |
| GPT-6 Luna | true | 96% | 77% | 56% | 52% | 33% | 36% |
| false | 85% | 61% | 46% | 52% | 27% | 35% | |
| unknown | 99% | 95% | 79% | 80% | 74% | 71% | |
| Kev-9B | true | 97% | 78% | 61% | 53% | 30% | 36% |
| false | 94% | 73% | 44% | 40% | 19% | 23% | |
| unknown | 77% | 70% | 70% | 55% | 67% | 69% | |
| Kev-4B | true | 96% | 64% | 44% | 35% | 20% | 22% |
| false | 82% | 48% | 32% | 27% | 16% | 12% | |
| unknown | 85% | 77% | 81% | 72% | 76% | 82% | |
| Kev-0.8B | true | 91% | 61% | 48% | 48% | 50% | 52% |
| false | 92% | 88% | 87% | 89% | 92% | 89% | |
| unknown | 28% | 12% | 6% | 10% | 10% | 5% | |
| Laya | true | 89% | 37% | 38% | 33% | 29% | 17% |
| false | 97% | 88% | 84% | 81% | 79% | 81% | |
| unknown | 2% | 1% | 2% | 1% | 0% | 0% |
About 100 items per cell, so each number is uncertain by roughly ±10 points.