Explainer · multi-hop reasoning examples

What proof depth means: multi-hop reasoning, step by step, with examples

Proof depth is how many rules must be chained to reach an answer. Here is what that means, from a statement written in the text (depth 0) to one that needs five chained inferences (depth 5), with a real ProofWriter problem and its proof at every depth.

Facts, rules and new facts

Every problem in this benchmark is a short text and a statement to judge. The text has two kinds of sentence.

A fact
Something stated as true about someone: "Erin is furry."
A rule
An "if this, then that": "If Erin is furry then Erin is not rough." Rules can need more than one condition: "Kind, young people are blue."
Applying a rule
When every condition of a rule is a fact, its conclusion becomes a new fact: from "Erin is furry" and that rule, "Erin is not rough".
A chain
A new fact can meet the condition of another rule, which gives another new fact, and so on.

What proof depth counts

Proof depth is how many rule applications the longest chain in the proof needs. At depth 0 the statement is written in the text. At depth 1 one rule gets you there. At depth 5 five rules must be applied one after another, each feeding the next.

Each extra step is another chance to go wrong: to miss a condition, apply a rule that does not fit, or lose track of a fact derived two steps earlier. A model that answers in one pass, without writing the steps down, has to do all of that internally. That is why accuracy falls with depth on the proof-depth breakdown, and why depth is the difficulty axis the benchmark controls: about 300 problems at each depth on each task.

Depth is not the length of the text. The depth 4 and depth 5 examples below share one text of 10 facts and 5 rules; only the statement changes. Longer texts mostly add distractors: facts and rules the proof never uses.

Open world and closed world: where false and unknown come from

The examples below are all true: their statement can be proved. The other answers depend on which rule the task uses for what cannot be proved.

Open world: true, false or unknown
A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does. A false needs a proof too, of the statement's negation; unknown is what is left when neither side can be proved.
Closed world: true or false
Anything that cannot be derived is false, so there is no unknown answer. "Not" in a rule's condition holds when the unnegated fact cannot be derived, so answering a closed-world question can mean checking that something cannot be proved.

How these examples were chosen

These are the shortest problems at each depth, so they are easier than typical problems at their depth. 3 of the 6 open-world examples were answered correctly by every model. They show what depth means; they do not show the accuracy gap. For that, read accuracy by proof depth, where each depth is about 300 problems.

The selection is fixed and blind to the answers. From the preregistered sample: Selection is fixed and blind to the answers: for each task and depth 0-5, among sampled items that are templated (not paraphrased), attribute theories, gold answer true (so a proof exists) and question strategy proof, take the one with the fewest words of theory, ties broken by id. The proof comes from the dataset's own proofsWithIntermediates (the first proof listed), parsed into steps in the order they're derived.

Depth 0: the answer is written down

The text (1 facts, 1 rules, 15 words)

  • fact Gary is not white.
  • rule If something is cold and big then it is not young.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Gary is not white.

Proof, no steps

"Gary is not white." is one of the facts.

Answer: true.

The statement is one of the facts, word for word, so no rule is needed: depth 0. One rule in the text is not needed: a distractor.

Percentages are each model's own stated confidence in its answer to this one problem, not its accuracy. GPT-6 Luna's scored run returned no probability, so its figure comes from the log-probability probe, which asked the same problem again and gave the same answer.

Across all 300 open-world problems at depth 0, not just this short one: Jev 97.7%, GLiDE 97.7%, GPT-6 Luna 93.3%, Kev-9B 89.3%, Kev-4B 87.7%, Kev-0.8B 70.3%, Laya 62.7%.

Depth 1: one rule

The text (3 facts, 1 rules, 19 words)

  • fact Erin is furry.
  • fact Erin is not kind.
  • fact Erin is round.
  • rule If Erin is furry then Erin is not rough.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Erin is not rough.

Proof, 1 step

  1. Erin is furry with If Erin is furry then Erin is not rough. gives Erin is not rough.

Answer: true.

  1. Erin is furry.If Erin is furry then Erin is not rough.Erin is not rough.

The answer takes one rule application, so it is depth 1. Two facts in the text are not needed: distractors.

Across all 302 open-world problems at depth 1, not just this short one: Jev 87.7%, GLiDE 87.7%, GPT-6 Luna 77.8%, Kev-9B 73.8%, Kev-4B 62.9%, Kev-0.8B 54.0%, Laya 42.1%.

Depth 2: two rules in a chain

The text (3 facts, 3 rules, 29 words)

  • fact Anne is rough.
  • fact Bob is not big.
  • fact Harry is big.
  • rule Nice, quiet people are not big.
  • rule If someone is big then they are quiet.
  • rule Quiet people are not nice.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Harry is not nice.

Proof, 2 steps

  1. Harry is big with If someone is big then they are quiet. gives Harry is quiet.
  2. Harry is quiet with Quiet people are not nice. gives Harry is not nice.

Answer: true.

  1. Harry is big.If someone is big then they are quiet.Harry is quiet.
  2. Quiet people are not nice.Harry is not nice.

The answer takes two rule applications in a chain: each new fact ("Harry is quiet") feeds the next rule, so it is depth 2. Two facts and one rule in the text are not needed: distractors.

Across all 303 open-world problems at depth 2, not just this short one: Jev 84.8%, GLiDE 84.2%, GPT-6 Luna 60.4%, Kev-9B 58.4%, Kev-4B 52.1%, Kev-0.8B 46.9%, Laya 41.3%.

Depth 3: three rules in a chain

The text (4 facts, 3 rules, 33 words)

  • fact Bob is not young.
  • fact Dave is not green.
  • fact Erin is nice.
  • fact Gary is white.
  • rule Round things are young.
  • rule If something is white and young then it is not green.
  • rule White things are round.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Gary is not green.

Proof, 3 steps

  1. Gary is white with White things are round. gives Gary is round.
  2. Gary is round with Round things are young. gives Gary is young.
  3. Gary is white and Gary is young with If something is white and young then it is not green. gives Gary is not green.

Answer: true.

  1. Gary is white.White things are round.Gary is round.
  2. Round things are young.Gary is young.
  3. + Gary is white.If something is white and young then it is not green.Gary is not green.

The answer takes three rule applications in a chain: each new fact ("Gary is round", "Gary is young") feeds the next rule, so it is depth 3. One step needs two premises at once, which the model must keep track of. Three facts in the text are not needed: distractors.

Across all 303 open-world problems at depth 3, not just this short one: Jev 79.9%, GLiDE 78.5%, GPT-6 Luna 61.7%, Kev-9B 49.5%, Kev-4B 44.6%, Kev-0.8B 48.8%, Laya 38.3%.

Depth 4: four rules in a chain

The text (10 facts, 5 rules, 57 words)

  • fact Anne is blue.
  • fact Anne is kind.
  • fact Anne is smart.
  • fact Anne is young.
  • fact Charlie is blue.
  • fact Charlie is green.
  • fact Charlie is kind.
  • fact Charlie is red.
  • fact Dave is quiet.
  • fact Harry is kind.
  • rule All kind, quiet people are young.
  • rule All kind people are quiet.
  • rule Kind, young people are blue.
  • rule All young, blue people are green.
  • rule Blue, green people are smart.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Harry is green.

Proof, 4 steps

  1. Harry is kind with All kind people are quiet. gives Harry is quiet.
  2. Harry is kind and Harry is quiet with All kind, quiet people are young. gives Harry is young.
  3. Harry is kind and Harry is young with Kind, young people are blue. gives Harry is blue.
  4. Harry is young and Harry is blue with All young, blue people are green. gives Harry is green.

Answer: true.

  1. Harry is kind.All kind people are quiet.Harry is quiet.
  2. + Harry is kind.All kind, quiet people are young.Harry is young.
  3. + Harry is kind.Kind, young people are blue.Harry is blue.
  4. All young, blue people are green.Harry is green.

The answer takes four rule applications in a chain: each new fact ("Harry is quiet", "Harry is young", "Harry is blue") feeds the next rule, so it is depth 4. Three steps need two premises at once, which the model must keep track of. Nine facts and one rule in the text are not needed: distractors.

Across all 303 open-world problems at depth 4, not just this short one: Jev 71.9%, GLiDE 73.9%, GPT-6 Luna 44.6%, Kev-9B 38.6%, Kev-4B 37.3%, Kev-0.8B 50.5%, Laya 36.0%.

Depth 5: five rules in a chain

The text (10 facts, 5 rules, 57 words)

  • fact Anne is blue.
  • fact Anne is kind.
  • fact Anne is smart.
  • fact Anne is young.
  • fact Charlie is blue.
  • fact Charlie is green.
  • fact Charlie is kind.
  • fact Charlie is red.
  • fact Dave is quiet.
  • fact Harry is kind.
  • rule All kind, quiet people are young.
  • rule All kind people are quiet.
  • rule Kind, young people are blue.
  • rule All young, blue people are green.
  • rule Blue, green people are smart.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Harry is smart.

Proof, 5 steps

  1. Harry is kind with All kind people are quiet. gives Harry is quiet.
  2. Harry is kind and Harry is quiet with All kind, quiet people are young. gives Harry is young.
  3. Harry is kind and Harry is young with Kind, young people are blue. gives Harry is blue.
  4. Harry is young and Harry is blue with All young, blue people are green. gives Harry is green.
  5. Harry is blue and Harry is green with Blue, green people are smart. gives Harry is smart.

Answer: true.

  1. Harry is kind.All kind people are quiet.Harry is quiet.
  2. + Harry is kind.All kind, quiet people are young.Harry is young.
  3. + Harry is kind.Kind, young people are blue.Harry is blue.
  4. All young, blue people are green.Harry is green.
  5. Blue, green people are smart.Harry is smart.

The answer takes five rule applications in a chain: each new fact ("Harry is quiet", "Harry is young", "Harry is blue", "Harry is green") feeds the next rule, so it is depth 5. Four steps need two premises at once, which the model must keep track of. Nine facts in the text are not needed: distractors. This is the same text as the depth 4 example, with 10 facts and 5 rules either way; only the statement changes, and it sits one step further down the same chain. Depth is the length of the chain, not the number of facts or rules.

Across all 289 open-world problems at depth 5, not just this short one: Jev 81.0%, GLiDE 75.1%, GPT-6 Luna 46.0%, Kev-9B 41.2%, Kev-4B 36.3%, Kev-0.8B 50.9%, Laya 34.3%.

The same idea under the closed world

Chosen by the same rule from the closed-world sample. The proofs work the same way; what changes is that "not" in a condition holds when something cannot be derived, and anything that cannot be derived is false. Depths 4 and 5 again share one text.

Six closed-world examples, depth 0 to 5

Depth 0: the answer is written down (Gary is blue.)

The text (3 facts, 3 rules, 30 words)

  • fact Gary is blue.
  • fact Gary is red.
  • fact Harry is young.
  • rule All young, kind people are blue.
  • rule Round people are kind.
  • rule If someone is young and not red then they are round.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Gary is blue.

Proof, no steps

"Gary is blue." is one of the facts.

Answer: true.

The statement is one of the facts, word for word, so no rule is needed: depth 0. Two facts and three rules in the text are not needed: distractors.

Percentages are each model's own stated confidence in its answer to this one problem, not its accuracy. GPT-6 Luna's scored run returned no probability, so its figure comes from the log-probability probe, which asked the same problem again and gave the same answer.

Depth 1: one rule (Anne is quiet.)

The text (10 facts, 1 rules, 35 words)

  • fact Anne is big.
  • fact Anne is cold.
  • fact Anne is nice.
  • fact Anne is rough.
  • fact Anne is round.
  • fact Bob is rough.
  • fact Charlie is nice.
  • fact Charlie is quiet.
  • fact Charlie is rough.
  • fact Charlie is round.
  • rule All round people are quiet.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Anne is quiet.

Proof, 1 step

  1. Anne is round with All round people are quiet. gives Anne is quiet.

Answer: true.

  1. Anne is round.All round people are quiet.Anne is quiet.

The answer takes one rule application, so it is depth 1. Nine facts in the text are not needed: distractors.

Depth 2: two rules in a chain (Erin is nice.)

The text (3 facts, 2 rules, 19 words)

  • fact Erin is blue.
  • fact Erin is cold.
  • fact Erin is young.
  • rule Young, blue people are big.
  • rule Big, young people are nice.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Erin is nice.

Proof, 2 steps

  1. Erin is young and Erin is blue with Young, blue people are big. gives Erin is big.
  2. Erin is big and Erin is young with Big, young people are nice. gives Erin is nice.

Answer: true.

  1. Erin is young. + Erin is blue.Young, blue people are big.Erin is big.
  2. + Erin is young.Big, young people are nice.Erin is nice.

The answer takes two rule applications in a chain: each new fact ("Erin is big") feeds the next rule, so it is depth 2. Every step needs two premises at once, which the model must keep track of. One fact in the text is not needed: a distractor.

Depth 3: three rules in a chain (Gary is kind.)

The text (2 facts, 3 rules, 23 words)

  • fact Erin is red.
  • fact Gary is red.
  • rule If Gary is quiet then Gary is kind.
  • rule Green things are quiet.
  • rule All red things are green.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Gary is kind.

Proof, 3 steps

  1. Gary is red with All red things are green. gives Gary is green.
  2. Gary is green with Green things are quiet. gives Gary is quiet.
  3. Gary is quiet with If Gary is quiet then Gary is kind. gives Gary is kind.

Answer: true.

  1. Gary is red.All red things are green.Gary is green.
  2. Green things are quiet.Gary is quiet.
  3. If Gary is quiet then Gary is kind.Gary is kind.

The answer takes three rule applications in a chain: each new fact ("Gary is green", "Gary is quiet") feeds the next rule, so it is depth 3. One fact in the text is not needed: a distractor.

Depth 4: four rules in a chain (Harry is green.)

The text (10 facts, 5 rules, 57 words)

  • fact Anne is blue.
  • fact Anne is kind.
  • fact Anne is smart.
  • fact Anne is young.
  • fact Charlie is blue.
  • fact Charlie is green.
  • fact Charlie is kind.
  • fact Charlie is red.
  • fact Dave is quiet.
  • fact Harry is kind.
  • rule All kind, quiet people are young.
  • rule All kind people are quiet.
  • rule Kind, young people are blue.
  • rule All young, blue people are green.
  • rule Blue, green people are smart.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Harry is green.

Proof, 4 steps

  1. Harry is kind with All kind people are quiet. gives Harry is quiet.
  2. Harry is kind and Harry is quiet with All kind, quiet people are young. gives Harry is young.
  3. Harry is kind and Harry is young with Kind, young people are blue. gives Harry is blue.
  4. Harry is young and Harry is blue with All young, blue people are green. gives Harry is green.

Answer: true.

  1. Harry is kind.All kind people are quiet.Harry is quiet.
  2. + Harry is kind.All kind, quiet people are young.Harry is young.
  3. + Harry is kind.Kind, young people are blue.Harry is blue.
  4. All young, blue people are green.Harry is green.

The answer takes four rule applications in a chain: each new fact ("Harry is quiet", "Harry is young", "Harry is blue") feeds the next rule, so it is depth 4. Three steps need two premises at once, which the model must keep track of. Nine facts and one rule in the text are not needed: distractors.

Depth 5: five rules in a chain (Harry is smart.)

The text (10 facts, 5 rules, 57 words)

  • fact Anne is blue.
  • fact Anne is kind.
  • fact Anne is smart.
  • fact Anne is young.
  • fact Charlie is blue.
  • fact Charlie is green.
  • fact Charlie is kind.
  • fact Charlie is red.
  • fact Dave is quiet.
  • fact Harry is kind.
  • rule All kind, quiet people are young.
  • rule All kind people are quiet.
  • rule Kind, young people are blue.
  • rule All young, blue people are green.
  • rule Blue, green people are smart.

Highlighted lines are the ones the proof uses; the muted ones are distractors.

Statement

Harry is smart.

Proof, 5 steps

  1. Harry is kind with All kind people are quiet. gives Harry is quiet.
  2. Harry is kind and Harry is quiet with All kind, quiet people are young. gives Harry is young.
  3. Harry is kind and Harry is young with Kind, young people are blue. gives Harry is blue.
  4. Harry is young and Harry is blue with All young, blue people are green. gives Harry is green.
  5. Harry is blue and Harry is green with Blue, green people are smart. gives Harry is smart.

Answer: true.

  1. Harry is kind.All kind people are quiet.Harry is quiet.
  2. + Harry is kind.All kind, quiet people are young.Harry is young.
  3. + Harry is kind.Kind, young people are blue.Harry is blue.
  4. All young, blue people are green.Harry is green.
  5. Blue, green people are smart.Harry is smart.

The answer takes five rule applications in a chain: each new fact ("Harry is quiet", "Harry is young", "Harry is blue", "Harry is green") feeds the next rule, so it is depth 5. Four steps need two premises at once, which the model must keep track of. Nine facts in the text are not needed: distractors. This is the same text as the depth 4 example, with 10 facts and 5 rules either way; only the statement changes, and it sits one step further down the same chain. Depth is the length of the chain, not the number of facts or rules.

Where to go next

On true-or-false questions, every model we tested except Jev falls to coin-flip accuracy by five chained inferences: the open decision models Kev and Laya, GPT-6 Luna used as a one-shot classifier, and GLiDE. Jev still answers 89.3% correctly at that depth.