Breakdown · descriptive

Paraphrased rules vs templated rules

Whether the theory comes from ProofWriter's NatLang set, where people reworded the templated sentences, or from the templated sets. The preregistration predicted paraphrased rules would be no easier.

8.0On the open-world task, Jev's accuracy ranges from 78.1% (paraphrased (NatLang)) to 86.1% (templated), a spread of 8.0 points. Jev is the most accurate model overall; every model is below.

Descriptive, not causal: the sample controls proof depth and the answer, not paraphrased rules. Across these slices the average proof depth runs from 2.0 to 2.7, so depth differs between slices and can add to or mask any gap. The answer mix and mean depth of every slice are in the tables.

Paraphrased rules, open world task

A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (33%); the grey tick is the score of always giving the slice's most common answer.

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
templated1,298 items
paraphrased (NatLang)502 items
Paraphrased rules, Open world
Paraphrased rulesItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
templated1,29886.1% 84.2–87.967.6% 64.8–70.259.6% 57.0–62.255.9% 53.2–58.653.8% 51.2–56.441.7% 39.2–44.334.5%false 33% · true 33% · unknown 35% · 2.7
paraphrased (NatLang)50278.1% 74.5–82.154.8% 50.4–59.455.8% 51.8–60.047.6% 43.4–52.053.0% 48.8–57.244.4% 40.2–48.835.9%false 36% · true 36% · unknown 28% · 2.0

Paraphrased rules, closed world task

Anything that cannot be derived is false, so there is no unknown answer. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (50%); the grey tick is the score of always giving the slice's most common answer.

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
templated1,306 items
paraphrased (NatLang)494 items
Paraphrased rules, Closed world
Paraphrased rulesItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
templated1,30694.0% 92.8–95.367.3% 64.8–69.964.8% 62.0–67.560.2% 57.7–62.756.6% 53.8–59.356.1% 53.4–58.950.2%false 50% · true 50% · 2.7
paraphrased (NatLang)49476.7% 73.1–80.458.9% 54.7–63.461.1% 56.7–65.455.1% 50.4–59.553.6% 49.6–58.154.3% 49.8–58.950.6%false 49% · true 51% · 2.0