Breakdown · descriptive
Paraphrased rules vs templated rules
Whether the theory comes from ProofWriter's NatLang set, where people reworded the templated sentences, or from the templated sets. The preregistration predicted paraphrased rules would be no easier.
8.0On the open-world task, Jev's accuracy ranges from 78.1% (paraphrased (NatLang)) to 86.1% (templated), a spread of 8.0 points. Jev is the most accurate model overall; every model is below.
Descriptive, not causal: the sample controls proof depth and the answer, not paraphrased rules. Across these slices the average proof depth runs from 2.0 to 2.7, so depth differs between slices and can add to or mask any gap. The answer mix and mean depth of every slice are in the tables.
Paraphrased rules, open world task
A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (33%); the grey tick is the score of always giving the slice's most common answer.
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Paraphrased rules | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| templated | 1,298 | 86.1% 84.2–87.9 | 67.6% 64.8–70.2 | 59.6% 57.0–62.2 | 55.9% 53.2–58.6 | 53.8% 51.2–56.4 | 41.7% 39.2–44.3 | 34.5% | false 33% · true 33% · unknown 35% · 2.7 |
| paraphrased (NatLang) | 502 | 78.1% 74.5–82.1 | 54.8% 50.4–59.4 | 55.8% 51.8–60.0 | 47.6% 43.4–52.0 | 53.0% 48.8–57.2 | 44.4% 40.2–48.8 | 35.9% | false 36% · true 36% · unknown 28% · 2.0 |
Paraphrased rules, closed world task
Anything that cannot be derived is false, so there is no unknown answer. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (50%); the grey tick is the score of always giving the slice's most common answer.
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Paraphrased rules | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| templated | 1,306 | 94.0% 92.8–95.3 | 67.3% 64.8–69.9 | 64.8% 62.0–67.5 | 60.2% 57.7–62.7 | 56.6% 53.8–59.3 | 56.1% 53.4–58.9 | 50.2% | false 50% · true 50% · 2.7 |
| paraphrased (NatLang) | 494 | 76.7% 73.1–80.4 | 58.9% 54.7–63.4 | 61.1% 56.7–65.4 | 55.1% 50.4–59.5 | 53.6% 49.6–58.1 | 54.3% 49.8–58.9 | 50.6% | false 49% · true 51% · 2.0 |