Breakdowns · ProofWriter results
Where decision models fail: every breakdown
Accuracy broken down by every property ProofWriter records about a problem. Proof depth is the controlled difficulty axis; the others are descriptive.
At the deep end: accuracy at proof depth 5
- Jev76.5 to 85.5 (95%), open world81.0%
- GPT-6 Luna40.1 to 51.6 (95%), open world46.0%
- Kev-9B35.3 to 46.7 (95%), open world41.2%
- Kev-4B31.1 to 41.9 (95%), open world36.3%
- Kev-0.8B45.3 to 56.7 (95%), open world50.9%
- Laya29.1 to 39.8 (95%), open world34.3%
Fully scored models only. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.
Every breakdown, with each model's weakest slice
Open-world task; weakest and strongest slices of at least 50 items.
| Breakdown | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya |
|---|---|---|---|---|---|---|
| Proof depth | depth 4 71.9% · depth 0 97.7% | depth 4 44.6% · depth 0 93.3% | depth 4 38.6% · depth 0 89.3% | depth 5 36.3% · depth 0 87.7% | depth 2 46.9% · depth 0 70.3% | depth 5 34.3% · depth 0 62.7% |
| Gold answer | unknown 74.2% · true 88.8% | false 50.9% · unknown 83.4% | false 48.6% · unknown 68.1% | false 35.9% · unknown 78.8% | unknown 11.9% · false 89.6% | unknown 1.0% · false 85.1% |
| Theory kind | attribute 82.8% · relation 85.3% | attribute 62.5% · relation 66.3% | relation 57.5% · attribute 59.3% | relation 52.4% · attribute 54.4% | relation 53.2% · attribute 53.8% | attribute 41.2% · relation 44.3% |
| Negation in the theory | without negation 80.3% · with negation 87.5% | without negation 61.2% · with negation 67.1% | without negation 56.8% · with negation 60.4% | without negation 49.9% · with negation 57.3% | without negation 50.6% · with negation 56.6% | without negation 38.8% · with negation 46.2% |
| Negated statement | negated statement 78.6% · plain statement 89.1% | plain statement 62.3% · negated statement 65.8% | negated statement 57.4% · plain statement 59.7% | negated statement 53.5% · plain statement 53.6% | plain statement 48.5% · negated statement 58.6% | plain statement 38.7% · negated statement 46.2% |
| Question strategy | inv-rconc 50.2% · random 98.0% | inv-proof 50.9% · inv-random 100.0% | inv-proof 48.6% · random 78.0% | inv-proof 35.9% · random 88.0% | inv-rconc 0.8% · inv-proof 89.6% | inv-rconc 0.8% · inv-proof 85.1% |
| Paraphrased rules | paraphrased (NatLang) 78.1% · templated 86.1% | paraphrased (NatLang) 54.8% · templated 67.6% | paraphrased (NatLang) 55.8% · templated 59.6% | paraphrased (NatLang) 47.6% · templated 55.9% | paraphrased (NatLang) 53.0% · templated 53.8% | templated 41.7% · paraphrased (NatLang) 44.4% |
| Theory's deepest proof | 4 67.9% · 0 90.1% | 4 40.7% · 0 87.3% | 4 42.6% · 0 87.3% | 5 38.0% · 0 85.9% | 0 42.3% · 5 58.0% | 0 39.4% · 1 48.3% |
| Theory length | 80-109 79.8% · 0-49 93.1% | 110+ 52.6% · 0-49 91.5% | 110+ 47.9% · 0-49 77.7% | 110+ 41.7% · 0-49 75.7% | 110+ 52.6% · 0-49 56.3% | 80-109 40.3% · 0-49 47.8% |
| Number of rules | 8+ 80.0% · 0-2 92.2% | 3-5 59.1% · 0-2 90.2% | 3-5 51.6% · 0-2 78.9% | 3-5 45.9% · 0-2 78.9% | 8+ 49.8% · 0-2 58.8% | 8+ 35.7% · 0-2 53.9% |
| Number of facts | 8-12 80.3% · 0-3 90.4% | 13+ 56.6% · 0-3 81.4% | 13+ 46.5% · 0-3 71.9% | 13+ 40.3% · 0-3 69.8% | 0-3 52.4% · 13+ 56.6% | 4-7 41.6% · 13+ 44.0% |
| Proof size | 0-1 80.4% · 2 98.1% | 5+ 33.6% · 0-1 85.2% | 5+ 31.8% · 2 84.1% | 5+ 20.3% · 0-1 81.4% | 0-1 32.0% · 2 77.6% | 0-1 24.3% · 2 63.6% |
How many rule applications the shortest proof of the answer needs. Depth 0 is a fact stated in the text; depth 5 needs five chained inferences. This is the benchmark's difficulty axis, and the sample was built to hold it fixed: about 300 items at each depth.
The correct answer. On the open-world task a third of the items are unknown: neither the statement nor its negation follows. That class is where several models fail.
Two kinds of "not": in the theory's facts and rules, and in the statement to judge. The preregistration predicted that negation in the rules would lower accuracy.
Whether the theory comes from ProofWriter's NatLang set, where people reworded the templated sentences, or from the templated sets. The preregistration predicted paraphrased rules would be no easier.
Every other property ProofWriter records about a problem: how long the theory is, how many rules and facts it has, how big the proof is, how deep the theory goes, whether it states attributes or relations, and how the question was generated. None of them was controlled when the problems were sampled.