Breakdowns · ProofWriter results

Where decision models fail: every breakdown

Accuracy broken down by every property ProofWriter records about a problem. Proof depth is the controlled difficulty axis; the others are descriptive.

At the deep end: accuracy at proof depth 5

Fully scored models only. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.

Every breakdown, with each model's weakest slice

Open-world task; weakest and strongest slices of at least 50 items.

BreakdownJevGPT-6 LunaKev-9BKev-4BKev-0.8BLaya
Proof depthdepth 4 71.9% · depth 0 97.7%depth 4 44.6% · depth 0 93.3%depth 4 38.6% · depth 0 89.3%depth 5 36.3% · depth 0 87.7%depth 2 46.9% · depth 0 70.3%depth 5 34.3% · depth 0 62.7%
Gold answerunknown 74.2% · true 88.8%false 50.9% · unknown 83.4%false 48.6% · unknown 68.1%false 35.9% · unknown 78.8%unknown 11.9% · false 89.6%unknown 1.0% · false 85.1%
Theory kindattribute 82.8% · relation 85.3%attribute 62.5% · relation 66.3%relation 57.5% · attribute 59.3%relation 52.4% · attribute 54.4%relation 53.2% · attribute 53.8%attribute 41.2% · relation 44.3%
Negation in the theorywithout negation 80.3% · with negation 87.5%without negation 61.2% · with negation 67.1%without negation 56.8% · with negation 60.4%without negation 49.9% · with negation 57.3%without negation 50.6% · with negation 56.6%without negation 38.8% · with negation 46.2%
Negated statementnegated statement 78.6% · plain statement 89.1%plain statement 62.3% · negated statement 65.8%negated statement 57.4% · plain statement 59.7%negated statement 53.5% · plain statement 53.6%plain statement 48.5% · negated statement 58.6%plain statement 38.7% · negated statement 46.2%
Question strategyinv-rconc 50.2% · random 98.0%inv-proof 50.9% · inv-random 100.0%inv-proof 48.6% · random 78.0%inv-proof 35.9% · random 88.0%inv-rconc 0.8% · inv-proof 89.6%inv-rconc 0.8% · inv-proof 85.1%
Paraphrased rulesparaphrased (NatLang) 78.1% · templated 86.1%paraphrased (NatLang) 54.8% · templated 67.6%paraphrased (NatLang) 55.8% · templated 59.6%paraphrased (NatLang) 47.6% · templated 55.9%paraphrased (NatLang) 53.0% · templated 53.8%templated 41.7% · paraphrased (NatLang) 44.4%
Theory's deepest proof4 67.9% · 0 90.1%4 40.7% · 0 87.3%4 42.6% · 0 87.3%5 38.0% · 0 85.9%0 42.3% · 5 58.0%0 39.4% · 1 48.3%
Theory length80-109 79.8% · 0-49 93.1%110+ 52.6% · 0-49 91.5%110+ 47.9% · 0-49 77.7%110+ 41.7% · 0-49 75.7%110+ 52.6% · 0-49 56.3%80-109 40.3% · 0-49 47.8%
Number of rules8+ 80.0% · 0-2 92.2%3-5 59.1% · 0-2 90.2%3-5 51.6% · 0-2 78.9%3-5 45.9% · 0-2 78.9%8+ 49.8% · 0-2 58.8%8+ 35.7% · 0-2 53.9%
Number of facts8-12 80.3% · 0-3 90.4%13+ 56.6% · 0-3 81.4%13+ 46.5% · 0-3 71.9%13+ 40.3% · 0-3 69.8%0-3 52.4% · 13+ 56.6%4-7 41.6% · 13+ 44.0%
Proof size0-1 80.4% · 2 98.1%5+ 33.6% · 0-1 85.2%5+ 31.8% · 2 84.1%5+ 20.3% · 0-1 81.4%0-1 32.0% · 2 77.6%0-1 24.3% · 2 63.6%