Benchmark for decision models and LLM classifiers · ProofWriter

Decision model benchmark: how accuracy falls as reasoning gets deeper

Fast decision models and LLM classifiers answer a question about a text in one step. We gave them 3,600 multi-hop reasoning problems from ProofWriter, from statements read straight off the page to ones that need five chained inferences, and measured where each one breaks.

6 models so far: hosted and open decision models, and a general LLM used as a classifier. Every result below links to the problems and the numbers behind it.

83.8%Jev answered 83.8% of open-world problems correctly and 89.3% of closed-world ones, the highest of the models tested. Its lead over GPT-6 Luna (reasoning off) is +4.3 points at depth 0 and +34.9 points at depth 5 on the same problems (+1.3 and +44.7 on the closed-world task).

Accuracy by proof depth, both tasksEach line is one model; each point is about 300 problems that need that many inference steps, with its 95% interval. The dashed line is chance: one in three on the open-world task, one in two on the closed-world task. Every number, both tasks.
  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya

Open world: true, false or unknown

0%25%50%75%100%012345proof depth (inference steps)chance 33%

Closed world: true or false

0%25%50%75%100%012345proof depth (inference steps)chance 50%

Overall accuracy on both tasks

The same 1,800 problems per task for every model. Models are ordered by accuracy; a model still being scored is listed last and marked partial.

Open world: true, false or unknown

A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does.

Overall accuracy, Open world: true, false or unknown
ModelAccuracy95% intervalMacro-F1Recall: trueRecall: falseRecall: unknownItems
Jev83.8%82.2–85.50.83888.8%88.3%74.2%1,800
GPT-6 Luna64.1%61.7–66.20.64758.3%50.9%83.4%1,800
Kev-9B58.6%56.3–60.90.58859.2%48.6%68.1%1,800
Kev-4B53.6%51.4–56.00.53246.6%35.9%78.8%1,800
Kev-0.8B53.6%51.4–55.90.47758.2%89.6%11.9%1,800
Laya42.4%40.3–44.50.33740.2%85.1%1.0%1,800

Chance is 33.3%; always giving the most common answer scores 33.6% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.

Closed world: true or false

Anything that cannot be derived is false, so there is no unknown answer.

Overall accuracy, Closed world: true or false
ModelAccuracy95% intervalMacro-F1Recall: trueRecall: falseItems
Jev89.3%87.8–90.70.89389.7%88.9%1,800
GPT-6 Luna65.0%62.8–67.40.65066.2%63.8%1,800
Kev-9B63.8%61.5–66.10.63864.7%62.9%1,800
Kev-4B58.8%56.6–61.10.57642.2%75.3%1,800
Kev-0.8B55.8%53.5–57.90.55345.4%66.1%1,800
Laya55.6%53.3–58.10.50924.8%86.4%1,800

Chance is 50.0%; always giving the most common answer scores 50.0% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.

Every model gets worse as proofs get deeper

As proofs get deeper, every model we tested except Jev falls to coin-flip accuracy by five chained inferences: the open decision models Kev and Laya, and GPT-6 Luna used as a one-shot classifier. Jev still answers 81.0–89.3% correctly at that depth.

Proof depth is the number of rules that must be chained to reach the answer. It is the one difficulty axis the sample controls.

Open-world task, change from depth 0 to depth 5 in points. Accuracy by proof depth has every depth, both tasks and the paired gaps.

Head-to-head

Where they fail

Accuracy by every property ProofWriter records. Only proof depth was controlled when the problems were sampled; the others move with depth and with each other, so they describe where errors fall rather than what causes them.

How we measured, consistency, speed and size

For teams evaluating or tuning a decision model

Not on the board yet

A model that has not been scored is never shown as zero. Every model.