Benchmark for decision models and LLM classifiers · ProofWriter
Decision model benchmark: how accuracy falls as reasoning gets deeper
Fast decision models and LLM classifiers answer a question about a text in one step. We gave them 3,600 multi-hop reasoning problems from ProofWriter, from statements read straight off the page to ones that need five chained inferences, and measured where each one breaks.
6 models so far: hosted and open decision models, and a general LLM used as a classifier. Every result below links to the problems and the numbers behind it.
83.8%Jev answered 83.8% of open-world problems correctly and 89.3% of closed-world ones, the highest of the models tested. Its lead over GPT-6 Luna (reasoning off) is +4.3 points at depth 0 and +34.9 points at depth 5 on the same problems (+1.3 and +44.7 on the closed-world task).
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
Open world: true, false or unknown
Closed world: true or false
Overall accuracy on both tasks
The same 1,800 problems per task for every model. Models are ordered by accuracy; a model still being scored is listed last and marked partial.
Open world: true, false or unknown
A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does.
| Model | Accuracy | 95% interval | Macro-F1 | Recall: true | Recall: false | Recall: unknown | Items |
|---|---|---|---|---|---|---|---|
| Jev | 83.8% | 82.2–85.5 | 0.838 | 88.8% | 88.3% | 74.2% | 1,800 |
| GPT-6 Luna | 64.1% | 61.7–66.2 | 0.647 | 58.3% | 50.9% | 83.4% | 1,800 |
| Kev-9B | 58.6% | 56.3–60.9 | 0.588 | 59.2% | 48.6% | 68.1% | 1,800 |
| Kev-4B | 53.6% | 51.4–56.0 | 0.532 | 46.6% | 35.9% | 78.8% | 1,800 |
| Kev-0.8B | 53.6% | 51.4–55.9 | 0.477 | 58.2% | 89.6% | 11.9% | 1,800 |
| Laya | 42.4% | 40.3–44.5 | 0.337 | 40.2% | 85.1% | 1.0% | 1,800 |
Chance is 33.3%; always giving the most common answer scores 33.6% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.
Closed world: true or false
Anything that cannot be derived is false, so there is no unknown answer.
| Model | Accuracy | 95% interval | Macro-F1 | Recall: true | Recall: false | Items |
|---|---|---|---|---|---|---|
| Jev | 89.3% | 87.8–90.7 | 0.893 | 89.7% | 88.9% | 1,800 |
| GPT-6 Luna | 65.0% | 62.8–67.4 | 0.650 | 66.2% | 63.8% | 1,800 |
| Kev-9B | 63.8% | 61.5–66.1 | 0.638 | 64.7% | 62.9% | 1,800 |
| Kev-4B | 58.8% | 56.6–61.1 | 0.576 | 42.2% | 75.3% | 1,800 |
| Kev-0.8B | 55.8% | 53.5–57.9 | 0.553 | 45.4% | 66.1% | 1,800 |
| Laya | 55.6% | 53.3–58.1 | 0.509 | 24.8% | 86.4% | 1,800 |
Chance is 50.0%; always giving the most common answer scores 50.0% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.
Every model gets worse as proofs get deeper
As proofs get deeper, every model we tested except Jev falls to coin-flip accuracy by five chained inferences: the open decision models Kev and Laya, and GPT-6 Luna used as a one-shot classifier. Jev still answers 81.0–89.3% correctly at that depth.
Proof depth is the number of rules that must be chained to reach the answer. It is the one difficulty axis the sample controls.
- Jev97.7% at depth 0, 81.0% at depth 5−16.7
- GPT-6 Luna93.3% at depth 0, 46.0% at depth 5−47.3
- Kev-9B89.3% at depth 0, 41.2% at depth 5−48.2
- Kev-4B87.7% at depth 0, 36.3% at depth 5−51.3
- Kev-0.8B70.3% at depth 0, 50.9% at depth 5−19.5
- Laya62.7% at depth 0, 34.3% at depth 5−28.4
Open-world task, change from depth 0 to depth 5 in points. Accuracy by proof depth has every depth, both tasks and the paired gaps.
Head-to-head
a decision model against an LLM classifier. OpenAI's new Decisions API is built on a version of GPT-6 Luna, so these results preview the accuracy of a Luna-based decision model on multi-step reasoning.
is an open-source decision model a Jev alternative?
hosted and open decision models on deep reasoning
GPT-6 Luna is the model behind OpenAI's new Decisions API: a preview of its accuracy and confidence on deep reasoning.
Where they fail
Accuracy by every property ProofWriter records. Only proof depth was controlled when the problems were sampled; the others move with depth and with each other, so they describe where errors fall rather than what causes them.
How many rule applications the shortest proof of the answer needs.
The correct answer.
Two kinds of "not": in the theory's facts and rules, and in the statement to judge.
Whether the theory comes from ProofWriter's NatLang set, where people reworded the templated sentences, or from the templated sets.
Every other property ProofWriter records about a problem: how long the theory is, how many rules and facts it has, how big the proof is, how deep the theory goes, whether it states attributes or relations, and how the question was generated.
How we measured, consistency, speed and size
The sample, the protocol and the scoring, and the predictions written before any model answered: 1 partly held, 14 held, 4 failed, 1 mixed, 2 pending.
Asked every question twice: Jev 97.2%, GPT-6 Luna 89.7%, Kev-4B 100.0%, Kev-0.8B 100.0% and Laya 100.0% the same on the open-world task.
Jev: a median 189 ms per open-world decision and 170 ms per closed-world one, over the network; open models were timed on a laptop, so not like for like. Parameters, weights on disk and the memory each open model used.
For teams evaluating or tuning a decision model
Overall accuracy hides most of what matters when you evaluate a decision model. This benchmark shows four things a single number would have missed, and how to check each one on your own decisions.
The failures in this benchmark are specific enough to plan training data around: deep proofs, the unknown answer, reworded rules and a lean toward one answer. Here is what the results show, model by model.
A decision model can be accurate and still answer a different question from the one you meant. The open-world and closed-world tasks ask the same problems under two rules for what "unknown" means; the gap between them is an alignment problem you can measure.
Not on the board yet
- Kev-27B: not run. does not fit this 32 GB machine; would need a rented 80 GB GPU.
A model that has not been scored is never shown as zero. Every model.