Guide · evaluating decision models
Evaluating decision models: what a deep-reasoning benchmark shows
Overall accuracy hides most of what matters when you evaluate a decision model. This benchmark shows four things a single number would have missed, and how to check each one on your own decisions.
1. Score by difficulty, not with one number
Averages hide where a model stops working. On the open-world task, every model's accuracy falls between depth 0 (the answer is stated) and depth 5 (five chained inferences):
- Jev97.7% at depth 0, 81.0% at depth 581.0%
- GPT-6 Luna93.3% at depth 0, 46.0% at depth 546.0%
- Kev-9B89.3% at depth 0, 41.2% at depth 5; within reach of chance from depth 441.2%
- Kev-4B87.7% at depth 0, 36.3% at depth 5; within reach of chance from depth 436.3%
- Kev-0.8B70.3% at depth 0, 50.9% at depth 550.9%
- Laya62.7% at depth 0, 34.3% at depth 5; within reach of chance from depth 334.3%
Two models can look close on easy items and far apart on hard ones. Jev and GPT-6 Luna (reasoning off) are +4.3 points apart at depth 0 and +34.9 points apart at depth 5 on the same open-world problems. If your decisions need several steps of reasoning, the easy-item number is the wrong one to choose on. Accuracy by proof depth.
2. Check every answer, especially "unknown"
A model can reach a respectable average by ignoring one answer entirely. On the open-world task a third of the correct answers are unknown:
| Model | Recall: true | Recall: false | Recall: unknown | Most common answer |
|---|---|---|---|---|
| Jev | 88.8% | 88.3% | 74.2% | true, 37.3% of answers (correct on 33.6% of items) |
| GPT-6 Luna | 58.3% | 50.9% | 83.4% | unknown, 57.3% of answers (correct on 32.8% of items) |
| Kev-9B | 59.2% | 48.6% | 68.1% | unknown, 47.2% of answers (correct on 32.8% of items) |
| Kev-4B | 46.6% | 35.9% | 78.8% | unknown, 62.2% of answers (correct on 32.8% of items) |
| Kev-0.8B | 58.2% | 89.6% | 11.9% | false, 62.3% of answers (correct on 33.6% of items) |
| Laya | 40.2% | 85.1% | 1.0% | false, 73.7% of answers (correct on 33.6% of items) |
Recall on each answer and the answer mix take one look and catch a model that has learned a default rather than the task. True, false and unknown.
3. Read every number against a floor
Chance is a third on the open-world task and a half on the closed-world one. A stricter floor is the best constant guess: always giving the most common answer in that slice. Jev is +50.2 points above it; GPT-6 Luna is +30.4 points above it; Kev-9B is +24.9 points above it; Kev-4B is +19.9 points above it; Kev-0.8B is +19.9 points above it; Laya is +8.8 points above it (open world). A model that is only a few points above the constant guess is mostly not using the rules.
4. Compare models on the same items
Every model here answered the same problems, so differences come with a paired interval that is narrower than two separate ones. Where two models are close, only a paired comparison can say whether the gap is real. Paired comparisons.
5. Ask twice
Asked every open-world question a second time, Jev gave the same answer on 97.2% of items; 50 answers changed, mostly on deeper proofs. A decision that changes when you ask again is hard to audit and hard to threshold. Measure it on your own decisions, and check calibration if the model returns probabilities. Consistency and calibration.
6. Time it where you will run it
Latency depends on where the model runs as much as on the model. Our hosted numbers include the network; our open-model numbers are one laptop. Neither predicts your deployment. Latency · Sizes and memory.
7. Write down what you expect first
We preregistered our predictions before any model answered. So far 14 held and 4 failed, with 2 mixed or partly held. The failures were the informative part: “By depth 5, Jev is within 10 points of its own best-constant floor on at least one task (that is, it has largely stopped using the rules).”, “Negation in the theory (`theory_negation = negation`) lowers accuracy relative to no negation.”, “On OWA, Luna's `unknown` recall is its lowest class recall.” and “Accuracy rises with size: Kev-4B is above Kev-0.8B overall on both tasks (paired interval excludes zero).”. An evaluation that only confirms what you expected is easy to fool yourself with. Every prediction.
Evaluating Jev: what this benchmark shows
Jev scored 83.8% on the open-world task and 89.3% on the closed-world one, the highest of the models tested. Its weakest answer is unknown (74.2% recall against 88.8% on true); its accuracy falls from 97.7% at depth 0 to its lowest, 71.9%, at depth 4, and is 81.0% at depth 5. Paraphrased rules cost it +8.0 points. Those are the places to test first on your own data. Jev's full results.
Evaluating the OpenAI Decisions API
OpenAI's Decisions API, announced at DevDay on 29 September 2026 and in limited preview, is built on a version of GPT-6 Luna (The Decoder). We have not tested the API itself, but we ran GPT-6 Luna in the same shape of task: reasoning off, one request per problem, a fixed list of answers. It answered 64.1% of open-world problems correctly overall and 46.0% at proof depth 5. That is a preview, not a verdict: OpenAI uses a specialized version of Luna, and its accuracy may differ. When you evaluate the API on your own decisions, the same checks apply: accuracy by difficulty, recall on every answer, paired comparisons on the same items, and a second run to check consistency. Jev vs GPT-6 Luna, and what it means for the OpenAI Decisions API. Before relying on any confidence it gives, check whether that confidence separates right answers from wrong ones on hard problems: our preview of GPT-6 Luna's log-probabilities shows why.