Comparison · Jev vs LLM
Jev vs GPT-6 Luna: a decision model against an LLM classifier
Jev is built to make one decision per request. GPT-6 Luna is a general large language model, used here as a classifier with reasoning off. Both answered the same 3,600 problems.
GPT-6 Luna ran with reasoning effort set to none, its lowest setting, under the same one-request protocol as every model here: the same question, the same options, one answer per item. A different reasoning setting would be a different model, scored separately.
OpenAI's new Decisions API is built on a version of GPT-6 Luna, so these results preview the accuracy of a Luna-based decision model on multi-step reasoning. They are not a test of the API itself. What this means for the OpenAI Decisions API.
Overall: Jev and GPT-6 Luna
Open world: true, false or unknown
- Jev82.2 to 85.5 (95%), 1,800 items83.8%
- GPT-6 Luna61.7 to 66.2 (95%), 1,800 items64.1%
Jev minus GPT-6 Luna: +19.8 points on the same 1,800 items, 95% interval +17.2 to +22.4: Jev is clearly ahead.
Closed world: true or false
- Jev87.8 to 90.7 (95%), 1,800 items89.3%
- GPT-6 Luna62.8 to 67.4 (95%), 1,800 items65.0%
Jev minus GPT-6 Luna: +24.3 points on the same 1,800 items, 95% interval +21.9 to +26.6: Jev is clearly ahead.
Jev vs GPT-6 Luna by proof depth
Each point is about 300 problems that need that many inference steps, with its 95% interval.
- Jev
- GPT-6 Luna(reasoning off)
- Jev
- GPT-6 Luna(reasoning off)
Jev minus GPT-6 Luna by depth, open world
Jev minus GPT-6 Luna by depth, closed world
Paired differences: positive means Jev was more accurate on the same items.
Accuracy by depth, open world
| Proof depth | Items | Jev | GPT-6 Luna | Best constant |
|---|---|---|---|---|
| depth 0 | 300 | 97.7% 96.0–99.3 | 93.3% 90.3–96.0 | 33.3% |
| depth 1 | 302 | 87.7% 83.8–91.4 | 77.8% 73.2–82.8 | 33.4% |
| depth 2 | 303 | 84.8% 80.2–88.4 | 60.4% 54.5–66.0 | 33.3% |
| depth 3 | 303 | 79.9% 75.2–84.5 | 61.7% 56.8–67.7 | 33.3% |
| depth 4 | 303 | 71.9% 66.7–77.2 | 44.6% 38.6–49.5 | 33.3% |
| depth 5 | 289 | 81.0% 76.5–85.5 | 46.0% 40.1–51.6 | 34.9% |
Accuracy by depth, closed world
| Proof depth | Items | Jev | GPT-6 Luna | Best constant |
|---|---|---|---|---|
| depth 0 | 300 | 99.3% 98.3–100.0 | 98.0% 96.3–99.3 | 50.0% |
| depth 1 | 300 | 93.7% 91.0–96.3 | 76.7% 72.3–81.0 | 50.0% |
| depth 2 | 300 | 92.0% 88.7–94.7 | 65.3% 60.0–70.7 | 50.0% |
| depth 3 | 300 | 85.0% 81.3–89.3 | 57.0% 51.7–63.0 | 50.0% |
| depth 4 | 300 | 76.3% 71.7–81.0 | 48.3% 43.0–53.7 | 50.0% |
| depth 5 | 300 | 89.3% 85.7–92.3 | 44.7% 38.7–50.3 | 50.0% |
What this means for the OpenAI Decisions API
At OpenAI DevDay on 29 September 2026, OpenAI announced the Decisions API, in limited preview with broader availability "in the coming days". OpenAI says it "uses a version of GPT-6 Luna": the model picks among answers the developer defines in advance, to classify, route or choose an agent's next action (The Decoder, Axios).
That is the shape of task we ran: GPT-6 Luna with reasoning off, one request per problem, and a fixed list of answers enforced with a strict JSON schema. So these results preview the accuracy a Luna-based decision model gets on multi-step reasoning. GPT-6 Luna answered 64.1% of open-world problems and 65.0% of closed-world ones correctly. On the open-world task it scored 93.3% when the answer is stated outright (depth 0) and 46.0% when it takes five chained inferences, where Jev scored 81.0% (+34.9 points on the same items). If your decisions need several steps of reasoning, the depth results above are the ones to read.
These are not results for the Decisions API itself. OpenAI uses a specialized version of Luna, and its accuracy may differ. As of 1 October 2026 OpenAI had published no Decisions API documentation. We will benchmark the API itself when we can get access.
On speed, OpenAI's own chart shows 150 ms per decision for the Decisions API, against 1.6 s for Luna through the regular API (The Decoder). Those are OpenAI's figures, not our measurements.
Confidence and log-probabilities: things to know
- The scored benchmark recorded no confidence. We asked every model for its answer only. A separate, preregistered probe then read GPT-6 Luna's log-probabilities on all 3,600 items: when Luna stated 99% or more, it was right 68.7% of the time on open-world problems. The full preview, with a verbatim request and response.
- With reasoning off, Luna exposes its log-probabilities. It returns the chosen answer's log-probability and the most likely alternatives, which cover 99.8% and 99.9% of the probability on average (open world, closed world). With reasoning on, the request is refused: "'logprobs' is not supported with this model". OpenAI's guide says to remove
logprobsandtop_logprobswhen reasoning effort is not none (OpenAI's latest-model guide; GPT-6 Luna's model page). - Overconfidence. OpenAI's GPT-4 Technical Report found the pre-trained model well calibrated, and that post-training reduced calibration markedly (its Figure 8). In our experience, where OpenAI models do expose log-probabilities, they tend to be too overconfident to calibrate into a useful confidence.
- Why log-probabilities are limited with reasoning on: our suspicion, with no proof. Researchers recovered the final layer of production OpenAI models from API log-probabilities and logit bias, and OpenAI changed its API in response (Carlini et al., 2024; see also Finlayson et al., 2024). OpenAI also chose not to show o1's raw chain of thought, citing user experience, competitive advantage and chain-of-thought monitoring (OpenAI on o1's chain of thought). None of this shows why Luna's log-probabilities are limited.
Where each one is better
Open world, Jev vs GPT-6 Luna. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. GPT-6 Luna is clearly ahead on unknown answers (+9.2 points).
Closed world, Jev vs GPT-6 Luna. Jev is clearly ahead on depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. GPT-6 Luna is not clearly ahead on any depth or answer class. Too close to call: depth 0.
A slice counts as clearly ahead when its paired 95% interval excludes zero.
True, false and unknown, open world
| Correct answer | Items | Jev | GPT-6 Luna | Best constant |
|---|---|---|---|---|
| false | 605 | 88.3% 85.5–90.9 | 50.9% 46.8–54.9 | one answer |
| true | 605 | 88.8% 86.1–91.1 | 58.3% 54.0–62.1 | one answer |
| unknown | 590 | 74.2% 70.7–77.6 | 83.4% 80.5–86.6 | one answer |
True, false and unknown, closed world
| Correct answer | Items | Jev | GPT-6 Luna | Best constant |
|---|---|---|---|---|
| false | 900 | 88.9% 86.9–90.9 | 63.8% 60.6–66.9 | one answer |
| true | 900 | 89.7% 87.7–91.6 | 66.2% 63.1–69.2 | one answer |
Paraphrased vs templated rules, open world
| Wording | Items | Jev | GPT-6 Luna | Best constant |
|---|---|---|---|---|
| templated | 1,298 | 86.1% 84.2–87.9 | 67.6% 64.8–70.2 | 34.5% |
| paraphrased (NatLang) | 502 | 78.1% 74.5–82.1 | 54.8% 50.4–59.4 | 35.9% |
Paraphrased vs templated rules, closed world
| Wording | Items | Jev | GPT-6 Luna | Best constant |
|---|---|---|---|---|
| templated | 1,306 | 94.0% 92.8–95.3 | 67.3% 64.8–69.9 | 50.2% |
| paraphrased (NatLang) | 494 | 76.7% 73.1–80.4 | 58.9% 54.7–63.4 | 50.6% |
Where GPT-6 Luna errs
Accuracy alone hides which answers a model gets wrong. Rows are the correct answer; columns are what the model said.
GPT-6 Luna, open-world task
| Correct answer | Said true | Said false | Said unknown | Recall |
|---|---|---|---|---|
| true | 353 | 3 | 249 | 58.3% |
| false | 7 | 308 | 290 | 50.9% |
| unknown | 45 | 53 | 492 | 83.4% |
GPT-6 Luna answered false 20.2%, true 22.5%, unknown 57.3% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.
GPT-6 Luna, closed-world task
| Correct answer | Said true | Said false | Recall |
|---|---|---|---|
| true | 596 | 304 | 66.2% |
| false | 326 | 574 | 63.8% |
GPT-6 Luna answered false 48.8%, true 51.2% of the time; the correct answers are false 50.0%, true 50.0%.
GPT-6 Luna across every breakdown
For each property of the problems, the weakest and strongest value on the open-world task (slices of at least 50 items). Only depth was controlled; the rest are descriptive.
| Breakdown | Weakest · strongest |
|---|---|
| Proof depth | depth 4 44.6% · depth 0 93.3% (48.8 points) |
| Gold answer | false 50.9% · unknown 83.4% (32.5 points) |
| Theory kind | attribute 62.5% · relation 66.3% (3.8 points) |
| Negation in the theory | without negation 61.2% · with negation 67.1% (5.9 points) |
| Negated statement | plain statement 62.3% · negated statement 65.8% (3.5 points) |
| Question strategy | inv-proof 50.9% · inv-random 100.0% (49.1 points) |
| Paraphrased rules | paraphrased (NatLang) 54.8% · templated 67.6% (12.9 points) |
| Theory's deepest proof | 4 40.7% · 0 87.3% (46.6 points) |
| Theory length | 110+ 52.6% · 0-49 91.5% (38.9 points) |
| Number of rules | 3-5 59.1% · 0-2 90.2% (31.1 points) |
| Number of facts | 13+ 56.6% · 0-3 81.4% (24.8 points) |
| Proof size | 5+ 33.6% · 0-1 85.2% (51.6 points) |
Consistency and calibration
Asked every open world question a second time, GPT-6 Luna gave the same answer on 89.7% of 1,800 items (185 changed; Gwet's AC1 0.855, interval 0.837 to 0.876). Accuracy was 64.1% the first time and 62.8% the second.
Asked every closed world question a second time, GPT-6 Luna gave the same answer on 90.6% of 1,800 items (169 changed; Gwet's AC1 0.812, interval 0.784 to 0.838). Accuracy was 65.0% the first time and 64.8% the second.
We asked GPT-6 Luna, like every model, for its answer only, so no confidence was recorded and its calibration is not measured here. A separate preregistered probe read its log-probabilities: see its calibration on the OpenAI Decisions API preview page.
Size, hardware and speed
| Model | Kind | Parameters | Ran on | Median time per decision | Memory peak |
|---|---|---|---|---|---|
| Jev | hosted decision model | undisclosed | vendor's servers | 189 ms | not applicable (hosted) |
| GPT-6 Luna | LLM classifier | undisclosed | vendor's servers | 825 ms | not applicable (hosted) |
Times are open-world, one request at a time. Hosted models were timed over the network against the vendor's servers and open models on one laptop, so speed is not like for like across the two. Speed, size and memory for every model.
GPT-6 Luna
- Parameters
- undisclosed
- Architecture
- undisclosed
- Weights
- not available (hosted)
- Where it ran
- OpenAI Chat Completions API, reasoning_effort none, strict JSON schema
- Maker's stated hardware
- not applicable (hosted)
- Price
- $0.10 input / $0.50 output per million tokens
- Time per decision, open world
- median 825 ms, 90th percentile 1,165 ms, over the network
- Time per decision, closed world
- median 929 ms, 90th percentile 1,268 ms, over the network
The models, and what we predicted
Jev Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.
GPT-6 Luna GPT-6 Luna is a hosted large language model made by OpenAI. We used it as a classifier: one request per item through the Chat Completions API, with reasoning effort set to none (its lowest setting) and a strict JSON schema that allows only the task's options. Its size and architecture are not published.
- 2 of 4 predictions about GPT-6 Luna held (1 failed, 1 mixed). See How we measured.
- 4 of 4 predictions about Luna confidence probe held. See How we measured.
Other comparisons: Kev vs Jev · Jev vs Laya · Every model