Comparison · Jev vs LLM

Jev vs GPT-6 Luna: a decision model against an LLM classifier

Jev is built to make one decision per request. GPT-6 Luna is a general large language model, used here as a classifier with reasoning off. Both answered the same 3,600 problems.

GPT-6 Luna ran with reasoning effort set to none, its lowest setting, under the same one-request protocol as every model here: the same question, the same options, one answer per item. A different reasoning setting would be a different model, scored separately.

OpenAI's new Decisions API is built on a version of GPT-6 Luna, so these results preview the accuracy of a Luna-based decision model on multi-step reasoning. They are not a test of the API itself. What this means for the OpenAI Decisions API.

Overall: Jev and GPT-6 Luna

Open world: true, false or unknown

Jev minus GPT-6 Luna: +19.8 points on the same 1,800 items, 95% interval +17.2 to +22.4: Jev is clearly ahead.

Closed world: true or false

Jev minus GPT-6 Luna: +24.3 points on the same 1,800 items, 95% interval +21.9 to +26.6: Jev is clearly ahead.

Jev vs GPT-6 Luna by proof depth

Each point is about 300 problems that need that many inference steps, with its 95% interval.

Open world: true, false or unknown
  • Jev
  • GPT-6 Luna(reasoning off)
0%25%50%75%100%012345proof depth (inference steps)chance 33%
Closed world: true or false
  • Jev
  • GPT-6 Luna(reasoning off)
0%25%50%75%100%012345proof depth (inference steps)chance 50%

Jev minus GPT-6 Luna by depth, open world

depth 0300 items
+4.3
depth 1302 items
+9.9
depth 2303 items
+24.4
depth 3303 items
+18.2
depth 4303 items
+27.4
depth 5289 items
+34.9

Jev minus GPT-6 Luna by depth, closed world

depth 0300 items
+1.3
depth 1300 items
+17.0
depth 2300 items
+26.7
depth 3300 items
+28.0
depth 4300 items
+28.0
depth 5300 items
+44.7

Paired differences: positive means Jev was more accurate on the same items.

Accuracy by depth, open world

By depth, Open world
Proof depthItemsJevGPT-6 LunaBest constant
depth 030097.7% 96.0–99.393.3% 90.3–96.033.3%
depth 130287.7% 83.8–91.477.8% 73.2–82.833.4%
depth 230384.8% 80.2–88.460.4% 54.5–66.033.3%
depth 330379.9% 75.2–84.561.7% 56.8–67.733.3%
depth 430371.9% 66.7–77.244.6% 38.6–49.533.3%
depth 528981.0% 76.5–85.546.0% 40.1–51.634.9%

Accuracy by depth, closed world

By depth, Closed world
Proof depthItemsJevGPT-6 LunaBest constant
depth 030099.3% 98.3–100.098.0% 96.3–99.350.0%
depth 130093.7% 91.0–96.376.7% 72.3–81.050.0%
depth 230092.0% 88.7–94.765.3% 60.0–70.750.0%
depth 330085.0% 81.3–89.357.0% 51.7–63.050.0%
depth 430076.3% 71.7–81.048.3% 43.0–53.750.0%
depth 530089.3% 85.7–92.344.7% 38.7–50.350.0%

What this means for the OpenAI Decisions API

At OpenAI DevDay on 29 September 2026, OpenAI announced the Decisions API, in limited preview with broader availability "in the coming days". OpenAI says it "uses a version of GPT-6 Luna": the model picks among answers the developer defines in advance, to classify, route or choose an agent's next action (The Decoder, Axios).

That is the shape of task we ran: GPT-6 Luna with reasoning off, one request per problem, and a fixed list of answers enforced with a strict JSON schema. So these results preview the accuracy a Luna-based decision model gets on multi-step reasoning. GPT-6 Luna answered 64.1% of open-world problems and 65.0% of closed-world ones correctly. On the open-world task it scored 93.3% when the answer is stated outright (depth 0) and 46.0% when it takes five chained inferences, where Jev scored 81.0% (+34.9 points on the same items). If your decisions need several steps of reasoning, the depth results above are the ones to read.

These are not results for the Decisions API itself. OpenAI uses a specialized version of Luna, and its accuracy may differ. As of 1 October 2026 OpenAI had published no Decisions API documentation. We will benchmark the API itself when we can get access.

On speed, OpenAI's own chart shows 150 ms per decision for the Decisions API, against 1.6 s for Luna through the regular API (The Decoder). Those are OpenAI's figures, not our measurements.

Confidence and log-probabilities: things to know

Where each one is better

Open world, Jev vs GPT-6 Luna. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. GPT-6 Luna is clearly ahead on unknown answers (+9.2 points).

Closed world, Jev vs GPT-6 Luna. Jev is clearly ahead on depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. GPT-6 Luna is not clearly ahead on any depth or answer class. Too close to call: depth 0.

A slice counts as clearly ahead when its paired 95% interval excludes zero.

True, false and unknown, open world

By correct answer, Open world
Correct answerItemsJevGPT-6 LunaBest constant
false60588.3% 85.5–90.950.9% 46.8–54.9one answer
true60588.8% 86.1–91.158.3% 54.0–62.1one answer
unknown59074.2% 70.7–77.683.4% 80.5–86.6one answer

True, false and unknown, closed world

By correct answer, Closed world
Correct answerItemsJevGPT-6 LunaBest constant
false90088.9% 86.9–90.963.8% 60.6–66.9one answer
true90089.7% 87.7–91.666.2% 63.1–69.2one answer

Paraphrased vs templated rules, open world

By wording, Open world
WordingItemsJevGPT-6 LunaBest constant
templated1,29886.1% 84.2–87.967.6% 64.8–70.234.5%
paraphrased (NatLang)50278.1% 74.5–82.154.8% 50.4–59.435.9%

Paraphrased vs templated rules, closed world

By wording, Closed world
WordingItemsJevGPT-6 LunaBest constant
templated1,30694.0% 92.8–95.367.3% 64.8–69.950.2%
paraphrased (NatLang)49476.7% 73.1–80.458.9% 54.7–63.450.6%

Where GPT-6 Luna errs

Accuracy alone hides which answers a model gets wrong. Rows are the correct answer; columns are what the model said.

GPT-6 Luna, open-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseSaid unknownRecall
true353324958.3%
false730829050.9%
unknown455349283.4%

GPT-6 Luna answered false 20.2%, true 22.5%, unknown 57.3% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.

GPT-6 Luna, closed-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseRecall
true59630466.2%
false32657463.8%

GPT-6 Luna answered false 48.8%, true 51.2% of the time; the correct answers are false 50.0%, true 50.0%.

GPT-6 Luna across every breakdown

For each property of the problems, the weakest and strongest value on the open-world task (slices of at least 50 items). Only depth was controlled; the rest are descriptive.

BreakdownWeakest · strongest
Proof depthdepth 4 44.6% · depth 0 93.3% (48.8 points)
Gold answerfalse 50.9% · unknown 83.4% (32.5 points)
Theory kindattribute 62.5% · relation 66.3% (3.8 points)
Negation in the theorywithout negation 61.2% · with negation 67.1% (5.9 points)
Negated statementplain statement 62.3% · negated statement 65.8% (3.5 points)
Question strategyinv-proof 50.9% · inv-random 100.0% (49.1 points)
Paraphrased rulesparaphrased (NatLang) 54.8% · templated 67.6% (12.9 points)
Theory's deepest proof4 40.7% · 0 87.3% (46.6 points)
Theory length110+ 52.6% · 0-49 91.5% (38.9 points)
Number of rules3-5 59.1% · 0-2 90.2% (31.1 points)
Number of facts13+ 56.6% · 0-3 81.4% (24.8 points)
Proof size5+ 33.6% · 0-1 85.2% (51.6 points)

Consistency and calibration

Asked every open world question a second time, GPT-6 Luna gave the same answer on 89.7% of 1,800 items (185 changed; Gwet's AC1 0.855, interval 0.837 to 0.876). Accuracy was 64.1% the first time and 62.8% the second.

Asked every closed world question a second time, GPT-6 Luna gave the same answer on 90.6% of 1,800 items (169 changed; Gwet's AC1 0.812, interval 0.784 to 0.838). Accuracy was 65.0% the first time and 64.8% the second.

We asked GPT-6 Luna, like every model, for its answer only, so no confidence was recorded and its calibration is not measured here. A separate preregistered probe read its log-probabilities: see its calibration on the OpenAI Decisions API preview page.

Size, hardware and speed

ModelKindParametersRan onMedian time per decisionMemory peak
Jevhosted decision modelundisclosedvendor's servers189 msnot applicable (hosted)
GPT-6 LunaLLM classifierundisclosedvendor's servers825 msnot applicable (hosted)

Times are open-world, one request at a time. Hosted models were timed over the network against the vendor's servers and open models on one laptop, so speed is not like for like across the two. Speed, size and memory for every model.

GPT-6 Luna

Parameters
undisclosed
Architecture
undisclosed
Weights
not available (hosted)
Where it ran
OpenAI Chat Completions API, reasoning_effort none, strict JSON schema
Maker's stated hardware
not applicable (hosted)
Price
$0.10 input / $0.50 output per million tokens
Time per decision, open world
median 825 ms, 90th percentile 1,165 ms, over the network
Time per decision, closed world
median 929 ms, 90th percentile 1,268 ms, over the network
Source
developers.openai.com/api/docs/models/gpt-6-luna

The models, and what we predicted

Jev Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.

GPT-6 Luna GPT-6 Luna is a hosted large language model made by OpenAI. We used it as a classifier: one request per item through the Chat Completions API, with reasoning effort set to none (its lowest setting) and a strict JSON schema that allows only the task's options. Its size and architecture are not published.