Preview · built on GPT-6 Luna, not on the API itself

Previewing the OpenAI Decisions API through GPT-6 Luna: accuracy and confidence on deep reasoning

OpenAI's new Decisions API is built on a version of GPT-6 Luna. We have not tested the API itself. We tested GPT-6 Luna the way such a decision model works: one request, a fixed list of answers, reasoning off. Here is how accurate it is on multi-step reasoning, and whether its own probabilities can tell you when to trust an answer.

68.7%When GPT-6 Luna stated 99% or more, it was right 68.7% of the time on open-world problems and 68.1% on closed-world ones, across 1,302 and 1,370 answers. Jev at 99% or more was right 99.4% and 98.5% of the time on the same problems.

We have not tested the Decisions API. OpenAI uses a specialized version of Luna, and its accuracy and confidence may differ from what is shown here. We will benchmark the API itself when we can get access.

What the OpenAI Decisions API is

At OpenAI DevDay on 29 September 2026, OpenAI announced the Decisions API, in limited preview with broader availability "in the coming days". OpenAI says it "uses a version of GPT-6 Luna": the model picks among answers the developer defines in advance, to classify, route or choose an agent's next action (The Decoder, Axios).

As of 1 October 2026 OpenAI had published no Decisions API documentation. What is public is the model it is built on. GPT-6 Luna, used with reasoning off, one request per problem and a fixed list of answers enforced by a strict JSON schema, is the same shape of task, and that is how it ran in this benchmark. So this page previews a Luna-based decision model; it does not measure the API.

Accuracy: a Luna-based decision model against Jev, by proof depth

From the scored benchmark: the same 1,800 problems per task for both models, one request each, the same question and options. GPT-6 Luna answered 64.1% of open-world problems correctly and Jev 83.8%; on the closed-world task, 65.0% and 89.3%.

Open world: true, false or unknown
  • Jev
  • GPT-6 Luna(reasoning off)
0%25%50%75%100%012345proof depth (inference steps)chance 33%
Closed world: true or false
  • Jev
  • GPT-6 Luna(reasoning off)
0%25%50%75%100%012345proof depth (inference steps)chance 50%

Jev minus GPT-6 Luna by depth, open-world task

depth 0300 items
+4.3
depth 1302 items
+9.9
depth 2303 items
+24.4
depth 3303 items
+18.2
depth 4303 items
+27.4
depth 5289 items
+34.9

Jev minus GPT-6 Luna by depth, closed-world task

depth 0300 items
+1.3
depth 1300 items
+17.0
depth 2300 items
+26.7
depth 3300 items
+28.0
depth 4300 items
+28.0
depth 5300 items
+44.7

Accuracy by depth, open-world task

By depth, open world
Proof depthItemsJevGPT-6 LunaBest constant
depth 030097.7% 96.0–99.393.3% 90.3–96.033.3%
depth 130287.7% 83.8–91.477.8% 73.2–82.833.4%
depth 230384.8% 80.2–88.460.4% 54.5–66.033.3%
depth 330379.9% 75.2–84.561.7% 56.8–67.733.3%
depth 430371.9% 66.7–77.244.6% 38.6–49.533.3%
depth 528981.0% 76.5–85.546.0% 40.1–51.634.9%

Can Luna give you a confidence?

A decision model is most useful when it also says how sure it is, so you can act on confident answers and send the rest to a person. The scored benchmark asked every model for its answer only. To see whether GPT-6 Luna could supply a confidence, we read its log-probabilities.

A language model writes its reply one token (a word or piece of a word) at a time, and at each step it assigns a probability to every possible next token. A log-probability is the natural logarithm of that probability: 0 means certain, and the more negative it is, the less likely. With logprobs turned on, the API returns the log-probability of each token it wrote, and with top_logprobs the most likely alternatives at each step.

Our requests used a strict JSON schema whose only allowed values are the task's options, so every reply is a small object such as {"answer":"false"}. We find the tokens that spell the answer inside that reply, add their log-probabilities and convert the sum back to a probability: that is the probability Luna gave its own answer. The options begin differently ("true", "false" and "unknown"), so the first answer token already identifies one option, and each alternative Luna lists at that position maps to exactly one other option.

With reasoning off, Luna exposed these probabilities almost completely. On most responses not every option appears among the alternatives, but the options it lists carry 99.8% and 99.9% of the probability on average (open world, closed world), so the ones left out carry almost none. With reasoning on, the API refuses the parameter: "'logprobs' is not supported with this model" (OpenAI's latest-model guide).

A real request and response, verbatim

One of the probe's 3,600 requests, exactly as sent and as returned: a open-world problem at proof depth 3 (AttNonegNatLang-OWA-479-Q8). The correct answer is false. Luna answered unknown, and the log-probability of that answer's token, -0.0018, is a probability of 99.82%. It listed no other option at that position.

That is a confidently wrong answer: stated as close to certain, and wrong. On its own it proves nothing; the results below show how often it happens.

The request body

{
  "logprobs": true,
  "max_completion_tokens": 1024,
  "messages": [
    {
      "content": "Alan is young, round, and kind, but that doesn't mean he isn't also rough and cold at times, as well. That guy Fred sure is nice. That guy Harry sure is nice. Young round people who are green are usually blue. When green, young and round fits a person, you'll see that rough will also fit. It's a certainty that any green, big and kind individual is going to be nice. Kind people with rough skin are usually red because it's wind burn. If someone shows that they are red, then they are also showing that they are green. Is a kind person if feeling cold and looking red then they are usually young.\n\nStatement: Alan is not blue.\n\nUsing only the facts and rules in the text, is the statement true, false, or unknown? A statement is true if it can be derived from the facts and rules, and false if its negation can be derived. If neither the statement nor its negation can be derived, it is unknown. A rule applies only when all of its conditions are established; 'not' in a condition requires the negation to be stated or derived.\n\nOptions:\n- true: The statement follows from the facts and rules.\n- false: The negation of the statement follows from the facts and rules.\n- unknown: Neither the statement nor its negation follows from the facts and rules.\n\nReply with a JSON object {\"answer\": <option>} where <option> is exactly one of: true, false, unknown.",
      "role": "user"
    }
  ],
  "model": "gpt-6-luna",
  "reasoning_effort": "none",
  "response_format": {
    "json_schema": {
      "name": "decision",
      "schema": {
        "additionalProperties": false,
        "properties": {
          "answer": {
            "enum": [
              "true",
              "false",
              "unknown"
            ],
            "type": "string"
          }
        },
        "required": [
          "answer"
        ],
        "type": "object"
      },
      "strict": true
    },
    "type": "json_schema"
  },
  "top_logprobs": 5
}

The response

{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "logprobs": {
        "content": [
          {
            "bytes": [
              123,
              34
            ],
            "logprob": 0,
            "token": "{\"",
            "top_logprobs": [
              {
                "bytes": [
                  123,
                  34
                ],
                "logprob": 0,
                "token": "{\""
              }
            ]
          },
          {
            "bytes": [
              97,
              110,
              115,
              119,
              101,
              114
            ],
            "logprob": 0,
            "token": "answer",
            "top_logprobs": [
              {
                "bytes": [
                  97,
                  110,
                  115,
                  119,
                  101,
                  114
                ],
                "logprob": 0,
                "token": "answer"
              }
            ]
          },
          {
            "bytes": [
              34,
              58,
              34
            ],
            "logprob": 0,
            "token": "\":\"",
            "top_logprobs": [
              {
                "bytes": [
                  34,
                  58,
                  34
                ],
                "logprob": 0,
                "token": "\":\""
              }
            ]
          },
          {
            "bytes": [
              117,
              110,
              107,
              110,
              111,
              119,
              110
            ],
            "logprob": -0.001819610595703125,
            "token": "unknown",
            "top_logprobs": [
              {
                "bytes": [
                  117,
                  110,
                  107,
                  110,
                  111,
                  119,
                  110
                ],
                "logprob": -0.001819610595703125,
                "token": "unknown"
              }
            ]
          },
          {
            "bytes": [
              34,
              125
            ],
            "logprob": 0,
            "token": "\"}",
            "top_logprobs": [
              {
                "bytes": [
                  34,
                  125
                ],
                "logprob": 0,
                "token": "\"}"
              }
            ]
          }
        ],
        "refusal": null
      },
      "message": {
        "annotations": [],
        "audio": null,
        "content": "{\"answer\":\"unknown\"}",
        "function_call": null,
        "refusal": null,
        "role": "assistant",
        "tool_calls": null
      }
    }
  ],
  "created": 1790866630,
  "id": "chatcmpl-EUCVa4ydOIAtLD1d5VSwqwdcWrIyc",
  "model": "gpt-6-luna",
  "moderation": null,
  "object": "chat.completion",
  "service_tier": "default",
  "system_fingerprint": null,
  "usage": {
    "completion_tokens": 11,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens": 340,
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cache_write_tokens": 0,
      "cached_tokens": 0
    },
    "total_tokens": 351
  }
}
A right answer, for comparison (AttNoneg-OWA-D5-924-Q18, stated 100.00%, correct answer unknown)
{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "logprobs": {
        "content": [
          {
            "bytes": [
              123,
              34
            ],
            "logprob": 0,
            "token": "{\"",
            "top_logprobs": [
              {
                "bytes": [
                  123,
                  34
                ],
                "logprob": 0,
                "token": "{\""
              }
            ]
          },
          {
            "bytes": [
              97,
              110,
              115,
              119,
              101,
              114
            ],
            "logprob": 0,
            "token": "answer",
            "top_logprobs": [
              {
                "bytes": [
                  97,
                  110,
                  115,
                  119,
                  101,
                  114
                ],
                "logprob": 0,
                "token": "answer"
              }
            ]
          },
          {
            "bytes": [
              34,
              58,
              34
            ],
            "logprob": 0,
            "token": "\":\"",
            "top_logprobs": [
              {
                "bytes": [
                  34,
                  58,
                  34
                ],
                "logprob": 0,
                "token": "\":\""
              }
            ]
          },
          {
            "bytes": [
              117,
              110,
              107,
              110,
              111,
              119,
              110
            ],
            "logprob": -0.000003814697265625,
            "token": "unknown",
            "top_logprobs": [
              {
                "bytes": [
                  117,
                  110,
                  107,
                  110,
                  111,
                  119,
                  110
                ],
                "logprob": -0.000003814697265625,
                "token": "unknown"
              }
            ]
          },
          {
            "bytes": [
              34,
              125
            ],
            "logprob": -0.000003814697265625,
            "token": "\"}",
            "top_logprobs": [
              {
                "bytes": [
                  34,
                  125
                ],
                "logprob": -0.000003814697265625,
                "token": "\"}"
              }
            ]
          }
        ],
        "refusal": null
      },
      "message": {
        "annotations": [],
        "audio": null,
        "content": "{\"answer\":\"unknown\"}",
        "function_call": null,
        "refusal": null,
        "role": "assistant",
        "tool_calls": null
      }
    }
  ],
  "created": 1790866630,
  "id": "chatcmpl-EUCVa0u5eg0DyNLwbbvk2m0sTjVW2",
  "model": "gpt-6-luna",
  "moderation": null,
  "object": "chat.completion",
  "service_tier": "default",
  "system_fingerprint": null,
  "usage": {
    "completion_tokens": 11,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "prompt_tokens": 320,
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cache_write_tokens": 0,
      "cached_tokens": 0
    },
    "total_tokens": 331
  }
}

Every request and response is kept verbatim in probes/luna-logprobs/.

The spot check that started it

Before writing the probe down, we sent a dozen deep problems to GPT-6 Luna with log-probabilities turned on. From the preregistration, word for word:

A spot check before this amendment (12 open-world items at depth 3 to 5, 2026-10-01) found Luna returns the chosen answer's log-probability and at most one or two alternatives with reasoning off, refuses logprobs with reasoning on, and gave 5 of 12 wrong answers at 97.8% or more. It motivated the probe and is not part of its results.

Twelve items prove nothing, so we preregistered a full probe (Amendment 6): every one of the 3,600 benchmark items, sent with exactly the benchmark's GPT-6 Luna request plus logprobs and top_logprobs: 5, with four predictions written down before it ran. It is outside the scored benchmark and changes no benchmark number.

Results: when Luna says 99%, how often is it right?

Each answer is grouped by the probability the model gave it. A well-calibrated model is right about as often as it says. Jev's probabilities are its own, from the scored run, on the same items.

Open world: true, false or unknown

Stated probabilityLuna answersLuna stated, averageLuna rightJev answersJev right
0% to 50%7432.0%40.5%3151.6%
50% to 80%9467.6%52.1%42557.4%
80% to 90%6086.1%43.3%19778.2%
90% to 95%7692.8%48.7%18883.5%
95% to 99%17997.5%50.8%25293.3%
99% or more1,30299.9%68.7%70799.4%

Closed world: true or false

Stated probabilityLuna answersLuna stated, averageLuna rightJev answersJev right
0% to 50%6529.1%49.2%0—
50% to 80%7468.2%52.7%27961.6%
80% to 90%5086.1%54.0%15778.3%
90% to 95%7092.8%54.3%13786.1%
95% to 99%14997.5%51.7%26792.9%
99% or more1,37099.9%68.1%96098.5%

Every model, same items

Expected calibration error (ECE) is the average gap between stated probability and accuracy; 0 is perfect. "Wrong at 95%" is the share of wrong answers stated at 95% or more. AUROC is how well the stated probability separates right answers from wrong ones: 1 is perfect and 0.5 is no better than a coin flip.

ModelTaskAccuracy, this runECEWrong at 95%+Wrong at 99%+AUROC
GPT-6 LunaOpen world63.3%0.31675.5%62.1%0.662
JevOpen world83.8%0.0417.2%1.4%0.849
Kev-0.8BOpen world53.6%0.1140.0%0.0%0.659
Kev-4BOpen world53.6%0.1640.5%0.0%0.711
LayaOpen world42.4%0.3236.2%0.1%0.612
GPT-6 LunaClosed world64.8%0.31880.6%69.2%0.679
JevClosed world89.3%0.02817.1%7.3%0.858
Kev-0.8BClosed world55.8%0.1701.9%0.0%0.573
Kev-4BClosed world58.8%0.0770.0%0.0%0.591
LayaClosed world55.6%0.31720.8%0.5%0.560

Luna's accuracy here is from the probe, a separate run at OpenAI's default sampling; the scored benchmark's figures (64.1% and 65.0%) come from the earlier run. In 137 of 3,600 responses (72 open world and 65 closed world), the answer returned was not Luna's most probable listed option: that is sampling at the default temperature.

Separating right from wrong, by proof depth (AUROC)

TaskDepthGPT-6 LunaJevKev-0.8BKev-4BLaya
Open world00.8610.9530.7150.8500.785
10.6820.8530.5840.7260.602
20.6280.8450.6470.6430.537
30.5570.8520.6650.6420.597
40.5200.7270.6380.5580.478
50.5210.8280.6850.6080.535
Closed world00.9420.9780.6610.8330.673
10.7510.8720.5360.6280.472
20.6380.8260.5480.5860.524
30.5630.7980.5360.4860.561
40.5560.7960.5210.5050.547
50.5020.8510.5970.4490.549

At proof depth 5, Luna's AUROC is 0.521 on the open-world task and 0.502 on the closed-world task: its probability says next to nothing about whether a deep answer is right. Jev's is 0.828 and 0.851. Jev's figures by depth are computed by this site from its saved probabilities, the same way as the analysis computes Luna's.

What we predicted

  1. HeldPrediction 1

    Luna's ECE exceeds 0.15 on both tasks.

    • Rule: Luna's expected calibration error is above 0.15 on each task.
    • OWA: ECE 0.316.
    • CWA: ECE 0.318.
  2. HeldPrediction 2

    At least half of Luna's wrong answers are stated at 95% or more, on both tasks.

    • Rule: at least half of Luna's wrong answers are stated at 95% or more, on each task.
    • OWA: 75.5% of 660 wrong answers.
    • CWA: 80.6% of 633 wrong answers.
  3. HeldPrediction 3

    Luna's AUROC is lower than Jev's on both tasks.

    • Rule: Luna's AUROC is below Jev's on each task, on the same items.
    • OWA: Luna 0.662, Jev 0.849.
    • CWA: Luna 0.679, Jev 0.858.
  4. HeldPrediction 4

    On most responses, Luna discloses fewer than all of the task's options.

    • Rule: on more than half of the responses, fewer than all of the task's options are disclosed.
    • OWA: 1,764 of 1,800 responses disclose fewer than all 3 options.
    • CWA: 1,455 of 1,800 responses disclose fewer than all 2 options.

What it means for a decision model built on Luna

A confidence is useful for routing: act on the sure answers, review the unsure ones. On these problems, Luna's stated confidence cannot do that without help.

Overconfidence can be calibrated. If a model's 99% answers are right 68.7% of the time, you can learn that from labelled data and relabel them. That fixes what the number says, as long as higher stated probability still means more likely right.

Calibration cannot add information. AUROC measures only the ordering: whether right answers tend to get higher probability than wrong ones. Any recalibration that keeps that order leaves AUROC unchanged. Where AUROC is near 0.5, as Luna's is at depth 5, right and wrong answers get the same spread of probabilities, so no remapping of that probability can pick out which deep answers to trust. You would need a different signal.

So a Luna-based decision model may be accurate enough on shallow problems, but on decisions that need several steps of reasoning, its own probability is not a usable confidence. Measure this on your own decisions before relying on it; the API's specialized model may behave differently.

Context: overconfidence and access to log-probabilities

Limits