A real request and response, verbatim
One of the probe's 3,600 requests, exactly as sent and as returned: a open-world problem at proof depth 3 (AttNonegNatLang-OWA-479-Q8). The correct answer is false. Luna answered unknown, and the log-probability of that answer's token, -0.0018, is a probability of 99.82%. It listed no other option at that position.
That is a confidently wrong answer: stated as close to certain, and wrong. On its own it proves nothing; the results below show how often it happens.
The request body
{
"logprobs": true,
"max_completion_tokens": 1024,
"messages": [
{
"content": "Alan is young, round, and kind, but that doesn't mean he isn't also rough and cold at times, as well. That guy Fred sure is nice. That guy Harry sure is nice. Young round people who are green are usually blue. When green, young and round fits a person, you'll see that rough will also fit. It's a certainty that any green, big and kind individual is going to be nice. Kind people with rough skin are usually red because it's wind burn. If someone shows that they are red, then they are also showing that they are green. Is a kind person if feeling cold and looking red then they are usually young.\n\nStatement: Alan is not blue.\n\nUsing only the facts and rules in the text, is the statement true, false, or unknown? A statement is true if it can be derived from the facts and rules, and false if its negation can be derived. If neither the statement nor its negation can be derived, it is unknown. A rule applies only when all of its conditions are established; 'not' in a condition requires the negation to be stated or derived.\n\nOptions:\n- true: The statement follows from the facts and rules.\n- false: The negation of the statement follows from the facts and rules.\n- unknown: Neither the statement nor its negation follows from the facts and rules.\n\nReply with a JSON object {\"answer\": <option>} where <option> is exactly one of: true, false, unknown.",
"role": "user"
}
],
"model": "gpt-6-luna",
"reasoning_effort": "none",
"response_format": {
"json_schema": {
"name": "decision",
"schema": {
"additionalProperties": false,
"properties": {
"answer": {
"enum": [
"true",
"false",
"unknown"
],
"type": "string"
}
},
"required": [
"answer"
],
"type": "object"
},
"strict": true
},
"type": "json_schema"
},
"top_logprobs": 5
}
The response
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"logprobs": {
"content": [
{
"bytes": [
123,
34
],
"logprob": 0,
"token": "{\"",
"top_logprobs": [
{
"bytes": [
123,
34
],
"logprob": 0,
"token": "{\""
}
]
},
{
"bytes": [
97,
110,
115,
119,
101,
114
],
"logprob": 0,
"token": "answer",
"top_logprobs": [
{
"bytes": [
97,
110,
115,
119,
101,
114
],
"logprob": 0,
"token": "answer"
}
]
},
{
"bytes": [
34,
58,
34
],
"logprob": 0,
"token": "\":\"",
"top_logprobs": [
{
"bytes": [
34,
58,
34
],
"logprob": 0,
"token": "\":\""
}
]
},
{
"bytes": [
117,
110,
107,
110,
111,
119,
110
],
"logprob": -0.001819610595703125,
"token": "unknown",
"top_logprobs": [
{
"bytes": [
117,
110,
107,
110,
111,
119,
110
],
"logprob": -0.001819610595703125,
"token": "unknown"
}
]
},
{
"bytes": [
34,
125
],
"logprob": 0,
"token": "\"}",
"top_logprobs": [
{
"bytes": [
34,
125
],
"logprob": 0,
"token": "\"}"
}
]
}
],
"refusal": null
},
"message": {
"annotations": [],
"audio": null,
"content": "{\"answer\":\"unknown\"}",
"function_call": null,
"refusal": null,
"role": "assistant",
"tool_calls": null
}
}
],
"created": 1790866630,
"id": "chatcmpl-EUCVa4ydOIAtLD1d5VSwqwdcWrIyc",
"model": "gpt-6-luna",
"moderation": null,
"object": "chat.completion",
"service_tier": "default",
"system_fingerprint": null,
"usage": {
"completion_tokens": 11,
"completion_tokens_details": {
"accepted_prediction_tokens": 0,
"audio_tokens": 0,
"reasoning_tokens": 0,
"rejected_prediction_tokens": 0
},
"prompt_tokens": 340,
"prompt_tokens_details": {
"audio_tokens": 0,
"cache_write_tokens": 0,
"cached_tokens": 0
},
"total_tokens": 351
}
}
A right answer, for comparison (AttNoneg-OWA-D5-924-Q18, stated 100.00%, correct answer unknown)
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"logprobs": {
"content": [
{
"bytes": [
123,
34
],
"logprob": 0,
"token": "{\"",
"top_logprobs": [
{
"bytes": [
123,
34
],
"logprob": 0,
"token": "{\""
}
]
},
{
"bytes": [
97,
110,
115,
119,
101,
114
],
"logprob": 0,
"token": "answer",
"top_logprobs": [
{
"bytes": [
97,
110,
115,
119,
101,
114
],
"logprob": 0,
"token": "answer"
}
]
},
{
"bytes": [
34,
58,
34
],
"logprob": 0,
"token": "\":\"",
"top_logprobs": [
{
"bytes": [
34,
58,
34
],
"logprob": 0,
"token": "\":\""
}
]
},
{
"bytes": [
117,
110,
107,
110,
111,
119,
110
],
"logprob": -0.000003814697265625,
"token": "unknown",
"top_logprobs": [
{
"bytes": [
117,
110,
107,
110,
111,
119,
110
],
"logprob": -0.000003814697265625,
"token": "unknown"
}
]
},
{
"bytes": [
34,
125
],
"logprob": -0.000003814697265625,
"token": "\"}",
"top_logprobs": [
{
"bytes": [
34,
125
],
"logprob": -0.000003814697265625,
"token": "\"}"
}
]
}
],
"refusal": null
},
"message": {
"annotations": [],
"audio": null,
"content": "{\"answer\":\"unknown\"}",
"function_call": null,
"refusal": null,
"role": "assistant",
"tool_calls": null
}
}
],
"created": 1790866630,
"id": "chatcmpl-EUCVa0u5eg0DyNLwbbvk2m0sTjVW2",
"model": "gpt-6-luna",
"moderation": null,
"object": "chat.completion",
"service_tier": "default",
"system_fingerprint": null,
"usage": {
"completion_tokens": 11,
"completion_tokens_details": {
"accepted_prediction_tokens": 0,
"audio_tokens": 0,
"reasoning_tokens": 0,
"rejected_prediction_tokens": 0
},
"prompt_tokens": 320,
"prompt_tokens_details": {
"audio_tokens": 0,
"cache_write_tokens": 0,
"cached_tokens": 0
},
"total_tokens": 331
}
}
Every request and response is kept verbatim in probes/luna-logprobs/.