Repeatability and calibration

Decision model consistency: does the same model give the same answer twice?

A decision you cannot repeat is hard to audit. We asked each model every question a second time and counted how many answers changed; for models that report probabilities, we also checked how well those probabilities match how often they are right.

97.2%of Jev's open-world answers were the same when asked twice: 1,750 of 1,800, with 50 changed. Gwet's AC1, which corrects for agreement by chance, is 0.958 (interval 0.948 to 0.969).

Test-retest agreement

The second run is a full rerun of every item under the same protocol, one request at a time. Agreement is the share of items with the same answer both times. AC1 corrects for chance; Cohen's kappa is shown beside it but understates agreement when a model gives one answer to most items.

ModelTaskAgreementAC1 (95%)KappaChangedAccuracy, run 1 / run 2Probability shift, mean / max
JevOpen world97.2%0.958 0.948–0.9690.95850 of 1,80083.8% / 84.2%0.023 / 0.36
JevClosed world97.9%0.958 0.944–0.9710.95838 of 1,80089.3% / 89.7%0.020 / 0.90
GPT-6 LunaOpen world89.7%0.855 0.837–0.8760.822185 of 1,80064.1% / 62.8%— / —
GPT-6 LunaClosed world90.6%0.812 0.784–0.8380.812169 of 1,80065.0% / 64.8%— / —
Kev-4BOpen world100.0%1.000 1.000–1.0001.0000 of 1,80053.6% / 53.6%0.000 / 0.00
Kev-4BClosed world100.0%1.000 1.000–1.0001.0000 of 1,80058.8% / 58.8%0.000 / 0.00
Kev-0.8BOpen world100.0%1.000 1.000–1.0001.0000 of 1,80053.6% / 53.6%0.000 / 0.00
Kev-0.8BClosed world100.0%1.000 1.000–1.0001.0000 of 1,80055.8% / 55.8%0.000 / 0.00
LayaOpen world100.0%1.000 1.000–1.0001.0000 of 1,80042.4% / 42.4%0.000 / 0.00
LayaClosed world100.0%1.000 1.000–1.0001.0000 of 1,80055.6% / 55.6%0.000 / 0.00

Probability shift is how far the probability of run 1's answer moved in run 2. Reruns not yet scored: Kev-9B.

Where Jev's answers changed, by proof depth

Taskdepth 0depth 1depth 2depth 3depth 4depth 5
Open world0.0% (0)1.0% (3)2.3% (7)4.3% (13)5.6% (17)3.5% (10)
Closed world0.0% (0)1.0% (3)1.7% (5)2.0% (6)4.0% (12)4.0% (12)

Where GPT-6 Luna's answers changed (reasoning off), by proof depth

Taskdepth 0depth 1depth 2depth 3depth 4depth 5
Open world3.7% (11)6.6% (20)12.9% (39)12.9% (39)12.2% (37)13.5% (39)
Closed world0.7% (2)7.3% (22)11.7% (35)12.0% (36)11.3% (34)13.3% (40)

Where Kev-4B's answers changed, by proof depth

Taskdepth 0depth 1depth 2depth 3depth 4depth 5
Open world0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)
Closed world0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)

Where Kev-0.8B's answers changed, by proof depth

Taskdepth 0depth 1depth 2depth 3depth 4depth 5
Open world0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)
Closed world0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)

Where Laya's answers changed, by proof depth

Taskdepth 0depth 1depth 2depth 3depth 4depth 5
Open world0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)
Closed world0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)0.0% (0)

So far, changes concentrate on deeper proofs: the problems a model finds hardest are also the ones it answers least consistently.

What we predicted

Calibration: do the probabilities mean what they say?

Several models return a probability with each answer. A well-calibrated model that says 80% is right about 80% of the time. We grouped each model's answers by the probability it gave its own answer and counted how often it was right.

Not preregistered and not part of the harness's report: computed by this site from the saved answers, for fully scored models only. Expected calibration error (ECE) is the average gap between stated probability and accuracy, weighted by how many answers fall in each group.

Jev, open world: stated 87.8%, right 83.8%, ECE 4.1 points

Stated probabilityAnswersAverage statedActually rightGap
0–50%3145.3%51.6%6.4
50–60%12754.3%50.4%-3.9
60–70%14664.6%57.5%-7.1
70–80%15274.6%63.2%-11.5
80–90%19784.7%78.2%-6.6
90–100%1,14797.8%95.5%-2.3

Jev, closed world: stated 92.1%, right 89.3%, ECE 2.9 points

Stated probabilityAnswersAverage statedActually rightGap
50–60%8054.9%46.3%-8.6
60–70%9564.7%60.0%-4.7
70–80%10474.6%75.0%0.4
80–90%15784.4%78.3%-6.0
90–100%1,36498.4%96.2%-2.2

Kev-9B, open world: stated 74.3%, right 58.6%, ECE 15.8 points

Stated probabilityAnswersAverage statedActually rightGap
0–50%20645.4%37.4%-8.0
50–60%27655.1%38.0%-17.1
60–70%26465.1%47.0%-18.1
70–80%27975.1%51.3%-23.8
80–90%29585.3%61.7%-23.6
90–100%48095.8%88.1%-7.7

Kev-9B, closed world: stated 77.3%, right 63.8%, ECE 13.5 points

Stated probabilityAnswersAverage statedActually rightGap
50–60%27054.9%49.6%-5.3
60–70%31065.1%56.1%-9.0
70–80%37675.2%57.7%-17.5
80–90%43985.2%68.8%-16.4
90–100%40594.9%79.3%-15.7

Kev-4B, open world: stated 69.9%, right 53.6%, ECE 16.4 points

Stated probabilityAnswersAverage statedActually rightGap
0–50%22146.3%33.5%-12.8
50–60%35854.8%39.1%-15.7
60–70%36165.0%40.4%-24.6
70–80%30675.0%55.6%-19.5
80–90%28084.6%64.3%-20.3
90–100%27494.5%92.7%-1.8

Kev-4B, closed world: stated 66.4%, right 58.8%, ECE 7.6 points

Stated probabilityAnswersAverage statedActually rightGap
50–60%70455.0%50.4%-4.6
60–70%44464.8%59.5%-5.4
70–80%36074.6%63.6%-11.0
80–90%23584.5%69.4%-15.2
90–100%5792.5%82.5%-10.0

Kev-0.8B, open world: stated 65.0%, right 53.6%, ECE 11.4 points

Stated probabilityAnswersAverage statedActually rightGap
0–50%44142.5%36.5%-6.0
50–60%26955.0%42.4%-12.6
60–70%30965.2%53.1%-12.1
70–80%32574.6%60.3%-14.3
80–90%38784.7%72.6%-12.1
90–100%6991.5%69.6%-21.9

Kev-0.8B, closed world: stated 72.8%, right 55.8%, ECE 17.0 points

Stated probabilityAnswersAverage statedActually rightGap
50–60%37455.1%50.0%-5.1
60–70%41365.2%47.0%-18.3
70–80%43475.0%58.5%-16.5
80–90%36384.9%64.5%-20.4
90–100%21693.3%62.5%-30.8

Laya, open world: stated 74.7%, right 42.4%, ECE 32.3 points

Stated probabilityAnswersAverage statedActually rightGap
0–50%21743.6%34.6%-9.1
50–60%18055.5%28.3%-27.1
60–70%25765.3%34.6%-30.6
70–80%31575.3%40.6%-34.7
80–90%41885.0%42.6%-42.5
90–100%41394.5%58.8%-35.7

Laya, closed world: stated 87.3%, right 55.6%, ECE 31.7 points

Stated probabilityAnswersAverage statedActually rightGap
50–60%9955.3%55.6%0.2
60–70%9465.1%53.2%-11.9
70–80%16875.8%53.6%-22.2
80–90%38485.8%50.3%-35.5
90–100%1,05594.7%58.1%-36.5

GPT-6 Luna (reasoning off): we asked for the answer only, as for every model, so no confidence was recorded in the scored run. A separate preregistered probe read GPT-6 Luna's log-probabilities: see its calibration on the OpenAI Decisions API preview page.