Decision model consistency: does the same model give the same answer twice?
A decision you cannot repeat is hard to audit. We asked each model every question a second time and counted how many answers changed; for models that report probabilities, we also checked how well those probabilities match how often they are right.
97.2%of Jev's open-world answers were the same when asked twice: 1,750 of 1,800, with 50 changed. Gwet's AC1, which corrects for agreement by chance, is 0.958 (interval 0.948 to 0.969).
Test-retest agreement
The second run is a full rerun of every item under the same protocol, one request at a time. Agreement is the share of items with the same answer both times. AC1 corrects for chance; Cohen's kappa is shown beside it but understates agreement when a model gives one answer to most items.
Probability shift is how far the probability of run 1's answer moved in run 2. Reruns not yet scored: Kev-9B.
Where Jev's answers changed, by proof depth
Task
depth 0
depth 1
depth 2
depth 3
depth 4
depth 5
Open world
0.0% (0)
1.0% (3)
2.3% (7)
4.3% (13)
5.6% (17)
3.5% (10)
Closed world
0.0% (0)
1.0% (3)
1.7% (5)
2.0% (6)
4.0% (12)
4.0% (12)
Where GPT-6 Luna's answers changed (reasoning off), by proof depth
Task
depth 0
depth 1
depth 2
depth 3
depth 4
depth 5
Open world
3.7% (11)
6.6% (20)
12.9% (39)
12.9% (39)
12.2% (37)
13.5% (39)
Closed world
0.7% (2)
7.3% (22)
11.7% (35)
12.0% (36)
11.3% (34)
13.3% (40)
Where Kev-4B's answers changed, by proof depth
Task
depth 0
depth 1
depth 2
depth 3
depth 4
depth 5
Open world
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
Closed world
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
Where Kev-0.8B's answers changed, by proof depth
Task
depth 0
depth 1
depth 2
depth 3
depth 4
depth 5
Open world
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
Closed world
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
Where Laya's answers changed, by proof depth
Task
depth 0
depth 1
depth 2
depth 3
depth 4
depth 5
Open world
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
Closed world
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
0.0% (0)
So far, changes concentrate on deeper proofs: the problems a model finds hardest are also the ones it answers least consistently.
What we predicted
Pending “Kev and Laya agree with themselves on at least 99.5% of items on both tasks (local, fixed weights).” Kev-0.8B on OWA: 100.0% agreement. Kev-0.8B on CWA: 100.0% agreement. Kev-4B on OWA: 100.0% agreement. Kev-4B on CWA: 100.0% agreement. Laya on OWA: 100.0% agreement. Laya on CWA: 100.0% agreement. No scored rerun yet for: Kev-9B on OWA, Kev-9B on CWA.
Held “Jev's AC1 is at least 0.90 on both tasks (observed: 0.958 on both, before this amendment).” OWA: AC1 0.958 [0.948–0.969], agreement 97.2%. CWA: AC1 0.958 [0.944–0.971], agreement 97.9%.
Held “Where answers change, they change more at depth 3 and above than at depth 0 to 2, for every engine with at least 20 changes.” Jev on OWA: 4.5% of answers changed at depths 3 to 5 (40 of 895) against 1.1% at depths 0 to 2 (10 of 905). Kev-0.8B on OWA: 0 changes, under the 20 the rule needs. Kev-4B on OWA: 0 changes, under the 20 the rule needs. Laya on OWA: 0 changes, under the 20 the rule needs. GPT-6 Luna on OWA: 12.8% of answers changed at depths 3 to 5 (115 of 895) against 7.7% at depths 0 to 2 (70 of 905). Jev on CWA: 3.3% of answers changed at depths 3 to 5 (30 of 900) against 0.9% at depths 0 to 2 (8 of 900). Kev-0.8B on CWA: 0 changes, under the 20 the rule needs. Kev-4B on CWA: 0 changes, under the 20 the rule needs. Laya on CWA: 0 changes, under the 20 the rule needs. GPT-6 Luna on CWA: 12.2% of answers changed at depths 3 to 5 (110 of 900) against 6.6% at depths 0 to 2 (59 of 900).
Calibration: do the probabilities mean what they say?
Several models return a probability with each answer. A well-calibrated model that says 80% is right about 80% of the time. We grouped each model's answers by the probability it gave its own answer and counted how often it was right.
Not preregistered and not part of the harness's report: computed by this site from the saved answers, for fully scored models only. Expected calibration error (ECE) is the average gap between stated probability and accuracy, weighted by how many answers fall in each group.
Jev, open world: stated 87.8%, right 83.8%, ECE 4.1 points
Stated probability
Answers
Average stated
Actually right
Gap
0–50%
31
45.3%
51.6%
6.4
50–60%
127
54.3%
50.4%
-3.9
60–70%
146
64.6%
57.5%
-7.1
70–80%
152
74.6%
63.2%
-11.5
80–90%
197
84.7%
78.2%
-6.6
90–100%
1,147
97.8%
95.5%
-2.3
Jev, closed world: stated 92.1%, right 89.3%, ECE 2.9 points
Stated probability
Answers
Average stated
Actually right
Gap
50–60%
80
54.9%
46.3%
-8.6
60–70%
95
64.7%
60.0%
-4.7
70–80%
104
74.6%
75.0%
0.4
80–90%
157
84.4%
78.3%
-6.0
90–100%
1,364
98.4%
96.2%
-2.2
Kev-9B, open world: stated 74.3%, right 58.6%, ECE 15.8 points
Stated probability
Answers
Average stated
Actually right
Gap
0–50%
206
45.4%
37.4%
-8.0
50–60%
276
55.1%
38.0%
-17.1
60–70%
264
65.1%
47.0%
-18.1
70–80%
279
75.1%
51.3%
-23.8
80–90%
295
85.3%
61.7%
-23.6
90–100%
480
95.8%
88.1%
-7.7
Kev-9B, closed world: stated 77.3%, right 63.8%, ECE 13.5 points
Stated probability
Answers
Average stated
Actually right
Gap
50–60%
270
54.9%
49.6%
-5.3
60–70%
310
65.1%
56.1%
-9.0
70–80%
376
75.2%
57.7%
-17.5
80–90%
439
85.2%
68.8%
-16.4
90–100%
405
94.9%
79.3%
-15.7
Kev-4B, open world: stated 69.9%, right 53.6%, ECE 16.4 points
Stated probability
Answers
Average stated
Actually right
Gap
0–50%
221
46.3%
33.5%
-12.8
50–60%
358
54.8%
39.1%
-15.7
60–70%
361
65.0%
40.4%
-24.6
70–80%
306
75.0%
55.6%
-19.5
80–90%
280
84.6%
64.3%
-20.3
90–100%
274
94.5%
92.7%
-1.8
Kev-4B, closed world: stated 66.4%, right 58.8%, ECE 7.6 points
Stated probability
Answers
Average stated
Actually right
Gap
50–60%
704
55.0%
50.4%
-4.6
60–70%
444
64.8%
59.5%
-5.4
70–80%
360
74.6%
63.6%
-11.0
80–90%
235
84.5%
69.4%
-15.2
90–100%
57
92.5%
82.5%
-10.0
Kev-0.8B, open world: stated 65.0%, right 53.6%, ECE 11.4 points
Stated probability
Answers
Average stated
Actually right
Gap
0–50%
441
42.5%
36.5%
-6.0
50–60%
269
55.0%
42.4%
-12.6
60–70%
309
65.2%
53.1%
-12.1
70–80%
325
74.6%
60.3%
-14.3
80–90%
387
84.7%
72.6%
-12.1
90–100%
69
91.5%
69.6%
-21.9
Kev-0.8B, closed world: stated 72.8%, right 55.8%, ECE 17.0 points
Stated probability
Answers
Average stated
Actually right
Gap
50–60%
374
55.1%
50.0%
-5.1
60–70%
413
65.2%
47.0%
-18.3
70–80%
434
75.0%
58.5%
-16.5
80–90%
363
84.9%
64.5%
-20.4
90–100%
216
93.3%
62.5%
-30.8
Laya, open world: stated 74.7%, right 42.4%, ECE 32.3 points
Stated probability
Answers
Average stated
Actually right
Gap
0–50%
217
43.6%
34.6%
-9.1
50–60%
180
55.5%
28.3%
-27.1
60–70%
257
65.3%
34.6%
-30.6
70–80%
315
75.3%
40.6%
-34.7
80–90%
418
85.0%
42.6%
-42.5
90–100%
413
94.5%
58.8%
-35.7
Laya, closed world: stated 87.3%, right 55.6%, ECE 31.7 points
Stated probability
Answers
Average stated
Actually right
Gap
50–60%
99
55.3%
55.6%
0.2
60–70%
94
65.1%
53.2%
-11.9
70–80%
168
75.8%
53.6%
-22.2
80–90%
384
85.8%
50.3%
-35.5
90–100%
1,055
94.7%
58.1%
-36.5
GPT-6 Luna (reasoning off): we asked for the answer only, as for every model, so no confidence was recorded in the scored run. A separate preregistered probe read GPT-6 Luna's log-probabilities: see its calibration on the OpenAI Decisions API preview page.