Breakdown · descriptive
Negation: in the rules and in the statement
Two kinds of "not": in the theory's facts and rules, and in the statement to judge. The preregistration predicted that negation in the rules would lower accuracy.
Descriptive, not causal: the sample controls proof depth and the answer, not these properties. Several of them grow with depth and with each other, so each section gives the average proof depth of its slices; the answer mix and mean depth of every slice are in the tables.
Does negation in the rules lower accuracy?
Whether the theory's facts and rules contain "not". The preregistration predicted that negation would lower accuracy.
On the open-world task, Jev's accuracy ranges from 80.3% (without negation) to 87.5% (with negation), a spread of 7.2 points.
Descriptive, not causal. The average proof depth of these slices runs from 2.5 to 2.5, so depth differs between them and can add to or mask any gap.
Negation in the theory, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Negation in the theory | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| with negation | 881 | 87.5% 85.4–89.7 | 67.1% 63.9–70.3 | 60.4% 57.1–63.8 | 57.3% 54.0–60.7 | 56.6% 53.6–59.6 | 46.2% 43.0–49.5 | 35.1% | false 35% · true 35% · unknown 30% · 2.5 |
| without negation | 919 | 80.3% 77.7–83.1 | 61.2% 58.1–64.1 | 56.8% 53.5–59.8 | 49.9% 46.9–53.1 | 50.6% 47.3–54.1 | 38.8% 35.5–41.9 | 35.5% | false 32% · true 32% · unknown 35% · 2.5 |
Negation in the theory, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Negation in the theory | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| with negation | 870 | 93.7% 92.1–95.3 | 68.6% 65.4–71.8 | 66.4% 63.3–69.7 | 61.4% 57.9–64.6 | 57.4% 54.1–60.7 | 57.4% 54.1–60.5 | 50.1% | false 50% · true 50% · 2.5 |
| without negation | 930 | 85.2% 82.8–87.5 | 61.6% 58.6–64.7 | 61.3% 58.3–64.6 | 56.3% 53.0–59.2 | 54.3% 51.0–57.4 | 54.0% 50.5–57.3 | 50.1% | false 50% · true 50% · 2.5 |
Accuracy on negated statements
Whether the statement to judge itself contains "not" ("The lion is not blue").
On the open-world task, Jev's accuracy ranges from 78.6% (negated statement) to 89.1% (plain statement), a spread of 10.5 points.
Descriptive, not causal. The average proof depth of these slices runs from 2.5 to 2.5, so depth differs between them and can add to or mask any gap.
Negated statement, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Negated statement | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| plain statement | 899 | 89.1% 87.1–91.0 | 62.3% 59.2–65.3 | 59.7% 56.4–63.0 | 53.6% 50.7–56.6 | 48.5% 45.4–51.7 | 38.7% 35.5–42.0 | 41.5% | false 26% · true 41% · unknown 32% · 2.5 |
| negated statement | 901 | 78.6% 75.9–81.2 | 65.8% 62.9–68.8 | 57.4% 53.8–60.4 | 53.5% 50.2–56.8 | 58.6% 55.5–61.7 | 46.2% 43.1–49.5 | 41.1% | false 41% · true 26% · unknown 33% · 2.5 |
Negated statement, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Negated statement | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| plain statement | 882 | 89.6% 87.5–91.6 | 67.5% 64.4–70.5 | 65.6% 62.6–68.7 | 58.6% 55.3–61.7 | 53.6% 50.3–56.9 | 56.7% 53.5–59.8 | 54.8% | false 45% · true 55% · 2.5 |
| negated statement | 918 | 89.0% 87.0–91.0 | 62.6% 59.4–65.8 | 62.0% 58.9–65.0 | 58.9% 55.7–62.2 | 57.8% 54.7–60.9 | 54.6% 51.3–57.7 | 54.6% | false 55% · true 45% · 2.5 |
In each chart a line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance; the grey tick is the score of always giving the slice's most common answer.