Breakdown · descriptive

Negation: in the rules and in the statement

Two kinds of "not": in the theory's facts and rules, and in the statement to judge. The preregistration predicted that negation in the rules would lower accuracy.

Descriptive, not causal: the sample controls proof depth and the answer, not these properties. Several of them grow with depth and with each other, so each section gives the average proof depth of its slices; the answer mix and mean depth of every slice are in the tables.

Does negation in the rules lower accuracy?

Whether the theory's facts and rules contain "not". The preregistration predicted that negation would lower accuracy.

On the open-world task, Jev's accuracy ranges from 80.3% (without negation) to 87.5% (with negation), a spread of 7.2 points.

Descriptive, not causal. The average proof depth of these slices runs from 2.5 to 2.5, so depth differs between them and can add to or mask any gap.

Negation in the theory, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
with negation881 items
without negation919 items
Negation in the theory, Open world
Negation in the theoryItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
with negation88187.5% 85.4–89.767.1% 63.9–70.360.4% 57.1–63.857.3% 54.0–60.756.6% 53.6–59.646.2% 43.0–49.535.1%false 35% · true 35% · unknown 30% · 2.5
without negation91980.3% 77.7–83.161.2% 58.1–64.156.8% 53.5–59.849.9% 46.9–53.150.6% 47.3–54.138.8% 35.5–41.935.5%false 32% · true 32% · unknown 35% · 2.5

Negation in the theory, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
with negation870 items
without negation930 items
Negation in the theory, Closed world
Negation in the theoryItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
with negation87093.7% 92.1–95.368.6% 65.4–71.866.4% 63.3–69.761.4% 57.9–64.657.4% 54.1–60.757.4% 54.1–60.550.1%false 50% · true 50% · 2.5
without negation93085.2% 82.8–87.561.6% 58.6–64.761.3% 58.3–64.656.3% 53.0–59.254.3% 51.0–57.454.0% 50.5–57.350.1%false 50% · true 50% · 2.5

Accuracy on negated statements

Whether the statement to judge itself contains "not" ("The lion is not blue").

On the open-world task, Jev's accuracy ranges from 78.6% (negated statement) to 89.1% (plain statement), a spread of 10.5 points.

Descriptive, not causal. The average proof depth of these slices runs from 2.5 to 2.5, so depth differs between them and can add to or mask any gap.

Negated statement, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
plain statement899 items
negated statement901 items
Negated statement, Open world
Negated statementItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
plain statement89989.1% 87.1–91.062.3% 59.2–65.359.7% 56.4–63.053.6% 50.7–56.648.5% 45.4–51.738.7% 35.5–42.041.5%false 26% · true 41% · unknown 32% · 2.5
negated statement90178.6% 75.9–81.265.8% 62.9–68.857.4% 53.8–60.453.5% 50.2–56.858.6% 55.5–61.746.2% 43.1–49.541.1%false 41% · true 26% · unknown 33% · 2.5

Negated statement, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
plain statement882 items
negated statement918 items
Negated statement, Closed world
Negated statementItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
plain statement88289.6% 87.5–91.667.5% 64.4–70.565.6% 62.6–68.758.6% 55.3–61.753.6% 50.3–56.956.7% 53.5–59.854.8%false 45% · true 55% · 2.5
negated statement91889.0% 87.0–91.062.6% 59.4–65.862.0% 58.9–65.058.9% 55.7–62.257.8% 54.7–60.954.6% 51.3–57.754.6%false 55% · true 45% · 2.5

In each chart a line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance; the grey tick is the score of always giving the slice's most common answer.