Breakdown · descriptive

Accuracy on true, false and unknown answers

The correct answer. On the open-world task a third of the items are unknown: neither the statement nor its negation follows. That class is where several models fail.

14.5On the open-world task, Jev's accuracy ranges from 74.2% (unknown) to 88.8% (true), a spread of 14.5 points. Jev is the most accurate model overall; every model is below.

Descriptive, not causal: the sample controls proof depth and the answer, not gold answer. Across these slices the average proof depth runs from 2.4 to 2.5, so depth differs between slices and can add to or mask any gap. The answer mix and mean depth of every slice are in the tables.

Gold answer, open world task

A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (33%); the grey tick is the score of always giving the slice's most common answer.

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
false605 items
true605 items
unknown590 items
Gold answer, Open world
Gold answerItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
false60588.3% 85.5–90.950.9% 46.8–54.948.6% 44.6–52.635.9% 31.9–39.589.6% 87.1–91.785.1% 82.1–87.9one answerfalse 100% · 2.5
true60588.8% 86.1–91.158.3% 54.0–62.159.2% 55.4–63.146.6% 42.8–50.658.2% 54.0–62.040.2% 36.2–44.1one answertrue 100% · 2.5
unknown59074.2% 70.7–77.683.4% 80.5–86.668.1% 64.6–71.778.8% 75.4–81.911.9% 9.3–14.61.0% 0.3–1.9one answerunknown 100% · 2.4

Gold answer, closed world task

Anything that cannot be derived is false, so there is no unknown answer. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (50%); the grey tick is the score of always giving the slice's most common answer.

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
false900 items
true900 items
Gold answer, Closed world
Gold answerItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
false90088.9% 86.9–90.963.8% 60.6–66.962.9% 59.9–65.975.3% 72.8–78.266.1% 63.0–69.386.4% 84.1–88.6one answerfalse 100% · 2.5
true90089.7% 87.7–91.666.2% 63.1–69.264.7% 61.7–67.642.2% 39.2–45.445.4% 42.1–48.624.8% 21.9–27.6one answertrue 100% · 2.5

How often each model says unknown

On the open-world task, about a third of the correct answers are unknown. A model that rarely says unknown cannot get those items right.