Breakdown · descriptive
Accuracy on true, false and unknown answers
The correct answer. On the open-world task a third of the items are unknown: neither the statement nor its negation follows. That class is where several models fail.
14.5On the open-world task, Jev's accuracy ranges from 74.2% (unknown) to 88.8% (true), a spread of 14.5 points. Jev is the most accurate model overall; every model is below.
Descriptive, not causal: the sample controls proof depth and the answer, not gold answer. Across these slices the average proof depth runs from 2.4 to 2.5, so depth differs between slices and can add to or mask any gap. The answer mix and mean depth of every slice are in the tables.
Gold answer, open world task
A statement is true if it follows from the facts and rules, false if its negation follows, and unknown if neither does. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (33%); the grey tick is the score of always giving the slice's most common answer.
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Gold answer | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| false | 605 | 88.3% 85.5–90.9 | 50.9% 46.8–54.9 | 48.6% 44.6–52.6 | 35.9% 31.9–39.5 | 89.6% 87.1–91.7 | 85.1% 82.1–87.9 | one answer | false 100% · 2.5 |
| true | 605 | 88.8% 86.1–91.1 | 58.3% 54.0–62.1 | 59.2% 55.4–63.1 | 46.6% 42.8–50.6 | 58.2% 54.0–62.0 | 40.2% 36.2–44.1 | one answer | true 100% · 2.5 |
| unknown | 590 | 74.2% 70.7–77.6 | 83.4% 80.5–86.6 | 68.1% 64.6–71.7 | 78.8% 75.4–81.9 | 11.9% 9.3–14.6 | 1.0% 0.3–1.9 | one answer | unknown 100% · 2.4 |
Gold answer, closed world task
Anything that cannot be derived is false, so there is no unknown answer. Each line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance (50%); the grey tick is the score of always giving the slice's most common answer.
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Gold answer | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| false | 900 | 88.9% 86.9–90.9 | 63.8% 60.6–66.9 | 62.9% 59.9–65.9 | 75.3% 72.8–78.2 | 66.1% 63.0–69.3 | 86.4% 84.1–88.6 | one answer | false 100% · 2.5 |
| true | 900 | 89.7% 87.7–91.6 | 66.2% 63.1–69.2 | 64.7% 61.7–67.6 | 42.2% 39.2–45.4 | 45.4% 42.1–48.6 | 24.8% 21.9–27.6 | one answer | true 100% · 2.5 |
How often each model says unknown
On the open-world task, about a third of the correct answers are unknown. A model that rarely says unknown cannot get those items right.