Breakdown · descriptive
Problem size and other parameters: length, rules, facts and proof size
Every other property ProofWriter records about a problem: how long the theory is, how many rules and facts it has, how big the proof is, how deep the theory goes, whether it states attributes or relations, and how the question was generated. None of them was controlled when the problems were sampled.
Descriptive, not causal: the sample controls proof depth and the answer, not these properties. Several of them grow with depth and with each other, so each section gives the average proof depth of its slices; the answer mix and mean depth of every slice are in the tables.
Accuracy by theory length
The theory's length in words. Longer theories carry more distractors, and they also tend to support deeper proofs.
On the open-world task, Jev's accuracy ranges from 79.8% (80-109) to 93.1% (0-49), a spread of 13.3 points.
Descriptive, not causal. The average proof depth of these slices runs from 1.5 to 2.7, so depth differs between them and can add to or mask any gap.
Theory length, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Theory length | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0-49 | 247 | 93.1% 89.9–96.4 | 91.5% 87.4–94.7 | 77.7% 72.5–82.6 | 75.7% 70.4–81.0 | 56.3% 49.8–62.3 | 47.8% 42.1–54.3 | 39.7% | false 30% · true 30% · unknown 40% · 1.5 |
| 50-79 | 416 | 89.4% 86.3–92.3 | 72.8% 68.8–76.9 | 63.2% 58.7–68.0 | 62.0% 57.0–66.3 | 54.6% 49.8–59.4 | 40.4% 35.6–45.4 | 36.8% | false 32% · true 31% · unknown 37% · 2.6 |
| 80-109 | 519 | 79.8% 76.1–83.0 | 57.6% 53.6–61.8 | 58.4% 54.1–62.6 | 50.3% 45.7–54.7 | 52.6% 48.6–56.6 | 40.3% 36.0–44.5 | 34.3% | false 33% · true 33% · unknown 34% · 2.6 |
| 110+ | 618 | 79.8% 76.7–83.0 | 52.6% 48.9–56.3 | 47.9% 44.0–51.9 | 41.7% 38.0–45.6 | 52.6% 48.7–56.5 | 43.5% 39.6–47.2 | 37.1% | false 37% · true 37% · unknown 26% · 2.7 |
Theory length, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Theory length | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0-49 | 256 | 99.2% 98.0–100.0 | 84.8% 80.5–89.1 | 73.8% 68.8–78.9 | 71.1% 65.6–76.6 | 64.5% 58.6–69.5 | 61.3% 55.9–67.2 | 50.8% | false 49% · true 51% · 1.4 |
| 50-79 | 432 | 94.9% 92.8–96.8 | 66.9% 62.5–71.3 | 67.6% 63.4–72.2 | 60.4% 55.6–65.0 | 54.4% 49.8–59.3 | 55.8% 50.7–60.4 | 50.2% | false 50% · true 50% · 2.6 |
| 80-109 | 482 | 85.7% 82.4–88.8 | 63.1% 58.7–67.4 | 65.8% 61.6–70.1 | 59.1% 55.0–63.3 | 55.6% 51.2–60.2 | 55.2% 51.0–59.8 | 50.8% | false 51% · true 49% · 2.6 |
| 110+ | 630 | 84.1% 81.3–87.0 | 57.1% 53.5–61.1 | 55.6% 51.6–59.2 | 52.4% 48.4–56.2 | 53.3% 49.5–57.1 | 53.5% 49.7–57.3 | 50.5% | false 50% · true 50% · 2.8 |
Accuracy by the number of rules
How many rules the theory states.
On the open-world task, Jev's accuracy ranges from 80.0% (8+) to 92.2% (0-2), a spread of 12.2 points.
Descriptive, not causal. The average proof depth of these slices runs from 0.9 to 3.0, so depth differs between them and can add to or mask any gap.
Number of rules, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Number of rules | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0-2 | 204 | 92.2% 88.2–95.6 | 90.2% 85.8–94.1 | 78.9% 73.5–84.8 | 78.9% 73.5–84.3 | 58.8% 52.0–65.2 | 53.9% 47.1–60.8 | 33.8% | false 32% · true 34% · unknown 34% · 0.9 |
| 3-5 | 438 | 83.8% 80.4–87.0 | 59.1% 54.6–63.7 | 51.6% 46.6–56.2 | 45.9% 40.9–50.2 | 57.8% 53.2–62.3 | 46.1% 41.6–50.9 | 38.4% | false 38% · true 38% · unknown 24% · 2.3 |
| 6-7 | 564 | 84.9% 81.9–88.1 | 61.3% 57.8–65.4 | 56.9% 53.0–60.8 | 49.5% 45.6–53.7 | 52.3% 48.2–56.4 | 42.6% 38.3–46.6 | 34.8% | false 35% · true 34% · unknown 31% · 2.6 |
| 8+ | 594 | 80.0% 76.4–83.2 | 61.3% 57.2–65.3 | 58.2% 54.5–62.1 | 54.4% 50.8–58.6 | 49.8% 46.0–53.9 | 35.7% 31.8–39.9 | 40.7% | false 29% · true 30% · unknown 41% · 3.0 |
Number of rules, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Number of rules | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0-2 | 208 | 98.6% 96.6–100.0 | 83.7% 78.8–88.5 | 74.5% 68.8–80.3 | 68.8% 63.0–75.5 | 67.3% 61.1–73.6 | 58.7% 51.9–65.4 | 50.0% | false 50% · true 50% · 1.0 |
| 3-5 | 405 | 85.9% 82.7–89.4 | 64.7% 60.2–69.1 | 60.0% 55.3–64.9 | 57.3% 52.6–62.0 | 62.2% 57.3–66.4 | 55.6% 51.1–60.5 | 50.4% | false 50% · true 50% · 2.2 |
| 6-7 | 605 | 84.8% 82.0–87.6 | 57.2% 53.4–61.2 | 54.7% 51.1–58.5 | 51.6% 47.8–55.5 | 50.6% 46.6–54.7 | 54.9% 50.9–58.7 | 50.6% | false 49% · true 51% · 2.8 |
| 8+ | 582 | 93.0% 90.7–95.0 | 66.7% 62.7–70.4 | 72.0% 68.2–75.6 | 63.7% 59.8–67.5 | 52.6% 48.3–56.4 | 55.3% 51.2–59.5 | 50.3% | false 50% · true 50% · 2.9 |
Accuracy by the number of facts
How many facts the theory states.
On the open-world task, Jev's accuracy ranges from 80.3% (8-12) to 90.4% (0-3), a spread of 10.1 points.
Descriptive, not causal. The average proof depth of these slices runs from 1.8 to 3.0, so depth differs between them and can add to or mask any gap.
Number of facts, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Number of facts | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0-3 | 334 | 90.4% 86.8–93.4 | 81.4% 77.2–85.3 | 71.9% 66.8–76.6 | 69.8% 65.3–74.6 | 52.4% 47.0–57.8 | 42.5% 37.1–47.9 | 42.8% | false 28% · true 29% · unknown 43% · 1.8 |
| 4-7 | 435 | 86.0% 82.5–89.0 | 66.9% 62.3–71.3 | 61.8% 57.5–66.4 | 59.3% 54.7–64.1 | 55.6% 51.0–60.2 | 41.6% 36.8–46.4 | 36.1% | false 33% · true 31% · unknown 36% · 2.7 |
| 8-12 | 872 | 80.3% 77.5–82.9 | 57.3% 54.0–60.7 | 54.0% 50.8–57.1 | 46.9% 43.6–49.9 | 52.4% 49.1–55.7 | 42.5% 39.4–45.6 | 35.8% | false 35% · true 36% · unknown 29% · 2.6 |
| 13+ | 159 | 83.6% 78.0–89.3 | 56.6% 49.1–64.8 | 46.5% 39.0–54.1 | 40.3% 32.7–48.4 | 56.6% 48.4–64.2 | 44.0% 36.5–51.6 | 38.4% | false 38% · true 38% · unknown 24% · 3.0 |
Number of facts, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Number of facts | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0-3 | 360 | 95.6% 93.3–97.5 | 71.9% 67.2–76.7 | 72.8% 68.3–77.8 | 68.1% 63.3–73.3 | 58.6% 53.6–63.9 | 58.6% 53.6–63.6 | 51.9% | false 48% · true 52% · 1.9 |
| 4-7 | 425 | 92.5% 90.1–94.8 | 71.8% 67.3–76.0 | 70.4% 65.9–74.6 | 63.8% 59.1–68.5 | 55.1% 50.1–59.5 | 59.1% 54.6–63.3 | 51.3% | false 51% · true 49% · 2.6 |
| 8-12 | 812 | 84.4% 81.9–86.9 | 61.6% 58.4–65.1 | 59.5% 56.0–63.1 | 55.2% 51.8–58.6 | 54.1% 50.9–57.4 | 52.7% 49.1–56.3 | 50.1% | false 50% · true 50% · 2.5 |
| 13+ | 203 | 91.1% 87.2–95.1 | 52.2% 45.3–58.6 | 51.2% 44.3–58.1 | 46.3% 39.4–53.2 | 59.1% 51.7–65.5 | 54.7% 47.8–61.6 | 50.2% | false 50% · true 50% · 3.3 |
Accuracy by proof size
The size of the answer's proof as ProofWriter records it (its QLen field). Unlike depth it also grows with branching: a proof can be shallow and wide.
On the open-world task, Jev's accuracy ranges from 80.4% (0-1) to 98.1% (2), a spread of 17.8 points.
Descriptive, not causal. The average proof depth of these slices runs from 1.0 to 3.8, so depth differs between them and can add to or mask any gap.
Proof size, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Proof size | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0-1 | 790 | 80.4% 77.5–82.8 | 85.2% 82.8–87.5 | 75.1% 72.0–78.1 | 81.4% 78.9–84.2 | 32.0% 28.9–35.1 | 24.3% 21.5–27.1 | 74.7% | false 13% · true 13% · unknown 75% · 1.8 |
| 2 | 107 | 98.1% 95.3–100.0 | 81.3% 73.8–88.8 | 84.1% 76.6–90.7 | 69.2% 59.8–77.6 | 77.6% 69.2–85.0 | 63.6% 54.2–72.9 | 50.5% | false 50% · true 50% · 1.0 |
| 3-4 | 296 | 91.2% 87.5–94.3 | 63.9% 58.4–69.3 | 60.1% 54.4–65.5 | 41.9% 36.8–47.3 | 69.9% 64.9–75.3 | 62.5% 57.4–68.2 | 50.3% | false 50% · true 50% · 2.0 |
| 5+ | 607 | 82.2% 79.1–85.2 | 33.6% 29.8–37.4 | 31.8% 28.3–35.4 | 20.3% 17.3–23.4 | 69.4% 65.9–73.0 | 52.6% 48.8–56.3 | 50.1% | false 50% · true 50% · 3.8 |
Proof size, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Proof size | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0-1 | 966 | 94.6% 93.2–96.0 | 80.1% 77.5–82.7 | 81.0% 78.6–83.4 | 72.6% 69.8–75.5 | 47.0% 44.1–50.2 | 59.1% 55.9–62.3 | 50.9% | false 49% · true 51% · 2.0 |
| 2 | 82 | 97.6% 93.9–100.0 | 78.0% 68.3–86.6 | 79.3% 70.7–87.8 | 79.3% 70.7–87.8 | 85.4% 78.0–92.7 | 56.1% 45.1–67.1 | 50.0% | false 50% · true 50% · 1.0 |
| 3-4 | 237 | 91.1% 87.8–94.5 | 63.3% 57.4–69.6 | 58.6% 52.3–65.0 | 51.9% 45.1–58.2 | 70.9% 64.6–76.4 | 52.3% 46.0–58.6 | 50.2% | false 50% · true 50% · 2.0 |
| 5+ | 515 | 77.1% 73.4–80.6 | 35.3% 31.5–39.8 | 31.5% 27.6–35.5 | 32.8% 28.7–37.1 | 60.6% 56.1–64.3 | 50.5% 46.4–55.0 | 51.8% | false 52% · true 48% · 3.9 |
Accuracy by the deepest proof in the theory
The depth of the deepest conclusion anywhere in the theory, whatever the question asks. It tells you how much inference the theory supports, not how much this question needs.
On the open-world task, Jev's accuracy ranges from 67.9% (4) to 90.1% (0), a spread of 22.2 points.
Descriptive, not causal. The average proof depth of these slices runs from 0.9 to 3.8, so depth differs between them and can add to or mask any gap.
Theory's deepest proof, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Theory's deepest proof | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 71 | 90.1% 83.1–97.2 | 87.3% 78.9–94.4 | 87.3% 78.9–94.4 | 85.9% 77.5–93.0 | 42.3% 31.0–53.5 | 39.4% 28.2–52.1 | 60.6% | false 18% · true 21% · unknown 61% · 0.9 |
| 1 | 145 | 88.3% 82.8–93.8 | 84.8% 79.3–90.3 | 81.4% 74.5–87.6 | 76.6% 69.7–82.8 | 53.8% 46.2–61.4 | 48.3% 40.0–56.6 | 43.4% | false 29% · true 28% · unknown 43% · 1.1 |
| 2 | 206 | 86.4% 81.6–91.3 | 82.0% 76.7–86.9 | 68.9% 62.1–74.8 | 69.9% 63.6–75.7 | 49.0% 42.2–56.3 | 43.2% 36.4–50.0 | 39.3% | false 30% · true 31% · unknown 39% · 1.7 |
| 3 | 620 | 84.4% 81.5–87.3 | 68.4% 64.8–72.3 | 64.4% 60.3–68.1 | 57.6% 53.7–61.3 | 53.5% 49.7–57.4 | 42.4% 38.4–46.5 | 33.9% | false 33% · true 33% · unknown 34% · 1.9 |
| 4 | 162 | 67.9% 61.1–75.3 | 40.7% 32.7–47.5 | 42.6% 35.2–50.0 | 38.9% 32.1–46.3 | 51.2% 43.2–58.6 | 43.2% 35.8–51.2 | 37.7% | false 38% · true 34% · unknown 28% · 2.8 |
| 5 | 571 | 85.5% 82.7–88.4 | 51.8% 47.6–55.9 | 44.3% 40.1–48.2 | 38.0% 33.6–42.2 | 58.0% 53.8–61.8 | 41.7% 37.8–45.5 | 38.4% | false 38% · true 38% · unknown 23% · 3.8 |
| 6 | 18 | 72.2% 50.0–94.4 | 55.6% 33.3–77.8 | 44.4% 22.2–66.7 | 33.3% 11.1–55.6 | 38.9% 16.7–66.7 | 22.2% 5.6–44.4 | 44.4% | false 28% · true 28% · unknown 44% · 3.3 |
| 7 | 5 | 60.0% 20.0–100.0 | 20.0% 0.0–60.0 | 20.0% 0.0–60.0 | 60.0% 20.0–100.0 | 20.0% 0.0–60.0 | 40.0% 0.0–80.0 | 60.0% | false 20% · true 20% · unknown 60% · 4.0 |
| 8 | 2 | 100.0% 100.0–100.0 | 100.0% 100.0–100.0 | 100.0% 100.0–100.0 | 100.0% 100.0–100.0 | 50.0% 0.0–100.0 | 0.0% 0.0–0.0 | one answer | unknown 100% · 4.0 |
Theory's deepest proof, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Theory's deepest proof | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 71 | 97.2% 93.0–100.0 | 81.7% 71.8–90.1 | 78.9% 69.0–87.3 | 60.6% 49.3–71.8 | 54.9% 42.3–66.2 | 47.9% 35.2–60.6 | 76.1% | false 24% · true 76% · 1.1 |
| 1 | 155 | 94.2% 90.3–98.1 | 72.3% 65.2–79.4 | 76.1% 69.0–82.6 | 68.4% 61.3–75.5 | 58.7% 51.0–67.1 | 61.3% 53.5–68.4 | 51.0% | false 49% · true 51% · 1.4 |
| 2 | 230 | 95.2% 92.2–97.8 | 80.0% 74.8–84.8 | 77.0% 71.3–82.2 | 67.8% 61.7–73.9 | 53.9% 47.0–60.9 | 56.5% 50.4–63.5 | 52.6% | false 47% · true 53% · 1.9 |
| 3 | 619 | 88.7% 86.3–91.3 | 69.1% 65.8–72.9 | 68.8% 65.1–72.7 | 63.7% 59.9–67.7 | 54.6% 50.6–58.5 | 55.7% 51.9–59.5 | 52.2% | false 52% · true 48% · 2.0 |
| 4 | 171 | 64.9% 57.9–71.3 | 48.0% 40.9–55.0 | 50.9% 43.3–59.1 | 48.5% 41.5–55.6 | 55.6% 48.0–63.2 | 57.3% 49.7–64.3 | 50.9% | false 49% · true 51% · 3.1 |
| 5 | 513 | 93.6% 91.2–95.5 | 55.0% 50.7–59.5 | 51.3% 47.2–55.8 | 49.7% 45.6–54.0 | 56.9% 52.8–61.2 | 53.2% 49.1–57.9 | 52.4% | false 52% · true 48% · 3.7 |
| 6 | 38 | 78.9% 65.8–92.1 | 57.9% 42.1–73.7 | 50.0% 34.2–65.8 | 52.6% 36.8–68.4 | 63.2% 47.4–78.9 | 65.8% 50.0–78.9 | 57.9% | false 58% · true 42% · 2.9 |
| 7 | 2 | 100.0% 100.0–100.0 | 100.0% 100.0–100.0 | 50.0% 0.0–100.0 | 50.0% 0.0–100.0 | 50.0% 0.0–100.0 | 50.0% 0.0–100.0 | one answer | true 100% · 3.5 |
| 8 | 1 | 100.0% 100.0–100.0 | 0.0% 0.0–0.0 | 100.0% 100.0–100.0 | 0.0% 0.0–0.0 | 0.0% 0.0–0.0 | 0.0% 0.0–0.0 | one answer | true 100% · 5.0 |
Attribute theories vs relation theories
Whether the theory states attributes of entities ("Bob is kind") or relations between them ("The bear eats the squirrel").
On the open-world task, Jev's accuracy ranges from 82.8% (attribute) to 85.3% (relation), a spread of 2.5 points.
Descriptive, not causal. The average proof depth of these slices runs from 2.2 to 2.7, so depth differs between them and can add to or mask any gap.
Theory kind, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Theory kind | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| attribute | 1,059 | 82.8% 80.5–85.2 | 62.5% 59.8–65.3 | 59.3% 56.4–62.3 | 54.4% 51.4–57.2 | 53.8% 50.7–56.6 | 41.2% 38.1–44.1 | 33.8% | false 33% · true 33% · unknown 34% · 2.7 |
| relation | 741 | 85.3% 82.6–87.7 | 66.3% 62.9–69.6 | 57.5% 53.8–61.0 | 52.4% 49.0–55.9 | 53.2% 49.4–56.7 | 44.3% 40.9–47.8 | 34.4% | false 34% · true 34% · unknown 31% · 2.2 |
Theory kind, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Theory kind | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| attribute | 1,016 | 85.8% 83.7–87.7 | 63.5% 60.6–66.3 | 65.6% 62.8–68.5 | 58.8% 56.0–61.9 | 52.3% 49.3–55.5 | 54.8% 51.9–57.9 | 50.2% | false 50% · true 50% · 2.6 |
| relation | 784 | 93.8% 92.1–95.4 | 67.0% 63.9–70.2 | 61.5% 58.2–64.9 | 58.8% 55.5–62.4 | 60.3% 56.9–63.6 | 56.6% 53.2–59.9 | 50.3% | false 50% · true 50% · 2.3 |
Accuracy by how each question was generated
ProofWriter records how each statement was generated (its strategy field): from a proof (proof), from a rule's conclusion (rconc) or at random (random), each also in an inverted form (inv-). In practice the strategy fixes the gold answer, so this page mostly repeats the true, false and unknown breakdown.
On the open-world task, Jev's accuracy ranges from 50.2% (inv-rconc) to 98.0% (random), a spread of 47.8 points.
Descriptive, not causal. The average proof depth of these slices runs from 0.0 to 3.0, so depth differs between them and can add to or mask any gap.
Question strategy, open world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Question strategy | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| inv-proof | 605 | 88.3% 85.5–90.9 | 50.9% 46.8–54.9 | 48.6% 44.6–52.6 | 35.9% 31.9–39.5 | 89.6% 87.1–91.7 | 85.1% 82.1–87.9 | one answer | false 100% · 2.5 |
| inv-random | 50 | 94.0% 86.0–100.0 | 100.0% 100.0–100.0 | 76.0% 64.0–86.0 | 82.0% 70.0–92.0 | 22.0% 12.0–34.0 | 2.0% 0.0–6.0 | one answer | unknown 100% · 0.0 |
| inv-rconc | 249 | 50.2% 43.8–56.6 | 78.3% 73.1–83.5 | 62.7% 56.2–68.3 | 72.3% 66.3–77.5 | 0.8% 0.0–2.0 | 0.8% 0.0–2.0 | one answer | unknown 100% · 3.0 |
| proof | 605 | 88.8% 86.1–91.1 | 58.3% 54.0–62.1 | 59.2% 55.4–63.1 | 46.6% 42.8–50.6 | 58.2% 54.0–62.0 | 40.2% 36.2–44.1 | one answer | true 100% · 2.5 |
| random | 50 | 98.0% 94.0–100.0 | 98.0% 94.0–100.0 | 78.0% 66.0–88.0 | 88.0% 78.0–96.0 | 34.0% 22.0–46.0 | 2.0% 0.0–6.0 | one answer | unknown 100% · 0.0 |
| rconc | 241 | 90.0% 85.9–93.8 | 82.2% 77.2–87.1 | 70.1% 64.7–75.9 | 83.0% 78.0–87.6 | 16.6% 12.4–21.6 | 0.8% 0.0–2.1 | one answer | unknown 100% · 2.9 |
Question strategy, closed world task
- Jev
- GPT-6 Luna(reasoning off)
- Kev-9B
- Kev-4B
- Kev-0.8B
- Laya
| Question strategy | Items | Jev | GPT-6 Luna | Kev-9B | Kev-4B | Kev-0.8B | Laya | Best constant | Answer mix · mean depth |
|---|---|---|---|---|---|---|---|---|---|
| inv-proof | 501 | 85.0% 81.6–88.4 | 51.9% 47.7–56.5 | 49.9% 45.5–54.3 | 65.9% 61.9–69.9 | 81.4% 78.2–84.6 | 84.4% 81.0–87.6 | one answer | false 100% · 2.6 |
| inv-random | 75 | 100.0% 100.0–100.0 | 98.7% 96.0–100.0 | 69.3% 58.7–80.0 | 60.0% 49.3–70.7 | 17.3% 9.3–26.7 | 17.3% 9.3–26.7 | one answer | true 100% · 0.0 |
| inv-rconc | 342 | 92.4% 89.8–95.0 | 70.5% 65.2–75.1 | 78.1% 73.4–82.5 | 48.5% 43.0–53.5 | 32.2% 26.9–36.8 | 19.0% 15.2–23.1 | one answer | true 100% · 3.0 |
| proof | 483 | 86.1% 83.0–89.0 | 58.2% 53.8–62.9 | 54.5% 49.7–59.0 | 35.0% 30.8–39.1 | 59.2% 54.7–64.0 | 30.0% 25.9–34.0 | one answer | true 100% · 2.5 |
| random | 75 | 98.7% 96.0–100.0 | 100.0% 100.0–100.0 | 93.3% 88.0–98.7 | 93.3% 88.0–98.7 | 62.7% 52.0–72.0 | 86.7% 78.7–93.3 | one answer | false 100% · 0.0 |
| rconc | 324 | 92.6% 89.8–95.4 | 73.8% 68.8–78.7 | 75.9% 71.3–80.6 | 85.8% 81.8–89.5 | 43.2% 38.0–48.8 | 89.5% 85.8–92.6 | one answer | false 100% · 2.9 |
In each chart a line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance; the grey tick is the score of always giving the slice's most common answer.