Breakdown · descriptive

Problem size and other parameters: length, rules, facts and proof size

Every other property ProofWriter records about a problem: how long the theory is, how many rules and facts it has, how big the proof is, how deep the theory goes, whether it states attributes or relations, and how the question was generated. None of them was controlled when the problems were sampled.

Descriptive, not causal: the sample controls proof depth and the answer, not these properties. Several of them grow with depth and with each other, so each section gives the average proof depth of its slices; the answer mix and mean depth of every slice are in the tables.

Accuracy by theory length

The theory's length in words. Longer theories carry more distractors, and they also tend to support deeper proofs.

On the open-world task, Jev's accuracy ranges from 79.8% (80-109) to 93.1% (0-49), a spread of 13.3 points.

Descriptive, not causal. The average proof depth of these slices runs from 1.5 to 2.7, so depth differs between them and can add to or mask any gap.

Theory length, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
0-49247 items
50-79416 items
80-109519 items
110+618 items
Theory length, Open world
Theory lengthItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
0-4924793.1% 89.9–96.491.5% 87.4–94.777.7% 72.5–82.675.7% 70.4–81.056.3% 49.8–62.347.8% 42.1–54.339.7%false 30% · true 30% · unknown 40% · 1.5
50-7941689.4% 86.3–92.372.8% 68.8–76.963.2% 58.7–68.062.0% 57.0–66.354.6% 49.8–59.440.4% 35.6–45.436.8%false 32% · true 31% · unknown 37% · 2.6
80-10951979.8% 76.1–83.057.6% 53.6–61.858.4% 54.1–62.650.3% 45.7–54.752.6% 48.6–56.640.3% 36.0–44.534.3%false 33% · true 33% · unknown 34% · 2.6
110+61879.8% 76.7–83.052.6% 48.9–56.347.9% 44.0–51.941.7% 38.0–45.652.6% 48.7–56.543.5% 39.6–47.237.1%false 37% · true 37% · unknown 26% · 2.7

Theory length, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
0-49256 items
50-79432 items
80-109482 items
110+630 items
Theory length, Closed world
Theory lengthItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
0-4925699.2% 98.0–100.084.8% 80.5–89.173.8% 68.8–78.971.1% 65.6–76.664.5% 58.6–69.561.3% 55.9–67.250.8%false 49% · true 51% · 1.4
50-7943294.9% 92.8–96.866.9% 62.5–71.367.6% 63.4–72.260.4% 55.6–65.054.4% 49.8–59.355.8% 50.7–60.450.2%false 50% · true 50% · 2.6
80-10948285.7% 82.4–88.863.1% 58.7–67.465.8% 61.6–70.159.1% 55.0–63.355.6% 51.2–60.255.2% 51.0–59.850.8%false 51% · true 49% · 2.6
110+63084.1% 81.3–87.057.1% 53.5–61.155.6% 51.6–59.252.4% 48.4–56.253.3% 49.5–57.153.5% 49.7–57.350.5%false 50% · true 50% · 2.8

Accuracy by the number of rules

How many rules the theory states.

On the open-world task, Jev's accuracy ranges from 80.0% (8+) to 92.2% (0-2), a spread of 12.2 points.

Descriptive, not causal. The average proof depth of these slices runs from 0.9 to 3.0, so depth differs between them and can add to or mask any gap.

Number of rules, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
0-2204 items
3-5438 items
6-7564 items
8+594 items
Number of rules, Open world
Number of rulesItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
0-220492.2% 88.2–95.690.2% 85.8–94.178.9% 73.5–84.878.9% 73.5–84.358.8% 52.0–65.253.9% 47.1–60.833.8%false 32% · true 34% · unknown 34% · 0.9
3-543883.8% 80.4–87.059.1% 54.6–63.751.6% 46.6–56.245.9% 40.9–50.257.8% 53.2–62.346.1% 41.6–50.938.4%false 38% · true 38% · unknown 24% · 2.3
6-756484.9% 81.9–88.161.3% 57.8–65.456.9% 53.0–60.849.5% 45.6–53.752.3% 48.2–56.442.6% 38.3–46.634.8%false 35% · true 34% · unknown 31% · 2.6
8+59480.0% 76.4–83.261.3% 57.2–65.358.2% 54.5–62.154.4% 50.8–58.649.8% 46.0–53.935.7% 31.8–39.940.7%false 29% · true 30% · unknown 41% · 3.0

Number of rules, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
0-2208 items
3-5405 items
6-7605 items
8+582 items
Number of rules, Closed world
Number of rulesItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
0-220898.6% 96.6–100.083.7% 78.8–88.574.5% 68.8–80.368.8% 63.0–75.567.3% 61.1–73.658.7% 51.9–65.450.0%false 50% · true 50% · 1.0
3-540585.9% 82.7–89.464.7% 60.2–69.160.0% 55.3–64.957.3% 52.6–62.062.2% 57.3–66.455.6% 51.1–60.550.4%false 50% · true 50% · 2.2
6-760584.8% 82.0–87.657.2% 53.4–61.254.7% 51.1–58.551.6% 47.8–55.550.6% 46.6–54.754.9% 50.9–58.750.6%false 49% · true 51% · 2.8
8+58293.0% 90.7–95.066.7% 62.7–70.472.0% 68.2–75.663.7% 59.8–67.552.6% 48.3–56.455.3% 51.2–59.550.3%false 50% · true 50% · 2.9

Accuracy by the number of facts

How many facts the theory states.

On the open-world task, Jev's accuracy ranges from 80.3% (8-12) to 90.4% (0-3), a spread of 10.1 points.

Descriptive, not causal. The average proof depth of these slices runs from 1.8 to 3.0, so depth differs between them and can add to or mask any gap.

Number of facts, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
0-3334 items
4-7435 items
8-12872 items
13+159 items
Number of facts, Open world
Number of factsItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
0-333490.4% 86.8–93.481.4% 77.2–85.371.9% 66.8–76.669.8% 65.3–74.652.4% 47.0–57.842.5% 37.1–47.942.8%false 28% · true 29% · unknown 43% · 1.8
4-743586.0% 82.5–89.066.9% 62.3–71.361.8% 57.5–66.459.3% 54.7–64.155.6% 51.0–60.241.6% 36.8–46.436.1%false 33% · true 31% · unknown 36% · 2.7
8-1287280.3% 77.5–82.957.3% 54.0–60.754.0% 50.8–57.146.9% 43.6–49.952.4% 49.1–55.742.5% 39.4–45.635.8%false 35% · true 36% · unknown 29% · 2.6
13+15983.6% 78.0–89.356.6% 49.1–64.846.5% 39.0–54.140.3% 32.7–48.456.6% 48.4–64.244.0% 36.5–51.638.4%false 38% · true 38% · unknown 24% · 3.0

Number of facts, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
0-3360 items
4-7425 items
8-12812 items
13+203 items
Number of facts, Closed world
Number of factsItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
0-336095.6% 93.3–97.571.9% 67.2–76.772.8% 68.3–77.868.1% 63.3–73.358.6% 53.6–63.958.6% 53.6–63.651.9%false 48% · true 52% · 1.9
4-742592.5% 90.1–94.871.8% 67.3–76.070.4% 65.9–74.663.8% 59.1–68.555.1% 50.1–59.559.1% 54.6–63.351.3%false 51% · true 49% · 2.6
8-1281284.4% 81.9–86.961.6% 58.4–65.159.5% 56.0–63.155.2% 51.8–58.654.1% 50.9–57.452.7% 49.1–56.350.1%false 50% · true 50% · 2.5
13+20391.1% 87.2–95.152.2% 45.3–58.651.2% 44.3–58.146.3% 39.4–53.259.1% 51.7–65.554.7% 47.8–61.650.2%false 50% · true 50% · 3.3

Accuracy by proof size

The size of the answer's proof as ProofWriter records it (its QLen field). Unlike depth it also grows with branching: a proof can be shallow and wide.

On the open-world task, Jev's accuracy ranges from 80.4% (0-1) to 98.1% (2), a spread of 17.8 points.

Descriptive, not causal. The average proof depth of these slices runs from 1.0 to 3.8, so depth differs between them and can add to or mask any gap.

Proof size, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
0-1790 items
2107 items
3-4296 items
5+607 items
Proof size, Open world
Proof sizeItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
0-179080.4% 77.5–82.885.2% 82.8–87.575.1% 72.0–78.181.4% 78.9–84.232.0% 28.9–35.124.3% 21.5–27.174.7%false 13% · true 13% · unknown 75% · 1.8
210798.1% 95.3–100.081.3% 73.8–88.884.1% 76.6–90.769.2% 59.8–77.677.6% 69.2–85.063.6% 54.2–72.950.5%false 50% · true 50% · 1.0
3-429691.2% 87.5–94.363.9% 58.4–69.360.1% 54.4–65.541.9% 36.8–47.369.9% 64.9–75.362.5% 57.4–68.250.3%false 50% · true 50% · 2.0
5+60782.2% 79.1–85.233.6% 29.8–37.431.8% 28.3–35.420.3% 17.3–23.469.4% 65.9–73.052.6% 48.8–56.350.1%false 50% · true 50% · 3.8

Proof size, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
0-1966 items
282 items
3-4237 items
5+515 items
Proof size, Closed world
Proof sizeItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
0-196694.6% 93.2–96.080.1% 77.5–82.781.0% 78.6–83.472.6% 69.8–75.547.0% 44.1–50.259.1% 55.9–62.350.9%false 49% · true 51% · 2.0
28297.6% 93.9–100.078.0% 68.3–86.679.3% 70.7–87.879.3% 70.7–87.885.4% 78.0–92.756.1% 45.1–67.150.0%false 50% · true 50% · 1.0
3-423791.1% 87.8–94.563.3% 57.4–69.658.6% 52.3–65.051.9% 45.1–58.270.9% 64.6–76.452.3% 46.0–58.650.2%false 50% · true 50% · 2.0
5+51577.1% 73.4–80.635.3% 31.5–39.831.5% 27.6–35.532.8% 28.7–37.160.6% 56.1–64.350.5% 46.4–55.051.8%false 52% · true 48% · 3.9

Accuracy by the deepest proof in the theory

The depth of the deepest conclusion anywhere in the theory, whatever the question asks. It tells you how much inference the theory supports, not how much this question needs.

On the open-world task, Jev's accuracy ranges from 67.9% (4) to 90.1% (0), a spread of 22.2 points.

Descriptive, not causal. The average proof depth of these slices runs from 0.9 to 3.8, so depth differs between them and can add to or mask any gap.

Theory's deepest proof, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
071 items
1145 items
2206 items
3620 items
4162 items
5571 items
618 items
75 items
82 items
Theory's deepest proof, Open world
Theory's deepest proofItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
07190.1% 83.1–97.287.3% 78.9–94.487.3% 78.9–94.485.9% 77.5–93.042.3% 31.0–53.539.4% 28.2–52.160.6%false 18% · true 21% · unknown 61% · 0.9
114588.3% 82.8–93.884.8% 79.3–90.381.4% 74.5–87.676.6% 69.7–82.853.8% 46.2–61.448.3% 40.0–56.643.4%false 29% · true 28% · unknown 43% · 1.1
220686.4% 81.6–91.382.0% 76.7–86.968.9% 62.1–74.869.9% 63.6–75.749.0% 42.2–56.343.2% 36.4–50.039.3%false 30% · true 31% · unknown 39% · 1.7
362084.4% 81.5–87.368.4% 64.8–72.364.4% 60.3–68.157.6% 53.7–61.353.5% 49.7–57.442.4% 38.4–46.533.9%false 33% · true 33% · unknown 34% · 1.9
416267.9% 61.1–75.340.7% 32.7–47.542.6% 35.2–50.038.9% 32.1–46.351.2% 43.2–58.643.2% 35.8–51.237.7%false 38% · true 34% · unknown 28% · 2.8
557185.5% 82.7–88.451.8% 47.6–55.944.3% 40.1–48.238.0% 33.6–42.258.0% 53.8–61.841.7% 37.8–45.538.4%false 38% · true 38% · unknown 23% · 3.8
61872.2% 50.0–94.455.6% 33.3–77.844.4% 22.2–66.733.3% 11.1–55.638.9% 16.7–66.722.2% 5.6–44.444.4%false 28% · true 28% · unknown 44% · 3.3
7560.0% 20.0–100.020.0% 0.0–60.020.0% 0.0–60.060.0% 20.0–100.020.0% 0.0–60.040.0% 0.0–80.060.0%false 20% · true 20% · unknown 60% · 4.0
82100.0% 100.0–100.0100.0% 100.0–100.0100.0% 100.0–100.0100.0% 100.0–100.050.0% 0.0–100.00.0% 0.0–0.0one answerunknown 100% · 4.0

Theory's deepest proof, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
071 items
1155 items
2230 items
3619 items
4171 items
5513 items
638 items
72 items
81 items
Theory's deepest proof, Closed world
Theory's deepest proofItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
07197.2% 93.0–100.081.7% 71.8–90.178.9% 69.0–87.360.6% 49.3–71.854.9% 42.3–66.247.9% 35.2–60.676.1%false 24% · true 76% · 1.1
115594.2% 90.3–98.172.3% 65.2–79.476.1% 69.0–82.668.4% 61.3–75.558.7% 51.0–67.161.3% 53.5–68.451.0%false 49% · true 51% · 1.4
223095.2% 92.2–97.880.0% 74.8–84.877.0% 71.3–82.267.8% 61.7–73.953.9% 47.0–60.956.5% 50.4–63.552.6%false 47% · true 53% · 1.9
361988.7% 86.3–91.369.1% 65.8–72.968.8% 65.1–72.763.7% 59.9–67.754.6% 50.6–58.555.7% 51.9–59.552.2%false 52% · true 48% · 2.0
417164.9% 57.9–71.348.0% 40.9–55.050.9% 43.3–59.148.5% 41.5–55.655.6% 48.0–63.257.3% 49.7–64.350.9%false 49% · true 51% · 3.1
551393.6% 91.2–95.555.0% 50.7–59.551.3% 47.2–55.849.7% 45.6–54.056.9% 52.8–61.253.2% 49.1–57.952.4%false 52% · true 48% · 3.7
63878.9% 65.8–92.157.9% 42.1–73.750.0% 34.2–65.852.6% 36.8–68.463.2% 47.4–78.965.8% 50.0–78.957.9%false 58% · true 42% · 2.9
72100.0% 100.0–100.0100.0% 100.0–100.050.0% 0.0–100.050.0% 0.0–100.050.0% 0.0–100.050.0% 0.0–100.0one answertrue 100% · 3.5
81100.0% 100.0–100.00.0% 0.0–0.0100.0% 100.0–100.00.0% 0.0–0.00.0% 0.0–0.00.0% 0.0–0.0one answertrue 100% · 5.0

Attribute theories vs relation theories

Whether the theory states attributes of entities ("Bob is kind") or relations between them ("The bear eats the squirrel").

On the open-world task, Jev's accuracy ranges from 82.8% (attribute) to 85.3% (relation), a spread of 2.5 points.

Descriptive, not causal. The average proof depth of these slices runs from 2.2 to 2.7, so depth differs between them and can add to or mask any gap.

Theory kind, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
attribute1,059 items
relation741 items
Theory kind, Open world
Theory kindItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
attribute1,05982.8% 80.5–85.262.5% 59.8–65.359.3% 56.4–62.354.4% 51.4–57.253.8% 50.7–56.641.2% 38.1–44.133.8%false 33% · true 33% · unknown 34% · 2.7
relation74185.3% 82.6–87.766.3% 62.9–69.657.5% 53.8–61.052.4% 49.0–55.953.2% 49.4–56.744.3% 40.9–47.834.4%false 34% · true 34% · unknown 31% · 2.2

Theory kind, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
attribute1,016 items
relation784 items
Theory kind, Closed world
Theory kindItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
attribute1,01685.8% 83.7–87.763.5% 60.6–66.365.6% 62.8–68.558.8% 56.0–61.952.3% 49.3–55.554.8% 51.9–57.950.2%false 50% · true 50% · 2.6
relation78493.8% 92.1–95.467.0% 63.9–70.261.5% 58.2–64.958.8% 55.5–62.460.3% 56.9–63.656.6% 53.2–59.950.3%false 50% · true 50% · 2.3

Accuracy by how each question was generated

ProofWriter records how each statement was generated (its strategy field): from a proof (proof), from a rule's conclusion (rconc) or at random (random), each also in an inverted form (inv-). In practice the strategy fixes the gold answer, so this page mostly repeats the true, false and unknown breakdown.

On the open-world task, Jev's accuracy ranges from 50.2% (inv-rconc) to 98.0% (random), a spread of 47.8 points.

Descriptive, not causal. The average proof depth of these slices runs from 0.0 to 3.0, so depth differs between them and can add to or mask any gap.

Question strategy, open world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
inv-proof605 items
inv-random50 items
inv-rconc249 items
proof605 items
random50 items
rconc241 items
Question strategy, Open world
Question strategyItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
inv-proof60588.3% 85.5–90.950.9% 46.8–54.948.6% 44.6–52.635.9% 31.9–39.589.6% 87.1–91.785.1% 82.1–87.9one answerfalse 100% · 2.5
inv-random5094.0% 86.0–100.0100.0% 100.0–100.076.0% 64.0–86.082.0% 70.0–92.022.0% 12.0–34.02.0% 0.0–6.0one answerunknown 100% · 0.0
inv-rconc24950.2% 43.8–56.678.3% 73.1–83.562.7% 56.2–68.372.3% 66.3–77.50.8% 0.0–2.00.8% 0.0–2.0one answerunknown 100% · 3.0
proof60588.8% 86.1–91.158.3% 54.0–62.159.2% 55.4–63.146.6% 42.8–50.658.2% 54.0–62.040.2% 36.2–44.1one answertrue 100% · 2.5
random5098.0% 94.0–100.098.0% 94.0–100.078.0% 66.0–88.088.0% 78.0–96.034.0% 22.0–46.02.0% 0.0–6.0one answerunknown 100% · 0.0
rconc24190.0% 85.9–93.882.2% 77.2–87.170.1% 64.7–75.983.0% 78.0–87.616.6% 12.4–21.60.8% 0.0–2.1one answerunknown 100% · 2.9

Question strategy, closed world task

  • Jev
  • GPT-6 Luna(reasoning off)
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
  • Laya
inv-proof501 items
inv-random75 items
inv-rconc342 items
proof483 items
random75 items
rconc324 items
Question strategy, Closed world
Question strategyItemsJevGPT-6 LunaKev-9BKev-4BKev-0.8BLayaBest constantAnswer mix · mean depth
inv-proof50185.0% 81.6–88.451.9% 47.7–56.549.9% 45.5–54.365.9% 61.9–69.981.4% 78.2–84.684.4% 81.0–87.6one answerfalse 100% · 2.6
inv-random75100.0% 100.0–100.098.7% 96.0–100.069.3% 58.7–80.060.0% 49.3–70.717.3% 9.3–26.717.3% 9.3–26.7one answertrue 100% · 0.0
inv-rconc34292.4% 89.8–95.070.5% 65.2–75.178.1% 73.4–82.548.5% 43.0–53.532.2% 26.9–36.819.0% 15.2–23.1one answertrue 100% · 3.0
proof48386.1% 83.0–89.058.2% 53.8–62.954.5% 49.7–59.035.0% 30.8–39.159.2% 54.7–64.030.0% 25.9–34.0one answertrue 100% · 2.5
random7598.7% 96.0–100.0100.0% 100.0–100.093.3% 88.0–98.793.3% 88.0–98.762.7% 52.0–72.086.7% 78.7–93.3one answerfalse 100% · 0.0
rconc32492.6% 89.8–95.473.8% 68.8–78.775.9% 71.3–80.685.8% 81.8–89.543.2% 38.0–48.889.5% 85.8–92.6one answerfalse 100% · 2.9

In each chart a line is one model: its accuracy, with the 95% interval as a bar. The dashed line is chance; the grey tick is the score of always giving the slice's most common answer.