Comparison · Jev vs Laya

Jev vs Laya: hosted and open decision models on deep reasoning

Laya is a small open decision model built on an encoder. Both it and Jev answered the same problems.

Overall: Jev and Laya

Open world: true, false or unknown

Jev minus Laya: +41.4 points on the same 1,800 items, 95% interval +38.8 to +44.0: Jev is clearly ahead.

Closed world: true or false

Jev minus Laya: +33.7 points on the same 1,800 items, 95% interval +30.7 to +36.5: Jev is clearly ahead.

Jev vs Laya by proof depth

Each point is about 300 problems that need that many inference steps, with its 95% interval.

Open world: true, false or unknown
  • Jev
  • Laya
0%25%50%75%100%012345proof depth (inference steps)chance 33%
Closed world: true or false
  • Jev
  • Laya
0%25%50%75%100%012345proof depth (inference steps)chance 50%

Jev minus Laya by depth, open world

depth 0300 items
+35.0
depth 1302 items
+45.7
depth 2303 items
+43.6
depth 3303 items
+41.6
depth 4303 items
+36.0
depth 5289 items
+46.7

Jev minus Laya by depth, closed world

depth 0300 items
+27.3
depth 1300 items
+42.7
depth 2300 items
+37.3
depth 3300 items
+33.7
depth 4300 items
+17.7
depth 5300 items
+43.3

Paired differences: positive means Jev was more accurate on the same items.

Accuracy by depth, open world

By depth, Open world
Proof depthItemsJevLayaBest constant
depth 030097.7% 96.0–99.362.7% 57.0–68.333.3%
depth 130287.7% 83.8–91.442.1% 36.8–47.733.4%
depth 230384.8% 80.2–88.441.3% 36.0–46.933.3%
depth 330379.9% 75.2–84.538.3% 33.0–44.633.3%
depth 430371.9% 66.7–77.236.0% 31.0–41.633.3%
depth 528981.0% 76.5–85.534.3% 29.1–39.834.9%

Accuracy by depth, closed world

By depth, Closed world
Proof depthItemsJevLayaBest constant
depth 030099.3% 98.3–100.072.0% 66.7–76.750.0%
depth 130093.7% 91.0–96.351.0% 45.0–56.050.0%
depth 230092.0% 88.7–94.754.7% 48.7–59.750.0%
depth 330085.0% 81.3–89.351.3% 45.7–56.750.0%
depth 430076.3% 71.7–81.058.7% 53.0–64.350.0%
depth 530089.3% 85.7–92.346.0% 40.3–51.050.0%

Where each one is better

Open world, Jev vs Laya. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, true answers and unknown answers. Laya is not clearly ahead on any depth or answer class. Too close to call: false answers.

Closed world, Jev vs Laya. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5 and true answers. Laya is not clearly ahead on any depth or answer class. Too close to call: false answers.

A slice counts as clearly ahead when its paired 95% interval excludes zero.

True, false and unknown, open world

By correct answer, Open world
Correct answerItemsJevLayaBest constant
false60588.3% 85.5–90.985.1% 82.1–87.9one answer
true60588.8% 86.1–91.140.2% 36.2–44.1one answer
unknown59074.2% 70.7–77.61.0% 0.3–1.9one answer

True, false and unknown, closed world

By correct answer, Closed world
Correct answerItemsJevLayaBest constant
false90088.9% 86.9–90.986.4% 84.1–88.6one answer
true90089.7% 87.7–91.624.8% 21.9–27.6one answer

Paraphrased vs templated rules, open world

By wording, Open world
WordingItemsJevLayaBest constant
templated1,29886.1% 84.2–87.941.7% 39.2–44.334.5%
paraphrased (NatLang)50278.1% 74.5–82.144.4% 40.2–48.835.9%

Paraphrased vs templated rules, closed world

By wording, Closed world
WordingItemsJevLayaBest constant
templated1,30694.0% 92.8–95.356.1% 53.4–58.950.2%
paraphrased (NatLang)49476.7% 73.1–80.454.3% 49.8–58.950.6%

Where Laya errs

Accuracy alone hides which answers a model gets wrong. Rows are the correct answer; columns are what the model said.

Laya, open-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseSaid unknownRecall
true243354840.2%
false89515185.1%
unknown12745761.0%

Laya answered false 73.7%, true 25.5%, unknown 0.8% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.

Laya, closed-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseRecall
true22367724.8%
false12277886.4%

Laya answered false 80.8%, true 19.2% of the time; the correct answers are false 50.0%, true 50.0%.

Laya across every breakdown

For each property of the problems, the weakest and strongest value on the open-world task (slices of at least 50 items). Only depth was controlled; the rest are descriptive.

BreakdownWeakest · strongest
Proof depthdepth 5 34.3% · depth 0 62.7% (28.4 points)
Gold answerunknown 1.0% · false 85.1% (84.1 points)
Theory kindattribute 41.2% · relation 44.3% (3.1 points)
Negation in the theorywithout negation 38.8% · with negation 46.2% (7.4 points)
Negated statementplain statement 38.7% · negated statement 46.2% (7.5 points)
Question strategyinv-rconc 0.8% · inv-proof 85.1% (84.3 points)
Paraphrased rulestemplated 41.7% · paraphrased (NatLang) 44.4% (2.7 points)
Theory's deepest proof0 39.4% · 1 48.3% (8.8 points)
Theory length80-109 40.3% · 0-49 47.8% (7.5 points)
Number of rules8+ 35.7% · 0-2 53.9% (18.2 points)
Number of facts4-7 41.6% · 13+ 44.0% (2.4 points)
Proof size0-1 24.3% · 2 63.6% (39.2 points)

Consistency and calibration

Asked every open world question a second time, Laya gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 42.4% the first time and 42.4% the second.

Asked every closed world question a second time, Laya gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 55.6% the first time and 55.6% the second.

Calibration on the open world task: Laya's probability for its own answer averaged 74.7% against 42.4% accuracy; expected calibration error 32.3 points.

Calibration on the closed world task: Laya's probability for its own answer averaged 87.3% against 55.6% accuracy; expected calibration error 31.7 points.

Size, hardware and speed

ModelKindParametersRan onMedian time per decisionMemory peak
Jevhosted decision modelundisclosedvendor's servers189 msnot applicable (hosted)
Layaopen decision model421Mour laptop (Apple M1 Max, 32 GB)140 ms10.0 GB

Times are open-world, one request at a time. Hosted models were timed over the network against the vendor's servers and open models on one laptop, so speed is not like for like across the two. Speed, size and memory for every model.

Laya

Parameters
421M
Architecture
ModernBERT-large encoder with decision heads, trained with RLCD
Weights
0.84 GB (model.safetensors, convaiinnovations/laya repo root)
Precision
as loaded by the laya package (PyTorch)
Where it ran
this machine, laya 0.3.21 package, PyTorch on MPS
Maker's stated hardware
not stated by the maker
Memory measured
10.0 GB in use, 10.0 GB peak
Time per decision, open world
median 140 ms, 90th percentile 479 ms, on a laptop
Time per decision, closed world
median 121 ms, 90th percentile 398 ms, on a laptop
Source
huggingface.co/convaiinnovations/laya (model card checkpoint table)

The models, and what we predicted

Jev Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.

Laya Laya is an open decision model made by Convai Innovations: a 421M-parameter ModernBERT-large encoder with decision heads. We ran it on our own laptop with the laya package, PyTorch on Apple's GPU.