Comparison · Jev vs Laya
Jev vs Laya: hosted and open decision models on deep reasoning
Laya is a small open decision model built on an encoder. Both it and Jev answered the same problems.
Overall: Jev and Laya
Open world: true, false or unknown
- Jev82.2 to 85.5 (95%), 1,800 items83.8%
- Laya40.3 to 44.5 (95%), 1,800 items42.4%
Jev minus Laya: +41.4 points on the same 1,800 items, 95% interval +38.8 to +44.0: Jev is clearly ahead.
Closed world: true or false
- Jev87.8 to 90.7 (95%), 1,800 items89.3%
- Laya53.3 to 58.1 (95%), 1,800 items55.6%
Jev minus Laya: +33.7 points on the same 1,800 items, 95% interval +30.7 to +36.5: Jev is clearly ahead.
Jev vs Laya by proof depth
Each point is about 300 problems that need that many inference steps, with its 95% interval.
- Jev
- Laya
- Jev
- Laya
Jev minus Laya by depth, open world
Jev minus Laya by depth, closed world
Paired differences: positive means Jev was more accurate on the same items.
Accuracy by depth, open world
| Proof depth | Items | Jev | Laya | Best constant |
|---|---|---|---|---|
| depth 0 | 300 | 97.7% 96.0–99.3 | 62.7% 57.0–68.3 | 33.3% |
| depth 1 | 302 | 87.7% 83.8–91.4 | 42.1% 36.8–47.7 | 33.4% |
| depth 2 | 303 | 84.8% 80.2–88.4 | 41.3% 36.0–46.9 | 33.3% |
| depth 3 | 303 | 79.9% 75.2–84.5 | 38.3% 33.0–44.6 | 33.3% |
| depth 4 | 303 | 71.9% 66.7–77.2 | 36.0% 31.0–41.6 | 33.3% |
| depth 5 | 289 | 81.0% 76.5–85.5 | 34.3% 29.1–39.8 | 34.9% |
Accuracy by depth, closed world
| Proof depth | Items | Jev | Laya | Best constant |
|---|---|---|---|---|
| depth 0 | 300 | 99.3% 98.3–100.0 | 72.0% 66.7–76.7 | 50.0% |
| depth 1 | 300 | 93.7% 91.0–96.3 | 51.0% 45.0–56.0 | 50.0% |
| depth 2 | 300 | 92.0% 88.7–94.7 | 54.7% 48.7–59.7 | 50.0% |
| depth 3 | 300 | 85.0% 81.3–89.3 | 51.3% 45.7–56.7 | 50.0% |
| depth 4 | 300 | 76.3% 71.7–81.0 | 58.7% 53.0–64.3 | 50.0% |
| depth 5 | 300 | 89.3% 85.7–92.3 | 46.0% 40.3–51.0 | 50.0% |
Where each one is better
Open world, Jev vs Laya. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, true answers and unknown answers. Laya is not clearly ahead on any depth or answer class. Too close to call: false answers.
Closed world, Jev vs Laya. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5 and true answers. Laya is not clearly ahead on any depth or answer class. Too close to call: false answers.
A slice counts as clearly ahead when its paired 95% interval excludes zero.
True, false and unknown, open world
| Correct answer | Items | Jev | Laya | Best constant |
|---|---|---|---|---|
| false | 605 | 88.3% 85.5–90.9 | 85.1% 82.1–87.9 | one answer |
| true | 605 | 88.8% 86.1–91.1 | 40.2% 36.2–44.1 | one answer |
| unknown | 590 | 74.2% 70.7–77.6 | 1.0% 0.3–1.9 | one answer |
True, false and unknown, closed world
| Correct answer | Items | Jev | Laya | Best constant |
|---|---|---|---|---|
| false | 900 | 88.9% 86.9–90.9 | 86.4% 84.1–88.6 | one answer |
| true | 900 | 89.7% 87.7–91.6 | 24.8% 21.9–27.6 | one answer |
Paraphrased vs templated rules, open world
| Wording | Items | Jev | Laya | Best constant |
|---|---|---|---|---|
| templated | 1,298 | 86.1% 84.2–87.9 | 41.7% 39.2–44.3 | 34.5% |
| paraphrased (NatLang) | 502 | 78.1% 74.5–82.1 | 44.4% 40.2–48.8 | 35.9% |
Paraphrased vs templated rules, closed world
| Wording | Items | Jev | Laya | Best constant |
|---|---|---|---|---|
| templated | 1,306 | 94.0% 92.8–95.3 | 56.1% 53.4–58.9 | 50.2% |
| paraphrased (NatLang) | 494 | 76.7% 73.1–80.4 | 54.3% 49.8–58.9 | 50.6% |
Where Laya errs
Accuracy alone hides which answers a model gets wrong. Rows are the correct answer; columns are what the model said.
Laya, open-world task
| Correct answer | Said true | Said false | Said unknown | Recall |
|---|---|---|---|---|
| true | 243 | 354 | 8 | 40.2% |
| false | 89 | 515 | 1 | 85.1% |
| unknown | 127 | 457 | 6 | 1.0% |
Laya answered false 73.7%, true 25.5%, unknown 0.8% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.
Laya, closed-world task
| Correct answer | Said true | Said false | Recall |
|---|---|---|---|
| true | 223 | 677 | 24.8% |
| false | 122 | 778 | 86.4% |
Laya answered false 80.8%, true 19.2% of the time; the correct answers are false 50.0%, true 50.0%.
Laya across every breakdown
For each property of the problems, the weakest and strongest value on the open-world task (slices of at least 50 items). Only depth was controlled; the rest are descriptive.
| Breakdown | Weakest · strongest |
|---|---|
| Proof depth | depth 5 34.3% · depth 0 62.7% (28.4 points) |
| Gold answer | unknown 1.0% · false 85.1% (84.1 points) |
| Theory kind | attribute 41.2% · relation 44.3% (3.1 points) |
| Negation in the theory | without negation 38.8% · with negation 46.2% (7.4 points) |
| Negated statement | plain statement 38.7% · negated statement 46.2% (7.5 points) |
| Question strategy | inv-rconc 0.8% · inv-proof 85.1% (84.3 points) |
| Paraphrased rules | templated 41.7% · paraphrased (NatLang) 44.4% (2.7 points) |
| Theory's deepest proof | 0 39.4% · 1 48.3% (8.8 points) |
| Theory length | 80-109 40.3% · 0-49 47.8% (7.5 points) |
| Number of rules | 8+ 35.7% · 0-2 53.9% (18.2 points) |
| Number of facts | 4-7 41.6% · 13+ 44.0% (2.4 points) |
| Proof size | 0-1 24.3% · 2 63.6% (39.2 points) |
Consistency and calibration
Asked every open world question a second time, Laya gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 42.4% the first time and 42.4% the second.
Asked every closed world question a second time, Laya gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 55.6% the first time and 55.6% the second.
Calibration on the open world task: Laya's probability for its own answer averaged 74.7% against 42.4% accuracy; expected calibration error 32.3 points.
Calibration on the closed world task: Laya's probability for its own answer averaged 87.3% against 55.6% accuracy; expected calibration error 31.7 points.
Size, hardware and speed
| Model | Kind | Parameters | Ran on | Median time per decision | Memory peak |
|---|---|---|---|---|---|
| Jev | hosted decision model | undisclosed | vendor's servers | 189 ms | not applicable (hosted) |
| Laya | open decision model | 421M | our laptop (Apple M1 Max, 32 GB) | 140 ms | 10.0 GB |
Times are open-world, one request at a time. Hosted models were timed over the network against the vendor's servers and open models on one laptop, so speed is not like for like across the two. Speed, size and memory for every model.
Laya
- Parameters
- 421M
- Architecture
- ModernBERT-large encoder with decision heads, trained with RLCD
- Weights
- 0.84 GB (model.safetensors, convaiinnovations/laya repo root)
- Precision
- as loaded by the laya package (PyTorch)
- Where it ran
- this machine, laya 0.3.21 package, PyTorch on MPS
- Maker's stated hardware
- not stated by the maker
- Memory measured
- 10.0 GB in use, 10.0 GB peak
- Time per decision, open world
- median 140 ms, 90th percentile 479 ms, on a laptop
- Time per decision, closed world
- median 121 ms, 90th percentile 398 ms, on a laptop
The models, and what we predicted
Jev Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.
Laya Laya is an open decision model made by Convai Innovations: a 421M-parameter ModernBERT-large encoder with decision heads. We ran it on our own laptop with the laya package, PyTorch on Apple's GPU.
- Repeatability prediction 1, which names Laya: pending. See How we measured.
Other comparisons: Jev vs GPT-6 Luna · Kev vs Jev · Every model