Comparison · Kev vs Jev
Kev vs Jev: is an open-source decision model a Jev alternative?
Kev is an open decision model you can run on your own hardware; Jev is hosted. We ran every Kev size that fits a 32 GB laptop against Jev on the same problems, and asked whether a bigger Kev helps.
Overall: Jev, Kev-9B, Kev-4B and Kev-0.8B
Open world: true, false or unknown
- Jev82.2 to 85.5 (95%), 1,800 items83.8%
- Kev-9B56.3 to 60.9 (95%), 1,800 items58.6%
- Kev-4B51.4 to 56.0 (95%), 1,800 items53.6%
- Kev-0.8B51.4 to 55.9 (95%), 1,800 items53.6%
Jev minus Kev-9B: +25.3 points on the same 1,800 items, 95% interval +22.7 to +27.8: Jev is clearly ahead.
Jev minus Kev-4B: +30.3 points on the same 1,800 items, 95% interval +27.4 to +32.9: Jev is clearly ahead.
Jev minus Kev-0.8B: +30.3 points on the same 1,800 items, 95% interval +27.7 to +32.7: Jev is clearly ahead.
Closed world: true or false
- Jev87.8 to 90.7 (95%), 1,800 items89.3%
- Kev-9B61.5 to 66.1 (95%), 1,800 items63.8%
- Kev-4B56.6 to 61.1 (95%), 1,800 items58.8%
- Kev-0.8B53.5 to 57.9 (95%), 1,800 items55.8%
Jev minus Kev-9B: +25.5 points on the same 1,800 items, 95% interval +23.0 to +27.8: Jev is clearly ahead.
Jev minus Kev-4B: +30.5 points on the same 1,800 items, 95% interval +27.9 to +33.1: Jev is clearly ahead.
Jev minus Kev-0.8B: +33.5 points on the same 1,800 items, 95% interval +30.7 to +36.3: Jev is clearly ahead.
Jev vs Kev-9B, Kev-4B and Kev-0.8B by proof depth
Each point is about 300 problems that need that many inference steps, with its 95% interval.
- Jev
- Kev-9B
- Kev-4B
- Kev-0.8B
- Jev
- Kev-9B
- Kev-4B
- Kev-0.8B
Jev minus Kev-9B by depth, open world
Jev minus Kev-9B by depth, closed world
Jev minus Kev-4B by depth, open world
Jev minus Kev-4B by depth, closed world
Jev minus Kev-0.8B by depth, open world
Jev minus Kev-0.8B by depth, closed world
Paired differences: positive means Jev was more accurate on the same items.
Accuracy by depth, open world
| Proof depth | Items | Jev | Kev-9B | Kev-4B | Kev-0.8B | Best constant |
|---|---|---|---|---|---|---|
| depth 0 | 300 | 97.7% 96.0–99.3 | 89.3% 85.7–92.7 | 87.7% 84.0–91.3 | 70.3% 65.3–75.7 | 33.3% |
| depth 1 | 302 | 87.7% 83.8–91.4 | 73.8% 68.9–78.8 | 62.9% 57.3–68.2 | 54.0% 48.3–59.6 | 33.4% |
| depth 2 | 303 | 84.8% 80.2–88.4 | 58.4% 53.1–63.7 | 52.1% 46.5–57.8 | 46.9% 41.3–52.8 | 33.3% |
| depth 3 | 303 | 79.9% 75.2–84.5 | 49.5% 44.2–55.4 | 44.6% 39.3–49.8 | 48.8% 43.6–54.5 | 33.3% |
| depth 4 | 303 | 71.9% 66.7–77.2 | 38.6% 33.3–43.6 | 37.3% 31.7–42.2 | 50.5% 44.9–56.1 | 33.3% |
| depth 5 | 289 | 81.0% 76.5–85.5 | 41.2% 35.3–46.7 | 36.3% 31.1–41.9 | 50.9% 45.3–56.7 | 34.9% |
Accuracy by depth, closed world
| Proof depth | Items | Jev | Kev-9B | Kev-4B | Kev-0.8B | Best constant |
|---|---|---|---|---|---|---|
| depth 0 | 300 | 99.3% 98.3–100.0 | 89.7% 86.0–93.0 | 85.7% 81.7–89.7 | 68.0% 62.7–73.3 | 50.0% |
| depth 1 | 300 | 93.7% 91.0–96.3 | 77.0% 72.7–81.7 | 68.3% 63.0–73.7 | 59.7% 54.0–65.7 | 50.0% |
| depth 2 | 300 | 92.0% 88.7–94.7 | 64.0% 59.0–69.3 | 54.3% 49.0–59.7 | 51.3% 45.7–57.0 | 50.0% |
| depth 3 | 300 | 85.0% 81.3–89.3 | 58.7% 53.0–64.3 | 56.0% 50.7–61.3 | 53.0% 47.7–58.7 | 50.0% |
| depth 4 | 300 | 76.3% 71.7–81.0 | 50.3% 44.7–55.7 | 45.0% 39.0–50.3 | 51.0% 45.7–57.0 | 50.0% |
| depth 5 | 300 | 89.3% 85.7–92.3 | 43.0% 37.3–48.7 | 43.3% 38.3–49.0 | 51.7% 46.0–57.3 | 50.0% |
Does a bigger Kev help?
The Kev sizes are built the same way on bigger base models. Every size we have scored is listed, smallest first; a size that is not scored yet appears here when it is.
| Size | Parameters | Open world | Closed world | Depth 0, open world | Depth 5, open world | Memory peak |
|---|---|---|---|---|---|---|
| Kev-0.8B | 0.8B base | 53.6% | 55.8% | 70.3% | 50.9% | 3.2 GB |
| Kev-4B | 4B base | 53.6% | 58.8% | 87.7% | 36.3% | 17.0 GB |
| Kev-9B | 9B base | 58.6% | 63.8% | 89.3% | 41.2% | 18.0 GB |
Kev-9B against Kev-0.8B: +5.0 points on the open-world task and +8.0 points on the closed-world task overall. By depth on the open-world task the change is +19.0 points at depth 0 and −9.7 points at depth 5: the bigger model is better on easy problems and worse on the deepest ones. The harness compares each Kev with Jev on the same items, not the sizes with each other, so read the two sizes' separate intervals with care.
1 of 3 predictions about Kev-4B and Kev-9B held (1 failed, 1 pending). See How we measured.
Where each one is better
Open world, Jev vs Kev-9B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers, true answers and unknown answers. Kev-9B is not clearly ahead on any depth or answer class.
Closed world, Jev vs Kev-9B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. Kev-9B is not clearly ahead on any depth or answer class.
Open world, Jev vs Kev-4B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. Kev-4B is clearly ahead on unknown answers (+4.6 points).
Closed world, Jev vs Kev-4B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. Kev-4B is not clearly ahead on any depth or answer class.
Open world, Jev vs Kev-0.8B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, true answers and unknown answers. Kev-0.8B is not clearly ahead on any depth or answer class. Too close to call: false answers.
Closed world, Jev vs Kev-0.8B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. Kev-0.8B is not clearly ahead on any depth or answer class.
A slice counts as clearly ahead when its paired 95% interval excludes zero.
True, false and unknown, open world
| Correct answer | Items | Jev | Kev-9B | Kev-4B | Kev-0.8B | Best constant |
|---|---|---|---|---|---|---|
| false | 605 | 88.3% 85.5–90.9 | 48.6% 44.6–52.6 | 35.9% 31.9–39.5 | 89.6% 87.1–91.7 | one answer |
| true | 605 | 88.8% 86.1–91.1 | 59.2% 55.4–63.1 | 46.6% 42.8–50.6 | 58.2% 54.0–62.0 | one answer |
| unknown | 590 | 74.2% 70.7–77.6 | 68.1% 64.6–71.7 | 78.8% 75.4–81.9 | 11.9% 9.3–14.6 | one answer |
True, false and unknown, closed world
| Correct answer | Items | Jev | Kev-9B | Kev-4B | Kev-0.8B | Best constant |
|---|---|---|---|---|---|---|
| false | 900 | 88.9% 86.9–90.9 | 62.9% 59.9–65.9 | 75.3% 72.8–78.2 | 66.1% 63.0–69.3 | one answer |
| true | 900 | 89.7% 87.7–91.6 | 64.7% 61.7–67.6 | 42.2% 39.2–45.4 | 45.4% 42.1–48.6 | one answer |
Paraphrased vs templated rules, open world
| Wording | Items | Jev | Kev-9B | Kev-4B | Kev-0.8B | Best constant |
|---|---|---|---|---|---|---|
| templated | 1,298 | 86.1% 84.2–87.9 | 59.6% 57.0–62.2 | 55.9% 53.2–58.6 | 53.8% 51.2–56.4 | 34.5% |
| paraphrased (NatLang) | 502 | 78.1% 74.5–82.1 | 55.8% 51.8–60.0 | 47.6% 43.4–52.0 | 53.0% 48.8–57.2 | 35.9% |
Paraphrased vs templated rules, closed world
| Wording | Items | Jev | Kev-9B | Kev-4B | Kev-0.8B | Best constant |
|---|---|---|---|---|---|---|
| templated | 1,306 | 94.0% 92.8–95.3 | 64.8% 62.0–67.5 | 60.2% 57.7–62.7 | 56.6% 53.8–59.3 | 50.2% |
| paraphrased (NatLang) | 494 | 76.7% 73.1–80.4 | 61.1% 56.7–65.4 | 55.1% 50.4–59.5 | 53.6% 49.6–58.1 | 50.6% |
Where Kev-9B, Kev-4B and Kev-0.8B err
Accuracy alone hides which answers a model gets wrong. Rows are the correct answer; columns are what the model said.
Kev-9B, open-world task
| Correct answer | Said true | Said false | Said unknown | Recall |
|---|---|---|---|---|
| true | 358 | 37 | 210 | 59.2% |
| false | 74 | 294 | 237 | 48.6% |
| unknown | 131 | 57 | 402 | 68.1% |
Kev-9B answered false 21.6%, true 31.3%, unknown 47.2% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.
Kev-9B, closed-world task
| Correct answer | Said true | Said false | Recall |
|---|---|---|---|
| true | 582 | 318 | 64.7% |
| false | 334 | 566 | 62.9% |
Kev-9B answered false 49.1%, true 50.9% of the time; the correct answers are false 50.0%, true 50.0%.
Kev-4B, open-world task
| Correct answer | Said true | Said false | Said unknown | Recall |
|---|---|---|---|---|
| true | 282 | 12 | 311 | 46.6% |
| false | 45 | 217 | 343 | 35.9% |
| unknown | 89 | 36 | 465 | 78.8% |
Kev-4B answered false 14.7%, true 23.1%, unknown 62.2% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.
Kev-4B, closed-world task
| Correct answer | Said true | Said false | Recall |
|---|---|---|---|
| true | 380 | 520 | 42.2% |
| false | 222 | 678 | 75.3% |
Kev-4B answered false 66.6%, true 33.4% of the time; the correct answers are false 50.0%, true 50.0%.
Kev-0.8B, open-world task
| Correct answer | Said true | Said false | Said unknown | Recall |
|---|---|---|---|---|
| true | 352 | 209 | 44 | 58.2% |
| false | 59 | 542 | 4 | 89.6% |
| unknown | 150 | 370 | 70 | 11.9% |
Kev-0.8B answered false 62.3%, true 31.2%, unknown 6.6% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.
Kev-0.8B, closed-world task
| Correct answer | Said true | Said false | Recall |
|---|---|---|---|
| true | 409 | 491 | 45.4% |
| false | 305 | 595 | 66.1% |
Kev-0.8B answered false 60.3%, true 39.7% of the time; the correct answers are false 50.0%, true 50.0%.
Kev-9B, Kev-4B and Kev-0.8B across every breakdown
For each property of the problems, the weakest and strongest value on the open-world task (slices of at least 50 items). Only depth was controlled; the rest are descriptive.
| Breakdown | Kev-9B | Kev-4B | Kev-0.8B |
|---|---|---|---|
| Proof depth | depth 4 38.6% · depth 0 89.3% (50.7 points) | depth 5 36.3% · depth 0 87.7% (51.3 points) | depth 2 46.9% · depth 0 70.3% (23.5 points) |
| Gold answer | false 48.6% · unknown 68.1% (19.5 points) | false 35.9% · unknown 78.8% (42.9 points) | unknown 11.9% · false 89.6% (77.7 points) |
| Theory kind | relation 57.5% · attribute 59.3% (1.8 points) | relation 52.4% · attribute 54.4% (2.0 points) | relation 53.2% · attribute 53.8% (0.7 points) |
| Negation in the theory | without negation 56.8% · with negation 60.4% (3.6 points) | without negation 49.9% · with negation 57.3% (7.4 points) | without negation 50.6% · with negation 56.6% (6.0 points) |
| Negated statement | negated statement 57.4% · plain statement 59.7% (2.4 points) | negated statement 53.5% · plain statement 53.6% (0.1 points) | plain statement 48.5% · negated statement 58.6% (10.1 points) |
| Question strategy | inv-proof 48.6% · random 78.0% (29.4 points) | inv-proof 35.9% · random 88.0% (52.1 points) | inv-rconc 0.8% · inv-proof 89.6% (88.8 points) |
| Paraphrased rules | paraphrased (NatLang) 55.8% · templated 59.6% (3.9 points) | paraphrased (NatLang) 47.6% · templated 55.9% (8.2 points) | paraphrased (NatLang) 53.0% · templated 53.8% (0.8 points) |
| Theory's deepest proof | 4 42.6% · 0 87.3% (44.7 points) | 5 38.0% · 0 85.9% (47.9 points) | 0 42.3% · 5 58.0% (15.7 points) |
| Theory length | 110+ 47.9% · 0-49 77.7% (29.8 points) | 110+ 41.7% · 0-49 75.7% (34.0 points) | 110+ 52.6% · 0-49 56.3% (3.7 points) |
| Number of rules | 3-5 51.6% · 0-2 78.9% (27.3 points) | 3-5 45.9% · 0-2 78.9% (33.0 points) | 8+ 49.8% · 0-2 58.8% (9.0 points) |
| Number of facts | 13+ 46.5% · 0-3 71.9% (25.3 points) | 13+ 40.3% · 0-3 69.8% (29.5 points) | 0-3 52.4% · 13+ 56.6% (4.2 points) |
| Proof size | 5+ 31.8% · 2 84.1% (52.3 points) | 5+ 20.3% · 0-1 81.4% (61.1 points) | 0-1 32.0% · 2 77.6% (45.5 points) |
Consistency and calibration
Kev-9B
Kev-9B's rerun has not been scored yet. It is preregistered to agree with itself on at least 99.5% of items, since its weights are fixed and it runs locally.
Calibration on the open world task: Kev-9B's probability for its own answer averaged 74.3% against 58.6% accuracy; expected calibration error 15.8 points.
Calibration on the closed world task: Kev-9B's probability for its own answer averaged 77.3% against 63.8% accuracy; expected calibration error 13.5 points.
Kev-4B
Asked every open world question a second time, Kev-4B gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 53.6% the first time and 53.6% the second.
Asked every closed world question a second time, Kev-4B gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 58.8% the first time and 58.8% the second.
Calibration on the open world task: Kev-4B's probability for its own answer averaged 69.9% against 53.6% accuracy; expected calibration error 16.4 points.
Calibration on the closed world task: Kev-4B's probability for its own answer averaged 66.4% against 58.8% accuracy; expected calibration error 7.6 points.
Kev-0.8B
Asked every open world question a second time, Kev-0.8B gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 53.6% the first time and 53.6% the second.
Asked every closed world question a second time, Kev-0.8B gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 55.8% the first time and 55.8% the second.
Calibration on the open world task: Kev-0.8B's probability for its own answer averaged 65.0% against 53.6% accuracy; expected calibration error 11.4 points.
Calibration on the closed world task: Kev-0.8B's probability for its own answer averaged 72.8% against 55.8% accuracy; expected calibration error 17.0 points.
Size, hardware and speed
| Model | Kind | Parameters | Ran on | Median time per decision | Memory peak |
|---|---|---|---|---|---|
| Jev | hosted decision model | undisclosed | vendor's servers | 189 ms | not applicable (hosted) |
| Kev-9B | open decision model | 9B base (Qwen3.5-9B-Base) + rank-16 LoRA adapter and pointer head | our laptop (Apple M1 Max, 32 GB) | not yet timed | 18.0 GB |
| Kev-4B | open decision model | 4B base (Qwen3.5-4B-Base) + rank-16 LoRA adapter (33.8M trainable) and pointer head | our laptop (Apple M1 Max, 32 GB) | 897 ms | 17.0 GB |
| Kev-0.8B | open decision model | 0.8B base (Qwen3.5-0.8B-Base) + rank-16 LoRA adapter and pointer head | our laptop (Apple M1 Max, 32 GB) | 123 ms | 3.2 GB |
Times are open-world, one request at a time. Hosted models were timed over the network against the vendor's servers and open models on one laptop, so speed is not like for like across the two. Speed, size and memory for every model.
Kev-9B
- Parameters
- 9B base (Qwen3.5-9B-Base) + rank-16 LoRA adapter and pointer head
- Architecture
- as kev-0.8b
- Weights
- 19.31 GB base (Qwen/Qwen3.5-9B-Base@68c46c4b) + 0.18 GB adapter and head (jaredpalmer/kev-9b@b5d8c18e)
- Precision
- bfloat16 (MLX), fp32 pointer head
- Where it ran
- this machine, Kev server c9c1f855 on MLX
- Maker's stated hardware
- "32 GB Mac, L40S, H100" (Kev README)
- Memory measured
- 17.0 GB in use, 18.0 GB peak
Kev-4B
- Parameters
- 4B base (Qwen3.5-4B-Base) + rank-16 LoRA adapter (33.8M trainable) and pointer head
- Architecture
- as kev-0.8b
- Weights
- 9.32 GB base (Qwen/Qwen3.5-4B-Base@1001bb4d) + 0.14 GB adapter and head (jaredpalmer/kev-4b@139fdd94)
- Precision
- bfloat16 (MLX), fp32 pointer head
- Where it ran
- this machine, Kev server c9c1f855 on MLX
- Maker's stated hardware
- "32 GB Mac, L40S, H100" (Kev README)
- Memory measured
- 8.6 GB in use, 17.0 GB peak
- Time per decision, open world
- median 897 ms, 90th percentile 1,100 ms, on a laptop
- Time per decision, closed world
- median 802 ms, 90th percentile 1,044 ms, on a laptop
Kev-0.8B
- Parameters
- 0.8B base (Qwen3.5-0.8B-Base) + rank-16 LoRA adapter and pointer head
- Architecture
- Qwen3.5 hybrid (Gated DeltaNet + attention) base, frozen; LoRA r16; pointer head
- Weights
- 1.75 GB base (Qwen/Qwen3.5-0.8B-Base@dc7cdfe2) + 0.05 GB adapter and head (jaredpalmer/kev-0.8b@54f4f877)
- Precision
- bfloat16 (MLX), fp32 pointer head
- Where it ran
- this machine, Kev server c9c1f855 on MLX
- Maker's stated hardware
- "Any Apple Silicon Mac, L4" (Kev README)
- Memory measured
- 2.1 GB in use, 3.2 GB peak
- Time per decision, open world
- median 123 ms, 90th percentile 152 ms, on a laptop
- Time per decision, closed world
- median 121 ms, 90th percentile 145 ms, on a laptop
The models, and what we predicted
Jev Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.
Kev-9B Kev-9B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 9-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.
Kev-4B Kev-4B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 4-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.
Kev-0.8B Kev-0.8B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 0.8-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.
- 3 of 3 predictions about Kev-0.8B held. See How we measured.
- 1 of 3 predictions about Kev-4B and Kev-9B held (1 failed, 1 pending). See How we measured.
- Repeatability prediction 1, which names Kev: pending. See How we measured.
Other comparisons: Jev vs GPT-6 Luna · Jev vs Laya · Every model