Comparison · Kev vs Jev

Kev vs Jev: is an open-source decision model a Jev alternative?

Kev is an open decision model you can run on your own hardware; Jev is hosted. We ran every Kev size that fits a 32 GB laptop against Jev on the same problems, and asked whether a bigger Kev helps.

Overall: Jev, Kev-9B, Kev-4B and Kev-0.8B

Open world: true, false or unknown

Jev minus Kev-9B: +25.3 points on the same 1,800 items, 95% interval +22.7 to +27.8: Jev is clearly ahead.

Jev minus Kev-4B: +30.3 points on the same 1,800 items, 95% interval +27.4 to +32.9: Jev is clearly ahead.

Jev minus Kev-0.8B: +30.3 points on the same 1,800 items, 95% interval +27.7 to +32.7: Jev is clearly ahead.

Closed world: true or false

Jev minus Kev-9B: +25.5 points on the same 1,800 items, 95% interval +23.0 to +27.8: Jev is clearly ahead.

Jev minus Kev-4B: +30.5 points on the same 1,800 items, 95% interval +27.9 to +33.1: Jev is clearly ahead.

Jev minus Kev-0.8B: +33.5 points on the same 1,800 items, 95% interval +30.7 to +36.3: Jev is clearly ahead.

Jev vs Kev-9B, Kev-4B and Kev-0.8B by proof depth

Each point is about 300 problems that need that many inference steps, with its 95% interval.

Open world: true, false or unknown
  • Jev
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
0%25%50%75%100%012345proof depth (inference steps)chance 33%
Closed world: true or false
  • Jev
  • Kev-9B
  • Kev-4B
  • Kev-0.8B
0%25%50%75%100%012345proof depth (inference steps)chance 50%

Jev minus Kev-9B by depth, open world

depth 0300 items
+8.3
depth 1302 items
+13.9
depth 2303 items
+26.4
depth 3303 items
+30.4
depth 4303 items
+33.3
depth 5289 items
+39.8

Jev minus Kev-9B by depth, closed world

depth 0300 items
+9.7
depth 1300 items
+16.7
depth 2300 items
+28.0
depth 3300 items
+26.3
depth 4300 items
+26.0
depth 5300 items
+46.3

Jev minus Kev-4B by depth, open world

depth 0300 items
+10.0
depth 1302 items
+24.8
depth 2303 items
+32.7
depth 3303 items
+35.3
depth 4303 items
+34.7
depth 5289 items
+44.6

Jev minus Kev-4B by depth, closed world

depth 0300 items
+13.7
depth 1300 items
+25.3
depth 2300 items
+37.7
depth 3300 items
+29.0
depth 4300 items
+31.3
depth 5300 items
+46.0

Jev minus Kev-0.8B by depth, open world

depth 0300 items
+27.3
depth 1302 items
+33.8
depth 2303 items
+38.0
depth 3303 items
+31.0
depth 4303 items
+21.5
depth 5289 items
+30.1

Jev minus Kev-0.8B by depth, closed world

depth 0300 items
+31.3
depth 1300 items
+34.0
depth 2300 items
+40.7
depth 3300 items
+32.0
depth 4300 items
+25.3
depth 5300 items
+37.7

Paired differences: positive means Jev was more accurate on the same items.

Accuracy by depth, open world

By depth, Open world
Proof depthItemsJevKev-9BKev-4BKev-0.8BBest constant
depth 030097.7% 96.0–99.389.3% 85.7–92.787.7% 84.0–91.370.3% 65.3–75.733.3%
depth 130287.7% 83.8–91.473.8% 68.9–78.862.9% 57.3–68.254.0% 48.3–59.633.4%
depth 230384.8% 80.2–88.458.4% 53.1–63.752.1% 46.5–57.846.9% 41.3–52.833.3%
depth 330379.9% 75.2–84.549.5% 44.2–55.444.6% 39.3–49.848.8% 43.6–54.533.3%
depth 430371.9% 66.7–77.238.6% 33.3–43.637.3% 31.7–42.250.5% 44.9–56.133.3%
depth 528981.0% 76.5–85.541.2% 35.3–46.736.3% 31.1–41.950.9% 45.3–56.734.9%

Accuracy by depth, closed world

By depth, Closed world
Proof depthItemsJevKev-9BKev-4BKev-0.8BBest constant
depth 030099.3% 98.3–100.089.7% 86.0–93.085.7% 81.7–89.768.0% 62.7–73.350.0%
depth 130093.7% 91.0–96.377.0% 72.7–81.768.3% 63.0–73.759.7% 54.0–65.750.0%
depth 230092.0% 88.7–94.764.0% 59.0–69.354.3% 49.0–59.751.3% 45.7–57.050.0%
depth 330085.0% 81.3–89.358.7% 53.0–64.356.0% 50.7–61.353.0% 47.7–58.750.0%
depth 430076.3% 71.7–81.050.3% 44.7–55.745.0% 39.0–50.351.0% 45.7–57.050.0%
depth 530089.3% 85.7–92.343.0% 37.3–48.743.3% 38.3–49.051.7% 46.0–57.350.0%

Does a bigger Kev help?

The Kev sizes are built the same way on bigger base models. Every size we have scored is listed, smallest first; a size that is not scored yet appears here when it is.

SizeParametersOpen worldClosed worldDepth 0, open worldDepth 5, open worldMemory peak
Kev-0.8B0.8B base53.6%55.8%70.3%50.9%3.2 GB
Kev-4B4B base53.6%58.8%87.7%36.3%17.0 GB
Kev-9B9B base58.6%63.8%89.3%41.2%18.0 GB

Kev-9B against Kev-0.8B: +5.0 points on the open-world task and +8.0 points on the closed-world task overall. By depth on the open-world task the change is +19.0 points at depth 0 and −9.7 points at depth 5: the bigger model is better on easy problems and worse on the deepest ones. The harness compares each Kev with Jev on the same items, not the sizes with each other, so read the two sizes' separate intervals with care.

1 of 3 predictions about Kev-4B and Kev-9B held (1 failed, 1 pending). See How we measured.

Where each one is better

Open world, Jev vs Kev-9B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers, true answers and unknown answers. Kev-9B is not clearly ahead on any depth or answer class.

Closed world, Jev vs Kev-9B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. Kev-9B is not clearly ahead on any depth or answer class.

Open world, Jev vs Kev-4B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. Kev-4B is clearly ahead on unknown answers (+4.6 points).

Closed world, Jev vs Kev-4B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. Kev-4B is not clearly ahead on any depth or answer class.

Open world, Jev vs Kev-0.8B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, true answers and unknown answers. Kev-0.8B is not clearly ahead on any depth or answer class. Too close to call: false answers.

Closed world, Jev vs Kev-0.8B. Jev is clearly ahead on depth 0, depth 1, depth 2, depth 3, depth 4, depth 5, false answers and true answers. Kev-0.8B is not clearly ahead on any depth or answer class.

A slice counts as clearly ahead when its paired 95% interval excludes zero.

True, false and unknown, open world

By correct answer, Open world
Correct answerItemsJevKev-9BKev-4BKev-0.8BBest constant
false60588.3% 85.5–90.948.6% 44.6–52.635.9% 31.9–39.589.6% 87.1–91.7one answer
true60588.8% 86.1–91.159.2% 55.4–63.146.6% 42.8–50.658.2% 54.0–62.0one answer
unknown59074.2% 70.7–77.668.1% 64.6–71.778.8% 75.4–81.911.9% 9.3–14.6one answer

True, false and unknown, closed world

By correct answer, Closed world
Correct answerItemsJevKev-9BKev-4BKev-0.8BBest constant
false90088.9% 86.9–90.962.9% 59.9–65.975.3% 72.8–78.266.1% 63.0–69.3one answer
true90089.7% 87.7–91.664.7% 61.7–67.642.2% 39.2–45.445.4% 42.1–48.6one answer

Paraphrased vs templated rules, open world

By wording, Open world
WordingItemsJevKev-9BKev-4BKev-0.8BBest constant
templated1,29886.1% 84.2–87.959.6% 57.0–62.255.9% 53.2–58.653.8% 51.2–56.434.5%
paraphrased (NatLang)50278.1% 74.5–82.155.8% 51.8–60.047.6% 43.4–52.053.0% 48.8–57.235.9%

Paraphrased vs templated rules, closed world

By wording, Closed world
WordingItemsJevKev-9BKev-4BKev-0.8BBest constant
templated1,30694.0% 92.8–95.364.8% 62.0–67.560.2% 57.7–62.756.6% 53.8–59.350.2%
paraphrased (NatLang)49476.7% 73.1–80.461.1% 56.7–65.455.1% 50.4–59.553.6% 49.6–58.150.6%

Where Kev-9B, Kev-4B and Kev-0.8B err

Accuracy alone hides which answers a model gets wrong. Rows are the correct answer; columns are what the model said.

Kev-9B, open-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseSaid unknownRecall
true3583721059.2%
false7429423748.6%
unknown1315740268.1%

Kev-9B answered false 21.6%, true 31.3%, unknown 47.2% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.

Kev-9B, closed-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseRecall
true58231864.7%
false33456662.9%

Kev-9B answered false 49.1%, true 50.9% of the time; the correct answers are false 50.0%, true 50.0%.

Kev-4B, open-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseSaid unknownRecall
true2821231146.6%
false4521734335.9%
unknown893646578.8%

Kev-4B answered false 14.7%, true 23.1%, unknown 62.2% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.

Kev-4B, closed-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseRecall
true38052042.2%
false22267875.3%

Kev-4B answered false 66.6%, true 33.4% of the time; the correct answers are false 50.0%, true 50.0%.

Kev-0.8B, open-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseSaid unknownRecall
true3522094458.2%
false59542489.6%
unknown1503707011.9%

Kev-0.8B answered false 62.3%, true 31.2%, unknown 6.6% of the time; the correct answers are false 33.6%, true 33.6%, unknown 32.8%.

Kev-0.8B, closed-world task

Correct answer by what the model said
Correct answerSaid trueSaid falseRecall
true40949145.4%
false30559566.1%

Kev-0.8B answered false 60.3%, true 39.7% of the time; the correct answers are false 50.0%, true 50.0%.

Kev-9B, Kev-4B and Kev-0.8B across every breakdown

For each property of the problems, the weakest and strongest value on the open-world task (slices of at least 50 items). Only depth was controlled; the rest are descriptive.

BreakdownKev-9BKev-4BKev-0.8B
Proof depthdepth 4 38.6% · depth 0 89.3% (50.7 points)depth 5 36.3% · depth 0 87.7% (51.3 points)depth 2 46.9% · depth 0 70.3% (23.5 points)
Gold answerfalse 48.6% · unknown 68.1% (19.5 points)false 35.9% · unknown 78.8% (42.9 points)unknown 11.9% · false 89.6% (77.7 points)
Theory kindrelation 57.5% · attribute 59.3% (1.8 points)relation 52.4% · attribute 54.4% (2.0 points)relation 53.2% · attribute 53.8% (0.7 points)
Negation in the theorywithout negation 56.8% · with negation 60.4% (3.6 points)without negation 49.9% · with negation 57.3% (7.4 points)without negation 50.6% · with negation 56.6% (6.0 points)
Negated statementnegated statement 57.4% · plain statement 59.7% (2.4 points)negated statement 53.5% · plain statement 53.6% (0.1 points)plain statement 48.5% · negated statement 58.6% (10.1 points)
Question strategyinv-proof 48.6% · random 78.0% (29.4 points)inv-proof 35.9% · random 88.0% (52.1 points)inv-rconc 0.8% · inv-proof 89.6% (88.8 points)
Paraphrased rulesparaphrased (NatLang) 55.8% · templated 59.6% (3.9 points)paraphrased (NatLang) 47.6% · templated 55.9% (8.2 points)paraphrased (NatLang) 53.0% · templated 53.8% (0.8 points)
Theory's deepest proof4 42.6% · 0 87.3% (44.7 points)5 38.0% · 0 85.9% (47.9 points)0 42.3% · 5 58.0% (15.7 points)
Theory length110+ 47.9% · 0-49 77.7% (29.8 points)110+ 41.7% · 0-49 75.7% (34.0 points)110+ 52.6% · 0-49 56.3% (3.7 points)
Number of rules3-5 51.6% · 0-2 78.9% (27.3 points)3-5 45.9% · 0-2 78.9% (33.0 points)8+ 49.8% · 0-2 58.8% (9.0 points)
Number of facts13+ 46.5% · 0-3 71.9% (25.3 points)13+ 40.3% · 0-3 69.8% (29.5 points)0-3 52.4% · 13+ 56.6% (4.2 points)
Proof size5+ 31.8% · 2 84.1% (52.3 points)5+ 20.3% · 0-1 81.4% (61.1 points)0-1 32.0% · 2 77.6% (45.5 points)

Consistency and calibration

Kev-9B

Kev-9B's rerun has not been scored yet. It is preregistered to agree with itself on at least 99.5% of items, since its weights are fixed and it runs locally.

Calibration on the open world task: Kev-9B's probability for its own answer averaged 74.3% against 58.6% accuracy; expected calibration error 15.8 points.

Calibration on the closed world task: Kev-9B's probability for its own answer averaged 77.3% against 63.8% accuracy; expected calibration error 13.5 points.

Kev-4B

Asked every open world question a second time, Kev-4B gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 53.6% the first time and 53.6% the second.

Asked every closed world question a second time, Kev-4B gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 58.8% the first time and 58.8% the second.

Calibration on the open world task: Kev-4B's probability for its own answer averaged 69.9% against 53.6% accuracy; expected calibration error 16.4 points.

Calibration on the closed world task: Kev-4B's probability for its own answer averaged 66.4% against 58.8% accuracy; expected calibration error 7.6 points.

Kev-0.8B

Asked every open world question a second time, Kev-0.8B gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 53.6% the first time and 53.6% the second.

Asked every closed world question a second time, Kev-0.8B gave the same answer on 100.0% of 1,800 items (0 changed; Gwet's AC1 1.000, interval 1.000 to 1.000). Accuracy was 55.8% the first time and 55.8% the second.

Calibration on the open world task: Kev-0.8B's probability for its own answer averaged 65.0% against 53.6% accuracy; expected calibration error 11.4 points.

Calibration on the closed world task: Kev-0.8B's probability for its own answer averaged 72.8% against 55.8% accuracy; expected calibration error 17.0 points.

Size, hardware and speed

ModelKindParametersRan onMedian time per decisionMemory peak
Jevhosted decision modelundisclosedvendor's servers189 msnot applicable (hosted)
Kev-9Bopen decision model9B base (Qwen3.5-9B-Base) + rank-16 LoRA adapter and pointer headour laptop (Apple M1 Max, 32 GB)not yet timed18.0 GB
Kev-4Bopen decision model4B base (Qwen3.5-4B-Base) + rank-16 LoRA adapter (33.8M trainable) and pointer headour laptop (Apple M1 Max, 32 GB)897 ms17.0 GB
Kev-0.8Bopen decision model0.8B base (Qwen3.5-0.8B-Base) + rank-16 LoRA adapter and pointer headour laptop (Apple M1 Max, 32 GB)123 ms3.2 GB

Times are open-world, one request at a time. Hosted models were timed over the network against the vendor's servers and open models on one laptop, so speed is not like for like across the two. Speed, size and memory for every model.

Kev-9B

Parameters
9B base (Qwen3.5-9B-Base) + rank-16 LoRA adapter and pointer head
Architecture
as kev-0.8b
Weights
19.31 GB base (Qwen/Qwen3.5-9B-Base@68c46c4b) + 0.18 GB adapter and head (jaredpalmer/kev-9b@b5d8c18e)
Precision
bfloat16 (MLX), fp32 pointer head
Where it ran
this machine, Kev server c9c1f855 on MLX
Maker's stated hardware
"32 GB Mac, L40S, H100" (Kev README)
Memory measured
17.0 GB in use, 18.0 GB peak
Source
huggingface.co/jaredpalmer/kev-9b

Kev-4B

Parameters
4B base (Qwen3.5-4B-Base) + rank-16 LoRA adapter (33.8M trainable) and pointer head
Architecture
as kev-0.8b
Weights
9.32 GB base (Qwen/Qwen3.5-4B-Base@1001bb4d) + 0.14 GB adapter and head (jaredpalmer/kev-4b@139fdd94)
Precision
bfloat16 (MLX), fp32 pointer head
Where it ran
this machine, Kev server c9c1f855 on MLX
Maker's stated hardware
"32 GB Mac, L40S, H100" (Kev README)
Memory measured
8.6 GB in use, 17.0 GB peak
Time per decision, open world
median 897 ms, 90th percentile 1,100 ms, on a laptop
Time per decision, closed world
median 802 ms, 90th percentile 1,044 ms, on a laptop
Source
huggingface.co/jaredpalmer/kev-4b

Kev-0.8B

Parameters
0.8B base (Qwen3.5-0.8B-Base) + rank-16 LoRA adapter and pointer head
Architecture
Qwen3.5 hybrid (Gated DeltaNet + attention) base, frozen; LoRA r16; pointer head
Weights
1.75 GB base (Qwen/Qwen3.5-0.8B-Base@dc7cdfe2) + 0.05 GB adapter and head (jaredpalmer/kev-0.8b@54f4f877)
Precision
bfloat16 (MLX), fp32 pointer head
Where it ran
this machine, Kev server c9c1f855 on MLX
Maker's stated hardware
"Any Apple Silicon Mac, L4" (Kev README)
Memory measured
2.1 GB in use, 3.2 GB peak
Time per decision, open world
median 123 ms, 90th percentile 152 ms, on a laptop
Time per decision, closed world
median 121 ms, 90th percentile 145 ms, on a laptop
Source
github.com/jaredpalmer/kev

The models, and what we predicted

Jev Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.

Kev-9B Kev-9B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 9-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.

Kev-4B Kev-4B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 4-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.

Kev-0.8B Kev-0.8B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 0.8-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.