Models · comparisons
Every decision model and LLM classifier we tested, compared with Jev
Each model answered the same 1,800 problems on each task, one request per problem. Here is how they compare overall and head to head with Jev, with what each one is and where it ran. Every rival has one page: its comparison with Jev.
Jev and its rivals, one page each
The reference model: every rival is compared with it on the same items.
Jev is built to make one decision per request. GPT-6 Luna is a general large language model, used here as a classifier with reasoning off. Both answered the same 3,600 problems. OpenAI's new Decisions API is built on a version of GPT-6 Luna, so these results preview the accuracy of a Luna-based decision model on multi-step reasoning. They are not a test of the API itself.
Kev is an open decision model you can run on your own hardware; Jev is hosted. We ran every Kev size that fits a 32 GB laptop against Jev on the same problems, and asked whether a bigger Kev helps.
Laya is a small open decision model built on an encoder. Both it and Jev answered the same problems.
GPT-6 Luna is the model behind OpenAI's new Decisions API: a preview of the OpenAI Decisions API through GPT-6 Luna, accuracy and confidence.
The models
Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.
GPT-6 Luna is a hosted large language model made by OpenAI. We used it as a classifier: one request per item through the Chat Completions API, with reasoning effort set to none (its lowest setting) and a strict JSON schema that allows only the task's options. Its size and architecture are not published.
- Kev-9Bopen decision model · ran on our laptop (Apple M1 Max, 32 GB)open world 58.6% · closed world 63.8%
Kev-9B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 9-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.
- Kev-4Bopen decision model · ran on our laptop (Apple M1 Max, 32 GB)open world 53.6% · closed world 58.8%
Kev-4B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 4-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.
- Kev-0.8Bopen decision model · ran on our laptop (Apple M1 Max, 32 GB)open world 53.6% · closed world 55.8%
Kev-0.8B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 0.8-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.
- Layaopen decision model · ran on our laptop (Apple M1 Max, 32 GB)open world 42.4% · closed world 55.6%
Laya is an open decision model made by Convai Innovations: a 421M-parameter ModernBERT-large encoder with decision heads. We ran it on our own laptop with the laya package, PyTorch on Apple's GPU.
- Kev-27BNot run
Kev-27B is the largest Kev: 27B, every weight fine-tuned from Qwen3.8-27B (post-trained), with 51.26 GB full weights (jaredpalmer/kev-27b). It was not run: does not fit this 32 GB machine; would need a rented 80 GB GPU.
Accuracy side by side
Open world: true, false or unknown
| Model | Accuracy | 95% interval | Macro-F1 | Recall: true | Recall: false | Recall: unknown | Items |
|---|---|---|---|---|---|---|---|
| Jev | 83.8% | 82.2–85.5 | 0.838 | 88.8% | 88.3% | 74.2% | 1,800 |
| GPT-6 Luna | 64.1% | 61.7–66.2 | 0.647 | 58.3% | 50.9% | 83.4% | 1,800 |
| Kev-9B | 58.6% | 56.3–60.9 | 0.588 | 59.2% | 48.6% | 68.1% | 1,800 |
| Kev-4B | 53.6% | 51.4–56.0 | 0.532 | 46.6% | 35.9% | 78.8% | 1,800 |
| Kev-0.8B | 53.6% | 51.4–55.9 | 0.477 | 58.2% | 89.6% | 11.9% | 1,800 |
| Laya | 42.4% | 40.3–44.5 | 0.337 | 40.2% | 85.1% | 1.0% | 1,800 |
Chance is 33.3%; always giving the most common answer scores 33.6% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.
Closed world: true or false
| Model | Accuracy | 95% interval | Macro-F1 | Recall: true | Recall: false | Items |
|---|---|---|---|---|---|---|
| Jev | 89.3% | 87.8–90.7 | 0.893 | 89.7% | 88.9% | 1,800 |
| GPT-6 Luna | 65.0% | 62.8–67.4 | 0.650 | 66.2% | 63.8% | 1,800 |
| Kev-9B | 63.8% | 61.5–66.1 | 0.638 | 64.7% | 62.9% | 1,800 |
| Kev-4B | 58.8% | 56.6–61.1 | 0.576 | 42.2% | 75.3% | 1,800 |
| Kev-0.8B | 55.8% | 53.5–57.9 | 0.553 | 45.4% | 66.1% | 1,800 |
| Laya | 55.6% | 53.3–58.1 | 0.509 | 24.8% | 86.4% | 1,800 |
Chance is 50.0%; always giving the most common answer scores 50.0% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.
Jev against every model, on the same items
Jev minus each model, in points, with the 95% interval of the paired difference. Positive means Jev was more accurate. Because both answered exactly the same items, the interval is narrower than comparing two separate scores.
| Model | Open world: overall | depth 0 | depth 5 | Closed world: overall | depth 0 | depth 5 |
|---|---|---|---|---|---|---|
| GPT-6 Luna | +19.8 +17.2 to +22.4 | +4.3 +1.0 to +7.7 | +34.9 +27.3 to +42.9 | +24.3 +21.9 to +26.6 | +1.3 0.0 to +3.0 | +44.7 +38.0 to +51.3 |
| Kev-9B | +25.3 +22.7 to +27.8 | +8.3 +5.3 to +11.7 | +39.8 +32.9 to +47.8 | +25.5 +23.0 to +27.8 | +9.7 +6.3 to +13.3 | +46.3 +40.3 to +52.3 |
| Kev-4B | +30.3 +27.4 to +32.9 | +10.0 +6.3 to +13.7 | +44.6 +37.0 to +51.9 | +30.5 +27.9 to +33.1 | +13.7 +9.7 to +18.0 | +46.0 +39.7 to +52.3 |
| Kev-0.8B | +30.3 +27.7 to +32.7 | +27.3 +21.7 to +32.7 | +30.1 +23.5 to +36.0 | +33.5 +30.7 to +36.3 | +31.3 +26.0 to +36.7 | +37.7 +31.0 to +44.0 |
| Laya | +41.4 +38.8 to +44.0 | +35.0 +29.3 to +40.7 | +46.7 +40.1 to +52.9 | +33.7 +30.7 to +36.5 | +27.3 +22.3 to +33.0 | +43.3 +37.0 to +50.3 |