Models · comparisons

Every decision model and LLM classifier we tested, compared with Jev

Each model answered the same 1,800 problems on each task, one request per problem. Here is how they compare overall and head to head with Jev, with what each one is and where it ran. Every rival has one page: its comparison with Jev.

Jev and its rivals, one page each

The models

  1. Jev
    hosted decision model · ran on vendor's servers
    open world 83.8% · closed world 89.3%

    Jev is a hosted decision model made by TypeSafe. We call it through its maker's software kit, typesafe-sdk, on paid API access; the version is recorded on every answer (jev-1.13.0). Its size and architecture are not published.

  2. GPT-6 Luna
    LLM classifier · ran on vendor's servers
    open world 64.1% · closed world 65.0%

    GPT-6 Luna is a hosted large language model made by OpenAI. We used it as a classifier: one request per item through the Chat Completions API, with reasoning effort set to none (its lowest setting) and a strict JSON schema that allows only the task's options. Its size and architecture are not published.

  3. Kev-9B
    open decision model · ran on our laptop (Apple M1 Max, 32 GB)
    open world 58.6% · closed world 63.8%

    Kev-9B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 9-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.

  4. Kev-4B
    open decision model · ran on our laptop (Apple M1 Max, 32 GB)
    open world 53.6% · closed world 58.8%

    Kev-4B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 4-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.

  5. Kev-0.8B
    open decision model · ran on our laptop (Apple M1 Max, 32 GB)
    open world 53.6% · closed world 55.8%

    Kev-0.8B is an open decision model made by Jared Palmer: a pointer head and a rank-16 LoRA adapter on a frozen 0.8-billion-parameter Qwen3.5 base model. We ran it on our own laptop through the Kev server on MLX, in bfloat16.

  6. Laya
    open decision model · ran on our laptop (Apple M1 Max, 32 GB)
    open world 42.4% · closed world 55.6%

    Laya is an open decision model made by Convai Innovations: a 421M-parameter ModernBERT-large encoder with decision heads. We ran it on our own laptop with the laya package, PyTorch on Apple's GPU.

  7. Kev-27B
    Not run

    Kev-27B is the largest Kev: 27B, every weight fine-tuned from Qwen3.8-27B (post-trained), with 51.26 GB full weights (jaredpalmer/kev-27b). It was not run: does not fit this 32 GB machine; would need a rented 80 GB GPU.

Accuracy side by side

Open world: true, false or unknown

Overall accuracy, Open world: true, false or unknown
ModelAccuracy95% intervalMacro-F1Recall: trueRecall: falseRecall: unknownItems
Jev83.8%82.2–85.50.83888.8%88.3%74.2%1,800
GPT-6 Luna64.1%61.7–66.20.64758.3%50.9%83.4%1,800
Kev-9B58.6%56.3–60.90.58859.2%48.6%68.1%1,800
Kev-4B53.6%51.4–56.00.53246.6%35.9%78.8%1,800
Kev-0.8B53.6%51.4–55.90.47758.2%89.6%11.9%1,800
Laya42.4%40.3–44.50.33740.2%85.1%1.0%1,800

Chance is 33.3%; always giving the most common answer scores 33.6% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.

Closed world: true or false

Overall accuracy, Closed world: true or false
ModelAccuracy95% intervalMacro-F1Recall: trueRecall: falseItems
Jev89.3%87.8–90.70.89389.7%88.9%1,800
GPT-6 Luna65.0%62.8–67.40.65066.2%63.8%1,800
Kev-9B63.8%61.5–66.10.63864.7%62.9%1,800
Kev-4B58.8%56.6–61.10.57642.2%75.3%1,800
Kev-0.8B55.8%53.5–57.90.55345.4%66.1%1,800
Laya55.6%53.3–58.10.50924.8%86.4%1,800

Chance is 50.0%; always giving the most common answer scores 50.0% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.

Jev against every model, on the same items

Jev minus each model, in points, with the 95% interval of the paired difference. Positive means Jev was more accurate. Because both answered exactly the same items, the interval is narrower than comparing two separate scores.

ModelOpen world: overalldepth 0depth 5Closed world: overalldepth 0depth 5
GPT-6 Luna+19.8 +17.2 to +22.4+4.3 +1.0 to +7.7+34.9 +27.3 to +42.9+24.3 +21.9 to +26.6+1.3 0.0 to +3.0+44.7 +38.0 to +51.3
Kev-9B+25.3 +22.7 to +27.8+8.3 +5.3 to +11.7+39.8 +32.9 to +47.8+25.5 +23.0 to +27.8+9.7 +6.3 to +13.3+46.3 +40.3 to +52.3
Kev-4B+30.3 +27.4 to +32.9+10.0 +6.3 to +13.7+44.6 +37.0 to +51.9+30.5 +27.9 to +33.1+13.7 +9.7 to +18.0+46.0 +39.7 to +52.3
Kev-0.8B+30.3 +27.7 to +32.7+27.3 +21.7 to +32.7+30.1 +23.5 to +36.0+33.5 +30.7 to +36.3+31.3 +26.0 to +36.7+37.7 +31.0 to +44.0
Laya+41.4 +38.8 to +44.0+35.0 +29.3 to +40.7+46.7 +40.1 to +52.9+33.7 +30.7 to +36.5+27.3 +22.3 to +33.0+43.3 +37.0 to +50.3