Guide · fine-tuning decision models
Fine-tuning decision models: where they fail and what to train on
The failures in this benchmark are specific enough to plan training data around: deep proofs, the unknown answer, reworded rules and a lean toward one answer. Here is what the results show, model by model.
This page describes failures we measured on ProofWriter, a synthetic dataset. It shows what to look for and test, not what any model will do on your decisions.
Failure 1: accuracy collapses past a few inference steps
On the open-world task, chance is a third. Here is each model at depth 3 and depth 5:
- Jevdepth 3: 79.9% · depth 5: 81.0%81.0%
- GPT-6 Lunadepth 3: 61.7% · depth 5: 46.0%46.0%
- Kev-9Bdepth 3: 49.5% · depth 5: 41.2% · interval reaches chance from depth 441.2%
- Kev-4Bdepth 3: 44.6% · depth 5: 36.3% · interval reaches chance from depth 436.3%
- Kev-0.8Bdepth 3: 48.8% · depth 5: 50.9%50.9%
- Layadepth 3: 38.3% · depth 5: 34.3% · interval reaches chance from depth 334.3%
A model that has only seen shallow examples has no reason to chain five rules. A training set for this failure needs items at every depth the decision can require, balanced so the deep ones are not a rounding error, and an evaluation split by depth so you can see whether tuning moved the deep end or only the easy one.
Failure 2: never saying "unknown"
- Jevrecall on unknown, open world74.2%
- GPT-6 Lunarecall on unknown, open world83.4%
- Kev-9Brecall on unknown, open world68.1%
- Kev-4Brecall on unknown, open world78.8%
- Kev-0.8Brecall on unknown, open world11.9%
- Layarecall on unknown, open world1.0%
Kev-0.8B and Laya almost never get an unknown item right, even though the question defines unknown and offers it as an option. Abstaining is a decision too: if your process needs "not enough information" as an answer, the training data has to contain it, labelled, in realistic proportion.
Failure 3: a lean toward one answer
On the closed-world task, true and false are exactly balanced, yet Kev-4B answered false 66.6% of the time, Kev-0.8B answered false 60.3% of the time and Laya answered false 80.8% of the time. A lean like this shows up as a large gap between per-answer recalls: Kev-4B 42.2% on true against 75.3% on false, Kev-0.8B 45.4% on true against 66.1% on false and Laya 24.8% on true against 86.4% on false. Rebalancing or reweighting the training data, and checking recall per answer after every run, catches it.
Failure 4: reworded rules
The same logic written by people instead of templates (ProofWriter's NatLang set):
- Jevtemplated 86.1%, paraphrased 78.1%−8.0
- GPT-6 Lunatemplated 67.6%, paraphrased 54.8%−12.9
- Kev-9Btemplated 59.6%, paraphrased 55.8%−3.9
- Kev-4Btemplated 55.9%, paraphrased 47.6%−8.2
- Kev-0.8Btemplated 53.8%, paraphrased 53.0%−0.8
- Layatemplated 41.7%, paraphrased 44.4%+2.7
Paraphrased items are shallower on average (depth 2.0 against 2.7), so depth does not explain the gap; if anything it hides some of it. A model tuned only on clean templates should be tested on the wording your users actually write. Paraphrased vs templated rules.
A surprise: negation did not hurt
We preregistered that negation in the rules would lower accuracy. For Jev it did the opposite: 87.5% with negation against 80.3% without, on the open-world task. Theories with negation differ in other ways too, so this is descriptive, but it is a reminder to measure a suspected weakness before spending training data on it. Negation.
Bigger is not automatically better
- Kev-9Bopen world; depth 0 89.3%, depth 5 41.2%58.6%
- Kev-4Bopen world; depth 0 87.7%, depth 5 36.3%53.6%
- Kev-0.8Bopen world; depth 0 70.3%, depth 5 50.9%53.6%
We preregistered that a bigger Kev would be more accurate overall. Verdict so far: failed. OWA: Kev-4B 53.6%, Kev-0.8B 53.6%. CWA: Kev-4B 58.8%, Kev-0.8B 55.8%. Kev-4B is not ahead on OWA, so its interval cannot exclude zero in its favour there. A larger base model changes where a model is strong, not only how strong it is. Kev model sizes.
Fine-tuning Jev, Kev or Laya
Kev-9B, Kev-4B, Kev-0.8B and Laya publish their weights (Kev-9B: 19.31 GB base, Kev-4B: 9.32 GB base, Kev-0.8B: 1.75 GB base and Laya: 0.84 GB), so they can be tuned and re-evaluated on your own hardware. Jev is hosted and its weights are not available; whether it can be adapted to a specific decision is a question for its maker, TypeSafe. Whichever you tune, the same evaluation applies before and after: by depth, by answer, paired on the same items.
What a fine-tuning set for this would need
- Items at every depth your decision can require, balanced across depth, so the deep end is trained and measured.
- Every answer the decision allows, including "not enough information", in realistic proportion and with recall checked per answer.
- The wording your users write, not only clean templates.
- A held-out set split the same way, scored before and after tuning with paired intervals, so an improvement is measured, not assumed.
- A rerun of the held-out set, to check the tuned model answers consistently.