Guide · fine-tuning decision models

Fine-tuning decision models: where they fail and what to train on

The failures in this benchmark are specific enough to plan training data around: deep proofs, the unknown answer, reworded rules and a lean toward one answer. Here is what the results show, model by model.

This page describes failures we measured on ProofWriter, a synthetic dataset. It shows what to look for and test, not what any model will do on your decisions.

Failure 1: accuracy collapses past a few inference steps

On the open-world task, chance is a third. Here is each model at depth 3 and depth 5:

A model that has only seen shallow examples has no reason to chain five rules. A training set for this failure needs items at every depth the decision can require, balanced so the deep ones are not a rounding error, and an evaluation split by depth so you can see whether tuning moved the deep end or only the easy one.

Failure 2: never saying "unknown"

Kev-0.8B and Laya almost never get an unknown item right, even though the question defines unknown and offers it as an option. Abstaining is a decision too: if your process needs "not enough information" as an answer, the training data has to contain it, labelled, in realistic proportion.

Failure 3: a lean toward one answer

On the closed-world task, true and false are exactly balanced, yet Kev-4B answered false 66.6% of the time, Kev-0.8B answered false 60.3% of the time and Laya answered false 80.8% of the time. A lean like this shows up as a large gap between per-answer recalls: Kev-4B 42.2% on true against 75.3% on false, Kev-0.8B 45.4% on true against 66.1% on false and Laya 24.8% on true against 86.4% on false. Rebalancing or reweighting the training data, and checking recall per answer after every run, catches it.

Failure 4: reworded rules

The same logic written by people instead of templates (ProofWriter's NatLang set):

Paraphrased items are shallower on average (depth 2.0 against 2.7), so depth does not explain the gap; if anything it hides some of it. A model tuned only on clean templates should be tested on the wording your users actually write. Paraphrased vs templated rules.

A surprise: negation did not hurt

We preregistered that negation in the rules would lower accuracy. For Jev it did the opposite: 87.5% with negation against 80.3% without, on the open-world task. Theories with negation differ in other ways too, so this is descriptive, but it is a reminder to measure a suspected weakness before spending training data on it. Negation.

Bigger is not automatically better

We preregistered that a bigger Kev would be more accurate overall. Verdict so far: failed. OWA: Kev-4B 53.6%, Kev-0.8B 53.6%. CWA: Kev-4B 58.8%, Kev-0.8B 55.8%. Kev-4B is not ahead on OWA, so its interval cannot exclude zero in its favour there. A larger base model changes where a model is strong, not only how strong it is. Kev model sizes.

Fine-tuning Jev, Kev or Laya

Kev-9B, Kev-4B, Kev-0.8B and Laya publish their weights (Kev-9B: 19.31 GB base, Kev-4B: 9.32 GB base, Kev-0.8B: 1.75 GB base and Laya: 0.84 GB), so they can be tuned and re-evaluated on your own hardware. Jev is hosted and its weights are not available; whether it can be adapted to a specific decision is a question for its maker, TypeSafe. Whichever you tune, the same evaluation applies before and after: by depth, by answer, paired on the same items.

What a fine-tuning set for this would need

  1. Items at every depth your decision can require, balanced across depth, so the deep end is trained and measured.
  2. Every answer the decision allows, including "not enough information", in realistic proportion and with recall checked per answer.
  3. The wording your users write, not only clean templates.
  4. A held-out set split the same way, scored before and after tuning with paired intervals, so an improvement is measured, not assumed.
  5. A rerun of the held-out set, to check the tuned model answers consistently.