Guide · aligning decision models

Aligning decision models with the decision you mean

A decision model can be accurate and still answer a different question from the one you meant. The open-world and closed-world tasks ask the same problems under two rules for what "unknown" means; the gap between them is an alignment problem you can measure.

Two rules for what cannot be proved

Under the open-world rule, a statement nothing supports is unknown. Under the closed-world rule, it is false. The two tasks state their rule in the question, and each model got the rule for its task.

Open world: true, false or unknown

Using only the facts and rules in the text, is the statement true, false, or unknown? A statement is true if it can be derived from the facts and rules, and false if its negation can be derived. If neither the statement nor its negation can be derived, it is unknown. A rule applies only when all of its conditions are established; 'not' in a condition requires the negation to be stated or derived.

Closed world: true or false

Using only the facts and rules in the text, is the statement true or false? Assume that anything which cannot be derived from the facts and rules is false: a positive statement is true only if it can be derived, and a statement with 'not' is true when the unnegated statement cannot be derived. A rule applies when all of its conditions hold; 'not' in a condition holds when the unnegated condition cannot be derived.

Accuracy under each rule

Closed-world minus open-world accuracy, in points. Chance is 50% on the closed-world task and 33% on the open-world one, so higher closed-world numbers are expected; the gap above chance is what to compare.

Open world: true, false or unknown

Open-world results
ModelAccuracy95% intervalMacro-F1Recall: trueRecall: falseRecall: unknownItems
Jev83.8%82.2–85.50.83888.8%88.3%74.2%1,800
GPT-6 Luna64.1%61.7–66.20.64758.3%50.9%83.4%1,800
Kev-9B58.6%56.3–60.90.58859.2%48.6%68.1%1,800
Kev-4B53.6%51.4–56.00.53246.6%35.9%78.8%1,800
Kev-0.8B53.6%51.4–55.90.47758.2%89.6%11.9%1,800
Laya42.4%40.3–44.50.33740.2%85.1%1.0%1,800

Chance is 33.3%; always giving the most common answer scores 33.6% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.

Closed world: true or false

Closed-world results
ModelAccuracy95% intervalMacro-F1Recall: trueRecall: falseItems
Jev89.3%87.8–90.70.89389.7%88.9%1,800
GPT-6 Luna65.0%62.8–67.40.65066.2%63.8%1,800
Kev-9B63.8%61.5–66.10.63864.7%62.9%1,800
Kev-4B58.8%56.6–61.10.57642.2%75.3%1,800
Kev-0.8B55.8%53.5–57.90.55345.4%66.1%1,800
Laya55.6%53.3–58.10.50924.8%86.4%1,800

Chance is 50.0%; always giving the most common answer scores 50.0% on these items. Recall is the share of items with that correct answer that the model got right. GPT-6 Luna ran with reasoning off, under the same one-request protocol as every model.

Does the model follow the rule it was given?

On the open-world task, 590 items have unknown as the correct answer. What each model said on them:

ModelSaid unknownSaid trueSaid false
Jev74.2%21.0%4.7%
GPT-6 Luna83.4%7.6%9.0%
Kev-9B68.1%22.2%9.7%
Kev-4B78.8%15.1%6.1%
Kev-0.8B11.9%25.4%62.7%
Laya1.0%21.5%77.5%

Kev-0.8B and Laya answer these items as if the closed-world rule applied, even though the question defines unknown and offers it: Kev-0.8B said false on 62.7% and Laya said false on 77.5%. That is not a reasoning failure so much as a different decision from the one asked for, and it is the kind of misalignment that a single accuracy number hides.

Aligning Jev

Jev scored 83.8% under the open-world rule and 89.3% under the closed-world one, the highest of the models here under both. On unknown items it said true 21.0% of the time and false 4.7%: when it errs on "not enough information", it leans toward yes. If your decision treats "cannot tell" as its own outcome, that is the error to test and correct for first. Jev's full results.

How to check alignment on your own decisions

  1. Write the decision rule down, including what "unknown" or "not enough information" means, and put it in the instructions the model gets.
  2. Build test items where the rule matters: the same facts that should lead to different answers under different rules.
  3. Look at what the model says on each correct answer, not only its accuracy; a lean toward one answer is a misalignment you can see in one table.
  4. Change the rule and check the answers move the way they should. A model that answers the same under both rules is not reading the rule.
  5. Repeat the test after any change of model version, instructions or tuning.