API cost · list price · per decision

Decision model cost per decision: Jev, GLiDE and GPT-6 Luna

What each hosted model cost to answer the same 3,600 problems, at its published list price and the tokens its own API reported, per decision and per correct decision.

9.3×GLiDE cost 9.3 times as much as Jev per decision on the same problems ($213 against $23 per million decisions), and 10.4 times as much per correct decision, because it was also less accurate.

Cost per decision, both tasks

Each model's published list price times the input and output tokens its own API reported on each of its 3,600 scored answers. A correct decision is one that matched the gold answer; the cost per correct decision spreads the cost of the wrong ones over the right ones.

ModelList price per million tokens, input / outputTokens billed, input / outputCost of these runsPer million decisionsAccuracyPer million correct decisions
Jev$0.042 / free1,973,258 / 124,200$0.083$2386.6%$27
GPT-6 Lunareasoning off$0.10 / $0.501,118,929 / 39,606$0.132$3764.5%$57
GLiDE$0.30 / free2,559,538 / 697,416$0.768$21377.3%$276

Prices: Jev, TypeSafe's published price; GPT-6 Luna, developers.openai.com/api/docs/models/gpt-6-luna, read 2026-10-01; GLiDE, docs.fastino.ai/pricing, read 2026-10-02. Output tokens that a price lists as free cost nothing even where the API reports them.

By task

The open-world task (true, false or unknown) and the closed-world task (true or false) use the same problems, so their texts are the same length; the difference in cost comes from the answers.

ModelTaskDecisionsCostPer million decisionsAccuracyPer million correct decisions
JevOpen world1,800$0.042$2483.8%$28
JevClosed world1,800$0.040$2289.3%$25
GPT-6 LunaOpen world1,800$0.068$3864.1%$59
GPT-6 LunaClosed world1,800$0.064$3565.0%$55
GLiDEOpen world1,800$0.501$27882.9%$336
GLiDEClosed world1,800$0.267$14871.8%$207

Cost by proof depth

The same cost per million decisions, split by how many inference steps each problem needs. A model that does the same work on every problem costs the same at every depth; one that spends more computation on problems it finds hard costs more where it does so.

  • Jev
  • GPT-6 Luna(reasoning off)
  • GLiDE

Open world: true, false or unknown

$0$100$200$300$400012345proof depth (inference steps)$23$23$23$24$24$24$36$37$37$38$39$39$97$213$288$374$391$305

Closed world: true or false

$0$100$200$300$400012345proof depth (inference steps)$22$22$22$23$23$23$34$35$35$36$37$37$84$157$166$167$159$157
ModelTaskDepth 0Depth 1Depth 2Depth 3Depth 4Depth 5
JevOpen world$23$23$23$24$24$24
JevClosed world$22$22$22$23$23$23
GPT-6 LunaOpen world$36$37$37$38$39$39
GPT-6 LunaClosed world$34$35$35$36$37$37
GLiDEOpen world$97$213$288$374$391$305
GLiDEClosed world$84$157$166$167$159$157

Fastino describes GLiDE as producing "a fast probability distribution," then allocating "additional reasoning when the leading result is uncertain," to "spend more computation on difficult decisions" (Fastino). TypeSafe describes Jev as generating all outputs "in a single query," in parallel (TypeSafe).

Tokens per request

Every model got the same text and question, but each vendor counts tokens its own way, so compare the dollars, not the token counts across models. Within one model, the spread shows whether some requests cost much more than others.

ModelTaskInput tokens, median90th percentileLargestOutput tokens, median
JevOpen world56062271538
JevClosed world53459565731
GPT-6 LunaOpen world32138347611
GPT-6 LunaClosed world29836042211
GLiDEOpen world3311,83811,6381
GLiDEClosed world2931,17410,3251

GLiDE's spread is its thinking: Fastino describes GLiDE as spending more computation when its first answer is uncertain, and those requests report several times the input tokens of the rest, which are billed. Its slow requests are the same ones on the speed page, where 90% of its open-world decisions took under 10.2 seconds.

What these numbers leave out