API cost · list price · per decision
Decision model cost per decision: Jev, GLiDE and GPT-6 Luna
What each hosted model cost to answer the same 3,600 problems, at its published list price and the tokens its own API reported, per decision and per correct decision.
9.3×GLiDE cost 9.3 times as much as Jev per decision on the same problems ($213 against $23 per million decisions), and 10.4 times as much per correct decision, because it was also less accurate.
Cost per decision, both tasks
Each model's published list price times the input and output tokens its own API reported on each of its 3,600 scored answers. A correct decision is one that matched the gold answer; the cost per correct decision spreads the cost of the wrong ones over the right ones.
| Model | List price per million tokens, input / output | Tokens billed, input / output | Cost of these runs | Per million decisions | Accuracy | Per million correct decisions |
|---|---|---|---|---|---|---|
| Jev | $0.042 / free | 1,973,258 / 124,200 | $0.083 | $23 | 86.6% | $27 |
| GPT-6 Lunareasoning off | $0.10 / $0.50 | 1,118,929 / 39,606 | $0.132 | $37 | 64.5% | $57 |
| GLiDE | $0.30 / free | 2,559,538 / 697,416 | $0.768 | $213 | 77.3% | $276 |
Prices: Jev, TypeSafe's published price; GPT-6 Luna, developers.openai.com/api/docs/models/gpt-6-luna, read 2026-10-01; GLiDE, docs.fastino.ai/pricing, read 2026-10-02. Output tokens that a price lists as free cost nothing even where the API reports them.
By task
The open-world task (true, false or unknown) and the closed-world task (true or false) use the same problems, so their texts are the same length; the difference in cost comes from the answers.
| Model | Task | Decisions | Cost | Per million decisions | Accuracy | Per million correct decisions |
|---|---|---|---|---|---|---|
| Jev | Open world | 1,800 | $0.042 | $24 | 83.8% | $28 |
| Jev | Closed world | 1,800 | $0.040 | $22 | 89.3% | $25 |
| GPT-6 Luna | Open world | 1,800 | $0.068 | $38 | 64.1% | $59 |
| GPT-6 Luna | Closed world | 1,800 | $0.064 | $35 | 65.0% | $55 |
| GLiDE | Open world | 1,800 | $0.501 | $278 | 82.9% | $336 |
| GLiDE | Closed world | 1,800 | $0.267 | $148 | 71.8% | $207 |
Cost by proof depth
The same cost per million decisions, split by how many inference steps each problem needs. A model that does the same work on every problem costs the same at every depth; one that spends more computation on problems it finds hard costs more where it does so.
- Jev
- GPT-6 Luna(reasoning off)
- GLiDE
Open world: true, false or unknown
Closed world: true or false
| Model | Task | Depth 0 | Depth 1 | Depth 2 | Depth 3 | Depth 4 | Depth 5 |
|---|---|---|---|---|---|---|---|
| Jev | Open world | $23 | $23 | $23 | $24 | $24 | $24 |
| Jev | Closed world | $22 | $22 | $22 | $23 | $23 | $23 |
| GPT-6 Luna | Open world | $36 | $37 | $37 | $38 | $39 | $39 |
| GPT-6 Luna | Closed world | $34 | $35 | $35 | $36 | $37 | $37 |
| GLiDE | Open world | $97 | $213 | $288 | $374 | $391 | $305 |
| GLiDE | Closed world | $84 | $157 | $166 | $167 | $159 | $157 |
Fastino describes GLiDE as producing "a fast probability distribution," then allocating "additional reasoning when the leading result is uncertain," to "spend more computation on difficult decisions" (Fastino). TypeSafe describes Jev as generating all outputs "in a single query," in parallel (TypeSafe).
Tokens per request
Every model got the same text and question, but each vendor counts tokens its own way, so compare the dollars, not the token counts across models. Within one model, the spread shows whether some requests cost much more than others.
| Model | Task | Input tokens, median | 90th percentile | Largest | Output tokens, median |
|---|---|---|---|---|---|
| Jev | Open world | 560 | 622 | 715 | 38 |
| Jev | Closed world | 534 | 595 | 657 | 31 |
| GPT-6 Luna | Open world | 321 | 383 | 476 | 11 |
| GPT-6 Luna | Closed world | 298 | 360 | 422 | 11 |
| GLiDE | Open world | 331 | 1,838 | 11,638 | 1 |
| GLiDE | Closed world | 293 | 1,174 | 10,325 | 1 |
GLiDE's spread is its thinking: Fastino describes GLiDE as spending more computation when its first answer is uncertain, and those requests report several times the input tokens of the rest, which are billed. Its slow requests are the same ones on the speed page, where 90% of its open-world decisions took under 10.2 seconds.
What these numbers leave out
- List prices, not invoices. Each cost is the vendor's published price times the tokens its API reported; volume discounts, free tiers and minimums are not included. Requests that failed and were retried are not counted.
- Open models have no API cost. Kev-9B, Kev-4B, Kev-0.8B and Laya ran on our own laptop, so there is no per-decision price to compare. Running them means paying for hardware and its time instead; their size, memory and speed show what that takes.
- The benchmark only. These are the scored runs. Reruns for timing and repeatability, and the GPT-6 Luna log-probability probe, are not included.
- Your text, your cost. Cost scales with the length of what you send. These problems average a few hundred tokens; a longer text costs more on every model.