Latency · model size · memory
Decision model speed, size and memory: what each model needs to run
How long one decision takes, how big each model is, what hardware its maker states, and how much memory the open models used on our laptop.
Speed is not like for like. Hosted models were timed over the network against the vendor's servers; open models on one laptop. A hosted model's time includes the network round trip and whatever the vendor's servers were doing; an open model's time is one Apple M1 Max laptop, which also ran other work. Read each number within its setting. A laptop number says nothing about the same model on a server GPU, and a hosted number says nothing about the model's raw speed.
189 msJev's median time per decision on the open-world task, over the network against TypeSafe's servers; 90% of decisions took under 241 ms.
Time per decision
Wall time around each request as the harness measured it, from the reruns made one request at a time. Only finished timing runs are summarised; a run still in progress shows how far it has got.
| Model | Task | Where | Median | 90th percentile | Slowest | Items | Load average, start / end |
|---|---|---|---|---|---|---|---|
| Jev | Open world | vendor's servers | 189 ms | 241 ms | 1,259 ms | 1,800 of 1,800 | — |
| Jev | Closed world | vendor's servers | 170 ms | 207 ms | 559 ms | 1,800 of 1,800 | — |
| GPT-6 Luna | Open world | vendor's servers | 825 ms | 1,165 ms | 5,836 ms | 1,800 of 1,800 | 15.51 / 7.88 |
| GPT-6 Luna | Closed world | vendor's servers | 929 ms | 1,268 ms | 10,025 ms | 1,800 of 1,800 | 7.88 / 4.64 |
| Kev-4B | Open world | our laptop (Apple M1 Max, 32 GB) | 897 ms | 1,100 ms | 1,431 ms | 1,800 of 1,800 | 7.26 / 11.14 |
| Kev-4B | Closed world | our laptop (Apple M1 Max, 32 GB) | 802 ms | 1,044 ms | 1,481 ms | 1,800 of 1,800 | 11.14 / 10.56 |
| Kev-0.8B | Open world | our laptop (Apple M1 Max, 32 GB) | 123 ms | 152 ms | 379 ms | 1,800 of 1,800 | 5.84 / 5.2 |
| Kev-0.8B | Closed world | our laptop (Apple M1 Max, 32 GB) | 121 ms | 145 ms | 281 ms | 1,800 of 1,800 | 5.2 / 7.16 |
| Laya | Open world | our laptop (Apple M1 Max, 32 GB) | 140 ms | 479 ms | 9,460 ms | 1,800 of 1,800 | 10.03 / 7.12 |
| Laya | Closed world | our laptop (Apple M1 Max, 32 GB) | 121 ms | 398 ms | 6,301 ms | 1,800 of 1,800 | 7.12 / 9.58 |
Not yet timed: Kev-9B. Load average is the machine's one-minute load when the run started and ended, from the run's manifest; it shows how busy the laptop was, which matters for local models.
How to read decision model latency
- Measure where you will deploy. The same open weights run at very different speeds on a laptop, a workstation GPU and a data-centre GPU; Kev's maker states hardware from "Any Apple Silicon Mac" for its smallest size up to an H100 for larger ones (sizes and hardware below).
- One request at a time is a floor, not throughput. These runs sent one request and waited for it. Batched or concurrent serving changes throughput, and for hosted models depends on rate limits.
- Tails matter for interactive use. The 90th percentile and the slowest request are listed beside the median for that reason.
Every model's size, weights and stated hardware
| Model | Parameters | Weights on disk | Maker's stated hardware | Memory used here | Open-world accuracy |
|---|---|---|---|---|---|
| Jevhosted decision model | undisclosed | not available (hosted) | not applicable (hosted) | not applicable (hosted) | 83.8% |
| GPT-6 Lunahosted LLM, used as a classifier with reasoning off | undisclosed | not available (hosted) | not applicable (hosted) | not applicable (hosted) | 64.1% |
| Kev-9Bopen decision model | 9B base (Qwen3.5-9B-Base) + rank-16 LoRA adapter and pointer head | 19.31 GB base (Qwen/Qwen3.5-9B-Base@68c46c4b) + 0.18 GB adapter and head (jaredpalmer/kev-9b@b5d8c18e) | "32 GB Mac, L40S, H100" (Kev README) | 17.0 GB in use, 18.0 GB peak | 58.6% |
| Kev-4Bopen decision model | 4B base (Qwen3.5-4B-Base) + rank-16 LoRA adapter (33.8M trainable) and pointer head | 9.32 GB base (Qwen/Qwen3.5-4B-Base@1001bb4d) + 0.14 GB adapter and head (jaredpalmer/kev-4b@139fdd94) | "32 GB Mac, L40S, H100" (Kev README) | 8.6 GB in use, 17.0 GB peak | 53.6% |
| Kev-0.8Bopen decision model | 0.8B base (Qwen3.5-0.8B-Base) + rank-16 LoRA adapter and pointer head | 1.75 GB base (Qwen/Qwen3.5-0.8B-Base@dc7cdfe2) + 0.05 GB adapter and head (jaredpalmer/kev-0.8b@54f4f877) | "Any Apple Silicon Mac, L4" (Kev README) | 2.1 GB in use, 3.2 GB peak | 53.6% |
| Layaopen decision model | 421M | 0.84 GB (model.safetensors, convaiinnovations/laya repo root) | not stated by the maker | 10.0 GB in use, 10.0 GB peak | 42.4% |
| Kev-27Bopen decision model (not run) | 27B, every weight fine-tuned from Qwen3.8-27B (post-trained) | 51.26 GB full weights (jaredpalmer/kev-27b) | "B200, H200, H100 80 GB; 96-128 GB Mac (expected)" (Kev README, later revision) | not run | not run |
Sizes and weight files are the makers' published figures and the pinned revisions on the Hugging Face Hub. Jev and GPT-6 Luna do not publish their size.
Memory on the test laptop
Every open model ran on one MacBookPro18,4 (32 GB unified memory, Apple M1 Max).
Memory below is the physical footprint of the process holding the model (the Kev server, or the harness for Laya), in use after answering and at its peak, which includes loading. On Apple silicon, model weights live in GPU memory, which a process's resident memory (RSS) does not count, so we use macOS's physical footprint.
Two sizes set the limits here. Kev-9B's base weights alone are 19.31 GB base, most of the laptop's memory. Kev-27B (51.26 GB full weights) does not fit at all and was not run: it would need a rented 80 GB GPU.