Latency · model size · memory

Decision model speed, size and memory: what each model needs to run

How long one decision takes, how big each model is, what hardware its maker states, and how much memory the open models used on our laptop.

Speed is not like for like. Hosted models were timed over the network against the vendor's servers; open models on one laptop. A hosted model's time includes the network round trip and whatever the vendor's servers were doing; an open model's time is one Apple M1 Max laptop, which also ran other work. Read each number within its setting. A laptop number says nothing about the same model on a server GPU, and a hosted number says nothing about the model's raw speed.

189 msJev's median time per decision on the open-world task, over the network against TypeSafe's servers; 90% of decisions took under 241 ms.

Time per decision

Wall time around each request as the harness measured it, from the reruns made one request at a time. Only finished timing runs are summarised; a run still in progress shows how far it has got.

ModelTaskWhereMedian90th percentileSlowestItemsLoad average, start / end
JevOpen worldvendor's servers189 ms241 ms1,259 ms1,800 of 1,800—
JevClosed worldvendor's servers170 ms207 ms559 ms1,800 of 1,800—
GPT-6 LunaOpen worldvendor's servers825 ms1,165 ms5,836 ms1,800 of 1,80015.51 / 7.88
GPT-6 LunaClosed worldvendor's servers929 ms1,268 ms10,025 ms1,800 of 1,8007.88 / 4.64
Kev-4BOpen worldour laptop (Apple M1 Max, 32 GB)897 ms1,100 ms1,431 ms1,800 of 1,8007.26 / 11.14
Kev-4BClosed worldour laptop (Apple M1 Max, 32 GB)802 ms1,044 ms1,481 ms1,800 of 1,80011.14 / 10.56
Kev-0.8BOpen worldour laptop (Apple M1 Max, 32 GB)123 ms152 ms379 ms1,800 of 1,8005.84 / 5.2
Kev-0.8BClosed worldour laptop (Apple M1 Max, 32 GB)121 ms145 ms281 ms1,800 of 1,8005.2 / 7.16
LayaOpen worldour laptop (Apple M1 Max, 32 GB)140 ms479 ms9,460 ms1,800 of 1,80010.03 / 7.12
LayaClosed worldour laptop (Apple M1 Max, 32 GB)121 ms398 ms6,301 ms1,800 of 1,8007.12 / 9.58

Not yet timed: Kev-9B. Load average is the machine's one-minute load when the run started and ended, from the run's manifest; it shows how busy the laptop was, which matters for local models.

How to read decision model latency

Every model's size, weights and stated hardware

ModelParametersWeights on diskMaker's stated hardwareMemory used hereOpen-world accuracy
Jevhosted decision modelundisclosednot available (hosted)not applicable (hosted)not applicable (hosted)83.8%
GPT-6 Lunahosted LLM, used as a classifier with reasoning offundisclosednot available (hosted)not applicable (hosted)not applicable (hosted)64.1%
Kev-9Bopen decision model9B base (Qwen3.5-9B-Base) + rank-16 LoRA adapter and pointer head19.31 GB base (Qwen/Qwen3.5-9B-Base@68c46c4b) + 0.18 GB adapter and head (jaredpalmer/kev-9b@b5d8c18e)"32 GB Mac, L40S, H100" (Kev README)17.0 GB in use, 18.0 GB peak58.6%
Kev-4Bopen decision model4B base (Qwen3.5-4B-Base) + rank-16 LoRA adapter (33.8M trainable) and pointer head9.32 GB base (Qwen/Qwen3.5-4B-Base@1001bb4d) + 0.14 GB adapter and head (jaredpalmer/kev-4b@139fdd94)"32 GB Mac, L40S, H100" (Kev README)8.6 GB in use, 17.0 GB peak53.6%
Kev-0.8Bopen decision model0.8B base (Qwen3.5-0.8B-Base) + rank-16 LoRA adapter and pointer head1.75 GB base (Qwen/Qwen3.5-0.8B-Base@dc7cdfe2) + 0.05 GB adapter and head (jaredpalmer/kev-0.8b@54f4f877)"Any Apple Silicon Mac, L4" (Kev README)2.1 GB in use, 3.2 GB peak53.6%
Layaopen decision model421M0.84 GB (model.safetensors, convaiinnovations/laya repo root)not stated by the maker10.0 GB in use, 10.0 GB peak42.4%
Kev-27Bopen decision model (not run)27B, every weight fine-tuned from Qwen3.8-27B (post-trained)51.26 GB full weights (jaredpalmer/kev-27b)"B200, H200, H100 80 GB; 96-128 GB Mac (expected)" (Kev README, later revision)not runnot run

Sizes and weight files are the makers' published figures and the pinned revisions on the Hugging Face Hub. Jev and GPT-6 Luna do not publish their size.

Memory on the test laptop

Every open model ran on one MacBookPro18,4 (32 GB unified memory, Apple M1 Max).

Memory below is the physical footprint of the process holding the model (the Kev server, or the harness for Laya), in use after answering and at its peak, which includes loading. On Apple silicon, model weights live in GPU memory, which a process's resident memory (RSS) does not count, so we use macOS's physical footprint.

Two sizes set the limits here. Kev-9B's base weights alone are 19.31 GB base, most of the laptop's memory. Kev-27B (51.26 GB full weights) does not fit at all and was not run: it would need a rented 80 GB GPU.