Writing
Notes with the numbers
Write-ups of the kernel and inference work, with the charts and the commands to reproduce them. Most of them started as posts on r/LocalLLaMA. Feed
-
Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and 35 tok/s with 400K tokens in the context
A 126B mixture-of-experts model on a 64 GB laptop, with kernels written for the M1, and how little the speed falls as the context grows.
-
Splash on M1, part 2: 35B-A3B at 144 tok/s and a head-to-head with oMLX and MTPLX
Attention and the MoE expert layers move off MPP too. Speed, GPU temperature, energy per token and memory against two MLX engines on the same Mac.
-
Splash on a 2021 M1 Max: Qwen3.8-27B at 39 tok/s with custom Metal kernels
Splash only runs on M3 and newer. I ported it to my M1 Max and rewrote the kernels that depended on bfloat16 and Apple's MPP library, which doubled decode speed.