Splash for M1/M2 Macs
A port of the Splash inference engine to M1 and M2 Macs. Upstream relies on Apple's MPP matrix library, which is emulated and hangs the GPU on Apple7/8, so this build ships its own Metal kernels.
- Simdgroup-matrix kernels for matmul, attention, MoE experts and the vision encoder
- GGUF models: Unsloth Qwen quantizations and Ternary Bonsai 2
- Speculative decoding: about 70 tok/s on Qwen3.8-27B code on an M1 Max
- Prefill split into ~5 s GPU commands, so long agent sessions do not abort
- Ternary Bonsai 2 27B runs on 16 GB Macs
- Prebuilt, one-line install, side by side with Homebrew's Splash