System Integration & Automation Engineer

Erik Paperniuk

I build backend systems, business automation and AI tools, and I run large language models locally on Apple Silicon, down to the GPU kernels.

Portrait of Erik Paperniuk

Selected work

Local LLM inference

Release 1.1.0-m1 Metal · C++ · Apple Silicon

Splash for M1/M2 Macs

A port of the Splash inference engine to M1 and M2 Macs. Upstream relies on Apple's MPP matrix library, which is emulated and hangs the GPU on Apple7/8, so this build ships its own Metal kernels.

  • Simdgroup-matrix kernels for matmul, attention, MoE experts and the vision encoder
  • GGUF models: Unsloth Qwen quantizations and Ternary Bonsai 2
  • Speculative decoding: about 70 tok/s on Qwen3.8-27B code on an M1 Max
  • Prefill split into ~5 s GPU commands, so long agent sessions do not abort
  • Ternary Bonsai 2 27B runs on 16 GB Macs
  • Prebuilt, one-line install, side by side with Homebrew's Splash
Installer LM Studio · llama.cpp

LM Studio + Ternary Bonsai 2

Makes LM Studio load Ternary Bonsai 2 GGUF models (PQ2_0 and PTQ1_0) through Prism's llama.cpp fork, with a reasoning-effort selector. Linux and Windows, double-click install.

Patch + guide oMLX · mlx-vlm

oMLX + Prism Hadamard packs

Makes oMLX load Prism Hadamard Qwen3.5 packs (Ternary Bonsai 2), plus a guide to adding an unrecognised architecture to mlx-vlm.

Contact

Get in touch

For collaborations, integrations and engineering conversations.