<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Erik Paperniuk, notes</title>
  <link href="https://paperniuk.github.io/blog/feed.xml" rel="self"/>
  <link href="https://paperniuk.github.io/blog/"/>
  <id>https://paperniuk.github.io/blog/</id>
  <updated>2026-10-04T00:00:00Z</updated>
  <author><name>Erik Paperniuk</name></author>
  <entry>
    <title>Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and 35 tok/s with 400K tokens in the context</title>
    <link href="https://paperniuk.github.io/blog/flash-next-on-m1-max/"/>
    <id>https://paperniuk.github.io/blog/flash-next-on-m1-max/</id>
    <updated>2026-10-04T00:00:00Z</updated>
    <summary>A 126B mixture-of-experts model on a 64 GB laptop, with kernels written for the M1, and how little the speed falls as the context grows.</summary>
  </entry>
  <entry>
    <title>Splash on M1, part 2: 35B-A3B at 144 tok/s and a head-to-head with oMLX and MTPLX</title>
    <link href="https://paperniuk.github.io/blog/splash-on-m1-part-2/"/>
    <id>https://paperniuk.github.io/blog/splash-on-m1-part-2/</id>
    <updated>2026-09-26T00:00:00Z</updated>
    <summary>Attention and the MoE expert layers move off MPP too. Speed, GPU temperature, energy per token and memory against two MLX engines on the same Mac.</summary>
  </entry>
  <entry>
    <title>Splash on a 2021 M1 Max: Qwen3.8-27B at 39 tok/s with custom Metal kernels</title>
    <link href="https://paperniuk.github.io/blog/splash-on-m1/"/>
    <id>https://paperniuk.github.io/blog/splash-on-m1/</id>
    <updated>2026-09-24T00:00:00Z</updated>
    <summary>Splash only runs on M3 and newer. I ported it to my M1 Max and rewrote the kernels that depended on bfloat16 and Apple&#x27;s MPP library, which doubled decode speed.</summary>
  </entry>
</feed>
