Runtimes effective-speed test

Qwen3.8-27B on Apple M4 Max

End-to-end speed of five runtime/quant stacks — oMLX, MTPLX and mlx-dspark — measured from 32K to the model's native 256K context, with lossless speculative decoding and prefix caching on each.

Charts

End-to-end speed to 256K

Decode, prefill, effective tps, e2e wait and peak RAM across contexts. Includes the per-config setup, a cache-hit heatmap, the rig spec and a comparison of how each runtime speculates and caches.

Dashboard

Campaign dashboard

Consolidated verdicts and telemetry across the full bake-off — cache reuse, speculation gates and runtime survivors at each context size.

RigApple M4 Max · 40-core GPU · 128 GB unified · macOS 26.5.2
ModelQwen3.8-27B · native context 262,144 tokens
RuntimesoMLX 0.6.3rc2 · MTPLX 2.9.2 · mlx-dspark 0.15.0