Aug 11, 2026 15:40 UTC
hacker news
https://news.ycombinator.com/item?id=49259339
An HN post details 11-16x faster LLM inference using Llama.cpp with Apple Silicon and macOS VMs.
READY TO POST:
llama.cpp users are seeing huge speedups (11-16x!) on apple silicon with macos vms. pretty wild performance gains.
if you're running llms locally on apple silicon, there's new research showing massive speed improvements with llama.cpp and macos vms. worth a look.