Resources › ai › language_models › inference Bookmarks 2026-07-16 bookmark lilianweng.github.io · 10 min Why We Think llm 2025-12-13 bookmark galacodes.hashnode.dev · 15 min Speculative Decoding: From Theory to Implementation efficiency 2025-10-13 bookmark andrewkchan.dev · 37 min Fast LLM Inference From Scratch cuda 2024-12-16 bookmark blog.steelph0enix.dev · 1h 9m llama.cpp guide - Running LLMs locally, on any hardware, from scratch efficiency Subcategories