Skip to content
All series

Hosting LLMs, From Scratch

How to host and serve large language models — following one request from the API down to the GPU and back. The mental model, the serving engine internals (batching, the KV cache, PagedAttention), quantization, parallelism, autoscaling, and cost. Built for real intuition and for the questions that come up on the job and in interviews.

0 of 2 posts complete0%