LLM Infra · Series
Hosting LLMs, From Scratch
How to host and serve large language models — following one request from the API down to the GPU and back. The mental model, the serving engine internals (batching, the KV cache, PagedAttention), quantization, parallelism, autoscaling, and cost. Built for real intuition and for the questions that come up on the job and in interviews.
0 of 2 posts complete0%