Serving & Inference Infrastructure
From a model file on disk to a fast, streamed answer — the KV cache, PagedAttention, batching, and the full request path.
A trained model saved to disk cannot do anything on its own. It is just a big file full of numbers. Turning that file into a service that answers thousands of people at once, in a fraction of a second each, is its own hard engineering problem. This post is about the software that does exactly that job. That software is called an inference engine ("inference" just means running a trained model to get an answer). We will cover the memory trick that makes serving affordable, the scheduling trick that lets one machine handle many users, the way work is split across several chips, and the whole path a request travels from the user to the answer and back.
The math inside the model never changes here. Everything in this post is about running that fixed math on real hardware, for real numbers of users, fast enough and cheap enough. If you want a refresher on how the model generates text, see the Transformer deep dive.
Part 1 · Serving
Why a raw model can't serve real traffic
You can run a model with a few lines of plain PyTorch code. It works fine for one user. It falls apart under ten. Looking at why it fails tells you exactly what a real inference engine is built to fix.
Here is the naive way. It loads the model and asks it to write a reply, one call at a time:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("some-7b-model").cuda()
tok = AutoTokenizer.from_pretrained("some-7b-model")
inputs = tok("Translate to French: good morning", return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=50) # handles ONE request
print(tok.decode(output[0]))
That generate() call can only work on one request at a time, it leaves the chip idle between steps, and it wastes memory. A real engine fixes each of these. Here is the side-by-side.
Plain PyTorch generate() | Inference engine (vLLM/SGLang/TGI) | |
|---|---|---|
| Memory | Grabs memory as it goes, leaves gaps it can't reuse | Reuses memory in small fixed blocks, almost no waste |
| Many users | One request at a time, or a rigid fixed group | Hundreds of requests mixed together and reshuffled every step |
| Chip use | Compute sits idle between word steps | Kept busy by processing many requests together |
| Speed of the math | Generic, unoptimized operations | Hand-tuned code written for this exact job |
Part 2 · Serving
The KV cache and the two phases of generating text
Memory is the central problem in serving, and to see why, you need one idea from how the model works. We will explain it plainly first, then name it.
2.1 The model saves its earlier work so it never repeats it
When a model writes text, it produces one word at a time. To pick the next word, it looks back at every word so far. For each earlier word, it has already computed two helper values (called the key and the value). If it recomputed those for every earlier word at every step, the work would grow with the square of the length — hugely wasteful. So instead the model saves each word's key and value the moment it computes them, and reuses them for every later step. That saved store of keys and values is called the KV cache (KV is just "keys and values").
2.2 Two phases: reading the prompt, then writing the answer
Generating an answer happens in two stages with very different behavior. First the model reads the whole prompt at once. Then it writes the answer one word at a time.
These two phases hit different limits. Reading the prompt (prefill) does a lot of math all at once, so it is limited by how fast the chip can compute. Writing the answer (decode) has to go word by word — word 100 needs word 99 to exist first — so it is limited by how fast the model's numbers and the growing cache can be moved through memory, not by raw compute. Because the phases behave so differently, we measure them with separate numbers: how long until the first word appears (mostly set by prefill) and how long between each following word (mostly set by decode). We give those numbers names in Part 6.
The cache size follows this formula:
The 2 is because we store both a key and a value. Let's plug in numbers for a 70-billion-parameter model: 80 layers, 64 heads, each head width \(d_{\text{head}}=128\), a sequence of 4,000 words, a batch of 16 users at once, and 2 bytes per number.
That is more than the model's own weights. This is the key fact of serving: how many users you can serve is often decided not by the model's size but by how much KV cache fits in the chip's memory (VRAM). The trick in the next section, PagedAttention, exists to squeeze far more of this cache into the same memory.
Part 3 · Serving
PagedAttention and continuous batching
Two ideas, both first shipped in a tool called vLLM, turned model serving from wasteful into efficient. Together they are why a modern engine handles several times more users on the same chip than a naive one. We take them one at a time.
3.1 Handing out memory in small blocks instead of one big slab
The naive approach reserves, up front, a single continuous chunk of memory big enough for each request's longest possible answer. But most answers are much shorter than the maximum, so most of that reserved space sits empty and unused. Wasting space you reserved but never filled is called internal fragmentation, and it can waste most of your cache memory.
The fix borrows an old operating-system idea: cut the cache into small, fixed-size blocks (called pages) and hand them out only as a request actually needs them, wherever there is free room. A small lookup table remembers where each request's pages live. This approach is called PagedAttention.
3.2 Never make a fast request wait for a slow one
A GPU is only efficient when it works on many requests together. Grouping requests together is called batching. The simple version, called static batching, waits until a fixed group of requests is ready, runs the whole group until every request in it is finished, and only then starts a new group. The problem: a request that finishes early is stuck waiting for the slowest one in its group, and the chip sits idle between groups.
The better version works at the level of a single word step. After every step it checks the group: any request that just finished is removed, and a waiting request takes its slot on the very next step. The group is constantly changing rather than fixed. This is called continuous batching (sometimes "in-flight batching").
Part 4 · Serving
When one GPU isn't enough: splitting across chips
A 70-billion-parameter model stored at 2 bytes per number needs about 140 GB of memory — more than fits on any single chip. And even models that do fit may need to serve more users than one chip can handle. The answer is to spread the work over several chips. There are three ways to split it, and they differ in what gets cut up and how much the chips must talk to each other.
| Way of splitting | What gets divided | Plain description | Main cost |
|---|---|---|---|
| Tensor parallelism (split each layer) | Each layer's big number-grids are sliced into pieces; every chip does a slice of every layer. | All chips work on the same layer at the same time, each doing part of it, then combine their partial answers. | The chips must exchange results at every layer, so they need a very fast connection between them (called NVLink). |
| Pipeline parallelism (split by layer) | Whole layers are placed on different chips; the data flows through them in order. | Chip 0 handles layers 1–16, then passes the result to chip 1 for layers 17–32, and so on, like stations on an assembly line. | Much less chatter between chips, but stations can sit idle waiting for the previous one (an "idle gap") unless carefully scheduled. |
| Expert parallelism (split the experts) | In a mixture-of-experts model (Part 3), the separate expert sub-models are placed on different chips. | Each word is sent to whichever chip holds the expert it needs, then the answer comes back. | Extra work to route each word to the right chip, but it spreads out the large memory a mixture-of-experts model needs. |
With tensor parallelism, every chip computes part of a layer and then all of them must combine their partial results before moving to the next layer. Combining partial results from all chips is called an all-reduce, and it can happen dozens of times for a single word. If that exchange has to travel over a slow connection, the waiting to exchange data takes longer than the actual math, and you lose all the speed you hoped to gain. Fast connections between chips (NVLink) exist for exactly this reason. Rule of thumb: use tensor parallelism only among chips inside one machine (fast links), use pipeline parallelism to reach across separate machines (slower links, but less talking needed), and a giant mixture-of-experts model often uses all three at once.
Part 5 · Serving
Running on your own machine vs running on a cluster
Which tool you pick depends on your goal. Are you running one model for yourself on a laptop, caring mainly about speed and privacy? Or are you serving many people from a cluster, caring mainly about total throughput? These are genuinely different problems with different tools.
| On your own machine (one user) | On a cluster (many users) | |
|---|---|---|
| Goal | Fast enough, low memory, works offline and private | Most words per second, many users at once |
| Hardware | A regular CPU, an Apple chip, or one consumer GPU | Many high-end GPUs linked with fast connections |
| Model file format | Shrunk-down (quantized) files, usually in a format called GGUF | Full-precision or lightly shrunk formats (FP16, FP8, AWQ, GPTQ) |
| Tools | llama.cpp, Ollama, LM Studio | vLLM, SGLang, TGI, TensorRT-LLM |
5.1 The tools you run on your own machine
Here is what each local tool is and when to reach for it:
- llama.cpp — a small, very fast engine written in C/C++ that runs models on a plain CPU, an Apple chip, or a GPU. It reads the shrunk-down GGUF file format directly. It sits underneath most other local tools. Pick it when you want the leanest, most direct option.
- Ollama — a friendly wrapper around llama.cpp that turns everything into one command. Typing
ollama run qwen3downloads the model, shrinks it, and starts serving it. Pick it when you want the easiest possible start. - LM Studio — the same idea as Ollama but with a full click-and-point desktop window instead of a command line. Pick it when you prefer a graphical app over the terminal.
On top of any of these you can add a chat window such as Open WebUI, Jan, or AnythingLLM, which gives you a ChatGPT-style interface and can answer questions about your own documents. Here is the whole local stack, top to bottom:
5.2 The tools for serving many users, and the pay-as-you-go middle option
For serving a crowd, here is what each big engine is and when to choose it:
- vLLM — the widely used open engine that introduced PagedAttention and continuous batching. Pick it as a strong general-purpose default for self-hosting.
- SGLang — an engine focused on reusing shared beginnings of prompts (RadixAttention). Pick it when many of your requests share a long common prefix, such as the same big system instruction.
- TGI (Text Generation Inference) — Hugging Face's serving engine, tightly tied to their model library. TGI stands for Text Generation Inference. Pick it if you already live in the Hugging Face ecosystem.
- TensorRT-LLM — NVIDIA's engine, tuned to squeeze the most out of NVIDIA chips. Pick it when you run on NVIDIA hardware and want the very last bit of speed, and don't mind extra setup.
There is also a middle path between "call someone else's private model over the internet" and "run your own cluster." Some companies run popular open models on their own clusters and let you send requests to them, paying by the amount of text. This is called managed open-weight inference ("open-weight" means the model's files are publicly available). Here is who they are:
- Together AI — hosts a broad menu of open models behind a simple pay-per-text interface. Pick it for wide model choice without running servers.
- Fireworks — similar, tuned for speed and production use. Pick it when you want fast managed open models.
- Groq — runs models on its own special chips (called an LPU, a Language Processing Unit) built to spit out words extremely fast. Pick it when very high words-per-second and low delay matter most.
- DeepInfra — hosts open models cheaply per unit of text. Pick it when low price per request is the priority.
Part 6 · Serving
The full path of a request, and the time budget
Now we put it all together. Here is the journey of one real request through a self-hosted setup, from the user's app to the answer streaming back — followed by the timing that decides whether it feels instant.
6.1 The time budget, with real numbers
What the user notices most is how long until the very first word shows up. The time from sending the request to seeing the first word is called Time To First Token, or TTFT (a "token" is roughly a word or word-piece). Here is a realistic breakdown for a well-tuned system:
The important detail: a well-built input safety check is a tiny, fast classifier (about 15 ms, covered in Part 6) that runs at the same time as the model reading the prompt. Because it finishes before the model has even finished reading, it adds no extra waiting. This is why a safety check does not have to slow things down, a point we back up with numbers in Part 6.
Appendix
References & further reading
Grouped by type; freely available online. Serving tools evolve fast — verify feature specifics against current docs.
Research papers & primary sources
- Efficient Memory Management for LLM Serving with PagedAttention. The vLLM paper. https://arxiv.org/abs/2309.06180
- SGLang: Efficient Execution with RadixAttention. https://arxiv.org/abs/2312.07104
- FlashAttention: Fast and Memory-Efficient Exact Attention. https://arxiv.org/abs/2205.14135
- Orca: A Distributed Serving System (continuous batching). https://www.usenix.org/conference/osdi22/presentation/yu
Tools & documentation
- Documentation. https://docs.vllm.ai/
- Documentation. https://docs.sglang.ai/
- Text Generation Inference (TGI). https://huggingface.co/docs/text-generation-inference
- llama.cpp. https://github.com/ggml-org/llama.cpp
- Run large language models locally. https://ollama.com/
Named engines and providers are illustrative of the state of the field as of early 2026. Latency figures are representative order-of-magnitude examples, not benchmarks for any specific system.