LLMOps, In Depth · Part 4 of 6

Serving & Inference Infrastructure

From a model file on disk to a fast, streamed answer — the KV cache, PagedAttention, batching, and the full request path.

By PrithvirajPart 4 of 6~32 min read

A trained model saved to disk cannot do anything on its own. It is just a big file full of numbers. Turning that file into a service that answers thousands of people at once, in a fraction of a second each, is its own hard engineering problem. This post is about the software that does exactly that job. That software is called an inference engine ("inference" just means running a trained model to get an answer). We will cover the memory trick that makes serving affordable, the scheduling trick that lets one machine handle many users, the way work is split across several chips, and the whole path a request travels from the user to the answer and back.

The math inside the model never changes here. Everything in this post is about running that fixed math on real hardware, for real numbers of users, fast enough and cheap enough. If you want a refresher on how the model generates text, see the Transformer deep dive.

The one goal that drives everything
Serving has a single aim: get as many words per second out of each expensive chip as possible, while making sure no single user waits too long. The number of words a chip produces per second is called throughput. Every technique in this post — the KV cache, PagedAttention, continuous batching, splitting across chips — exists to raise throughput without making any one user wait.

Part 1 · Serving

Why a raw model can't serve real traffic

You can run a model with a few lines of plain PyTorch code. It works fine for one user. It falls apart under ten. Looking at why it fails tells you exactly what a real inference engine is built to fix.

Here is the naive way. It loads the model and asks it to write a reply, one call at a time:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("some-7b-model").cuda()
tok   = AutoTokenizer.from_pretrained("some-7b-model")

inputs = tok("Translate to French: good morning", return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=50)   # handles ONE request
print(tok.decode(output[0]))

That generate() call can only work on one request at a time, it leaves the chip idle between steps, and it wastes memory. A real engine fixes each of these. Here is the side-by-side.

Plain PyTorch generate()Inference engine (vLLM/SGLang/TGI)
MemoryGrabs memory as it goes, leaves gaps it can't reuseReuses memory in small fixed blocks, almost no waste
Many usersOne request at a time, or a rigid fixed groupHundreds of requests mixed together and reshuffled every step
Chip useCompute sits idle between word stepsKept busy by processing many requests together
Speed of the mathGeneric, unoptimized operationsHand-tuned code written for this exact job

Part 2 · Serving

The KV cache and the two phases of generating text

Memory is the central problem in serving, and to see why, you need one idea from how the model works. We will explain it plainly first, then name it.

2.1 The model saves its earlier work so it never repeats it

When a model writes text, it produces one word at a time. To pick the next word, it looks back at every word so far. For each earlier word, it has already computed two helper values (called the key and the value). If it recomputed those for every earlier word at every step, the work would grow with the square of the length — hugely wasteful. So instead the model saves each word's key and value the moment it computes them, and reuses them for every later step. That saved store of keys and values is called the KV cache (KV is just "keys and values").

2.2 Two phases: reading the prompt, then writing the answer

Generating an answer happens in two stages with very different behavior. First the model reads the whole prompt at once. Then it writes the answer one word at a time.

Phase 1 · reading the prompt
The whole prompt is processed in one shot, filling the KV cache with keys and values for every prompt word. Heavy math, done all at once. This phase is called prefill.
↓ then, one word at a time
Phase 2 · writing the answer
Produce one word, add its key and value to the cache, repeat. Must be done in order, and is limited by how fast data moves through memory. This phase is called decode.

These two phases hit different limits. Reading the prompt (prefill) does a lot of math all at once, so it is limited by how fast the chip can compute. Writing the answer (decode) has to go word by word — word 100 needs word 99 to exist first — so it is limited by how fast the model's numbers and the growing cache can be moved through memory, not by raw compute. Because the phases behave so differently, we measure them with separate numbers: how long until the first word appears (mostly set by prefill) and how long between each following word (mostly set by decode). We give those numbers names in Part 6.

How big the KV cache gets — a worked number

The cache size follows this formula:

$$\text{cache bytes} = 2 \times \text{layers} \times \text{heads} \times d_{\text{head}} \times \text{seq len} \times \text{batch} \times \text{bytes per number}$$

The 2 is because we store both a key and a value. Let's plug in numbers for a 70-billion-parameter model: 80 layers, 64 heads, each head width \(d_{\text{head}}=128\), a sequence of 4,000 words, a batch of 16 users at once, and 2 bytes per number.

$$2 \times 80 \times 64 \times 128 \times 4000 \times 16 \times 2 \approx 1.7 \times 10^{11}\ \text{bytes} \approx 168\ \text{GB}$$

That is more than the model's own weights. This is the key fact of serving: how many users you can serve is often decided not by the model's size but by how much KV cache fits in the chip's memory (VRAM). The trick in the next section, PagedAttention, exists to squeeze far more of this cache into the same memory.

Part 3 · Serving

PagedAttention and continuous batching

Two ideas, both first shipped in a tool called vLLM, turned model serving from wasteful into efficient. Together they are why a modern engine handles several times more users on the same chip than a naive one. We take them one at a time.

3.1 Handing out memory in small blocks instead of one big slab

The naive approach reserves, up front, a single continuous chunk of memory big enough for each request's longest possible answer. But most answers are much shorter than the maximum, so most of that reserved space sits empty and unused. Wasting space you reserved but never filled is called internal fragmentation, and it can waste most of your cache memory.

The fix borrows an old operating-system idea: cut the cache into small, fixed-size blocks (called pages) and hand them out only as a request actually needs them, wherever there is free room. A small lookup table remembers where each request's pages live. This approach is called PagedAttention.

Naive: one big reserved slab per request
Reserve room for 2000 words  →  [used: 90 words ......... empty: 1910 words wasted]
↓ PagedAttention instead
Small blocks, given out on demand
Request A: [blk 3][blk 7][blk 1]   Request B: [blk 2][blk 5]   free: [blk 4][blk 6]...   almost nothing wasted
A bonus: sharing identical pages
Because pages are handed out individually, two requests that begin with the exact same text (say, the same system instruction at the top) can point at the very same cached pages instead of each storing their own copy. That saves both memory and repeated work. A related tool called SGLang extends this sharing to any shared beginning of text using a tree structure, and calls its version RadixAttention ("radix tree" is the data structure it uses to find shared prefixes quickly).

3.2 Never make a fast request wait for a slow one

A GPU is only efficient when it works on many requests together. Grouping requests together is called batching. The simple version, called static batching, waits until a fixed group of requests is ready, runs the whole group until every request in it is finished, and only then starts a new group. The problem: a request that finishes early is stuck waiting for the slowest one in its group, and the chip sits idle between groups.

The better version works at the level of a single word step. After every step it checks the group: any request that just finished is removed, and a waiting request takes its slot on the very next step. The group is constantly changing rather than fixed. This is called continuous batching (sometimes "in-flight batching").

Static batching
[ A B C D ] run to the end together → A finished at step 3 but idles until D finishes at step 40 → then start next group
↓ continuous batching
Continuous batching
step 3: A done → drop A, add E immediately   step 7: C done → drop C, add F   the group never stops churning
Why this matters to your bill
PagedAttention and continuous batching together are why a tool like vLLM can push several times more traffic through the same hardware than a naive setup. In practice that is the difference between needing eight GPUs and needing two for the same number of users — a large, direct saving on your infrastructure cost. When you pick a serving tool, both features are must-haves, not extras.

Part 4 · Serving

When one GPU isn't enough: splitting across chips

A 70-billion-parameter model stored at 2 bytes per number needs about 140 GB of memory — more than fits on any single chip. And even models that do fit may need to serve more users than one chip can handle. The answer is to spread the work over several chips. There are three ways to split it, and they differ in what gets cut up and how much the chips must talk to each other.

Way of splittingWhat gets dividedPlain descriptionMain cost
Tensor parallelism
(split each layer)
Each layer's big number-grids are sliced into pieces; every chip does a slice of every layer.All chips work on the same layer at the same time, each doing part of it, then combine their partial answers.The chips must exchange results at every layer, so they need a very fast connection between them (called NVLink).
Pipeline parallelism
(split by layer)
Whole layers are placed on different chips; the data flows through them in order.Chip 0 handles layers 1–16, then passes the result to chip 1 for layers 17–32, and so on, like stations on an assembly line.Much less chatter between chips, but stations can sit idle waiting for the previous one (an "idle gap") unless carefully scheduled.
Expert parallelism
(split the experts)
In a mixture-of-experts model (Part 3), the separate expert sub-models are placed on different chips.Each word is sent to whichever chip holds the expert it needs, then the answer comes back.Extra work to route each word to the right chip, but it spreads out the large memory a mixture-of-experts model needs.
Why tensor parallelism needs a very fast link

With tensor parallelism, every chip computes part of a layer and then all of them must combine their partial results before moving to the next layer. Combining partial results from all chips is called an all-reduce, and it can happen dozens of times for a single word. If that exchange has to travel over a slow connection, the waiting to exchange data takes longer than the actual math, and you lose all the speed you hoped to gain. Fast connections between chips (NVLink) exist for exactly this reason. Rule of thumb: use tensor parallelism only among chips inside one machine (fast links), use pipeline parallelism to reach across separate machines (slower links, but less talking needed), and a giant mixture-of-experts model often uses all three at once.

Check: Why does a mixture-of-experts model use a lot of memory even though each word only uses a small part of it (Part 3), and why does that make expert parallelism a good fit? (All the experts have to be loaded in memory at all times, even though each word only passes through a few of them. Spreading the experts across chips shares out that large memory load, so no single chip has to hold all of them.)

Part 5 · Serving

Running on your own machine vs running on a cluster

Which tool you pick depends on your goal. Are you running one model for yourself on a laptop, caring mainly about speed and privacy? Or are you serving many people from a cluster, caring mainly about total throughput? These are genuinely different problems with different tools.

On your own machine (one user)On a cluster (many users)
GoalFast enough, low memory, works offline and privateMost words per second, many users at once
HardwareA regular CPU, an Apple chip, or one consumer GPUMany high-end GPUs linked with fast connections
Model file formatShrunk-down (quantized) files, usually in a format called GGUFFull-precision or lightly shrunk formats (FP16, FP8, AWQ, GPTQ)
Toolsllama.cpp, Ollama, LM StudiovLLM, SGLang, TGI, TensorRT-LLM

5.1 The tools you run on your own machine

Here is what each local tool is and when to reach for it:

  • llama.cpp — a small, very fast engine written in C/C++ that runs models on a plain CPU, an Apple chip, or a GPU. It reads the shrunk-down GGUF file format directly. It sits underneath most other local tools. Pick it when you want the leanest, most direct option.
  • Ollama — a friendly wrapper around llama.cpp that turns everything into one command. Typing ollama run qwen3 downloads the model, shrinks it, and starts serving it. Pick it when you want the easiest possible start.
  • LM Studio — the same idea as Ollama but with a full click-and-point desktop window instead of a command line. Pick it when you prefer a graphical app over the terminal.

On top of any of these you can add a chat window such as Open WebUI, Jan, or AnythingLLM, which gives you a ChatGPT-style interface and can answer questions about your own documents. Here is the whole local stack, top to bottom:

Chat window
Open WebUI / Jan / AnythingLLM — the interface you type into, plus your-documents search
↓ sends your message to
Inference engine
Ollama or LM Studio (both run llama.cpp underneath)
↓ which uses
Your hardware's fast-math system
Metal (Apple), CUDA (NVIDIA), or ROCm (AMD)

5.2 The tools for serving many users, and the pay-as-you-go middle option

For serving a crowd, here is what each big engine is and when to choose it:

  • vLLM — the widely used open engine that introduced PagedAttention and continuous batching. Pick it as a strong general-purpose default for self-hosting.
  • SGLang — an engine focused on reusing shared beginnings of prompts (RadixAttention). Pick it when many of your requests share a long common prefix, such as the same big system instruction.
  • TGI (Text Generation Inference) — Hugging Face's serving engine, tightly tied to their model library. TGI stands for Text Generation Inference. Pick it if you already live in the Hugging Face ecosystem.
  • TensorRT-LLM — NVIDIA's engine, tuned to squeeze the most out of NVIDIA chips. Pick it when you run on NVIDIA hardware and want the very last bit of speed, and don't mind extra setup.

There is also a middle path between "call someone else's private model over the internet" and "run your own cluster." Some companies run popular open models on their own clusters and let you send requests to them, paying by the amount of text. This is called managed open-weight inference ("open-weight" means the model's files are publicly available). Here is who they are:

  • Together AI — hosts a broad menu of open models behind a simple pay-per-text interface. Pick it for wide model choice without running servers.
  • Fireworks — similar, tuned for speed and production use. Pick it when you want fast managed open models.
  • Groq — runs models on its own special chips (called an LPU, a Language Processing Unit) built to spit out words extremely fast. Pick it when very high words-per-second and low delay matter most.
  • DeepInfra — hosts open models cheaply per unit of text. Pick it when low price per request is the priority.
The three-way choice, plainly
Someone else's private model over the internet (Claude, GPT, Gemini): highest capability, no servers to run, you pay per unit of text, and your data goes to them. Managed open model (Together, Fireworks, Groq, DeepInfra): you choose the open model, little to run, pay per unit of text, and more control over your data. Your own cluster (self-hosted vLLM): full control and privacy, and the cheapest per unit of text once volume is high, but you have to run and maintain everything. The tipping point: running your own cluster only becomes cheaper once you have a lot of steady traffic, because then the fixed cost of the GPUs beats a bill that grows with every request. Below that level, paying per unit of text is both cheaper and far less work.

Part 6 · Serving

The full path of a request, and the time budget

Now we put it all together. Here is the journey of one real request through a self-hosted setup, from the user's app to the answer streaming back — followed by the timing that decides whether it feels instant.

1 · The user's app sends the request
A web request that keeps the connection open so words can stream back as they are produced
↓
2 · Front door
Nginx or Cloudflare — handles the secure connection, checks the user is allowed in, and limits abuse
↓
3 · Traffic director and waiting line
Sends the request to the least-busy machine; a quick safety check on the input runs here (Part 6)
↓
4 · The inference engine does the work
Scheduler groups requests → PagedAttention manages the KV cache → the work runs split across GPUs
↓
5 · Choosing each word
From the model's list of likely next words, pick one using the temperature and top-p settings
↓
6 · Words stream back
Each word is sent to the user as it is produced; a safety check on the output runs alongside

6.1 The time budget, with real numbers

What the user notices most is how long until the very first word shows up. The time from sending the request to seeing the first word is called Time To First Token, or TTFT (a "token" is roughly a word or word-piece). Here is a realistic breakdown for a well-tuned system:

Total time to first word = network trip + max(input safety check, reading the prompt) + producing the first word 0 ms ~20 ms ~300 ms ~350 ms |─────────────|─────────────────────────|────────────────────────|──▶ first word appears network trip input safety check model reads the prompt streams to and back (15 ms, runs AT THE SAME (the prefill phase) to the app TIME as reading the prompt)

The important detail: a well-built input safety check is a tiny, fast classifier (about 15 ms, covered in Part 6) that runs at the same time as the model reading the prompt. Because it finishes before the model has even finished reading, it adds no extra waiting. This is why a safety check does not have to slow things down, a point we back up with numbers in Part 6.

The four numbers that define a serving promise
A serving promise (an "SLA," service-level agreement, the target you commit to) rests on four numbers. TTFT — time to the first word, i.e. how responsive it feels. ITL — inter-token latency, the gap between each following word, which decides how smoothly the answer streams (also called TPOT, time per output token). Throughput — total words per second across all users at once, i.e. your capacity. Cost per word — how much each unit of text costs you. These pull against each other: bigger batches raise throughput and lower cost per word, but can raise TTFT because each request waits a little longer to be scheduled. Tuning a deployment means choosing where on that trade-off your product needs to sit.

Appendix

References & further reading

Grouped by type; freely available online. Serving tools evolve fast — verify feature specifics against current docs.

Research papers & primary sources

  1. Kwon et al. (2023). Efficient Memory Management for LLM Serving with PagedAttention. The vLLM paper. https://arxiv.org/abs/2309.06180
  2. Zheng et al. (2023). SGLang: Efficient Execution with RadixAttention. https://arxiv.org/abs/2312.07104
  3. Dao et al. (2022). FlashAttention: Fast and Memory-Efficient Exact Attention. https://arxiv.org/abs/2205.14135
  4. Yu et al. (2022). Orca: A Distributed Serving System (continuous batching). https://www.usenix.org/conference/osdi22/presentation/yu

Tools & documentation

  1. vLLM. Documentation. https://docs.vllm.ai/
  2. SGLang. Documentation. https://docs.sglang.ai/
  3. Hugging Face. Text Generation Inference (TGI). https://huggingface.co/docs/text-generation-inference
  4. ggml-org. llama.cpp. https://github.com/ggml-org/llama.cpp
  5. Ollama. Run large language models locally. https://ollama.com/

Named engines and providers are illustrative of the state of the field as of early 2026. Latency figures are representative order-of-magnitude examples, not benchmarks for any specific system.