From the Transformer to Claude, Gemini, and GPT — the upgrades, the intuition, and the strategy.
By PrithvirajPart 3 of 3~30 min read
You now understand the 2017 Transformer completely — attention, encoders, decoders, the whole flow. So here is the question this final post answers: what actually separates that textbook Transformer from the frontier models — Claude, Gemini, GPT-4-class systems — that people build companies on today?
The reassuring answer, and the thesis of this post, is that the gap is smaller than the hype suggests. Frontier models are not a different paradigm. They are the decoder you already know, made wider, deeper, cheaper to run, better aligned, and multimodal — through a series of individually understandable upgrades. Once you can name each upgrade and say what it replaces and why, the frontier stops looking like magic and starts looking like engineering.
I've written this post in a deliberately dual register. Each section has the technical core — the mechanism, the math, the reason it works — and a strategic lens: what this choice costs, what it buys, and what it means for anyone deciding how to build with or invest in these systems. That second lens is the one that matters if your interest in AI is eventually about leading the work, not only doing it.
The one-sentence map of a frontier model
A modern LLM is a deep stack of masked multi-head attention + feed-forward blocks (each wrapped in residual connections and normalization), trained by gradient descent to predict the next token, then aligned into a helpful assistant — and every frontier upgrade is an optimization of that single skeleton.
Part 1 · The base object
Building a GPT — a decoder-only Transformer at scale
Before the upgrades, we need the object they upgrade. A GPT is the decoder tower from Part 2 of this series, with one simplification: there is no separate encoder, so the cross-attention sublayer is removed. What remains is a stack of blocks, each doing masked self-attention followed by a feed-forward network, trained on exactly one objective — predict the next token.
Why "just predict the next token" is deceptively powerful
It sounds almost trivial. But to predict the next token well across all of human text, a model is quietly forced to learn grammar, facts, arithmetic, translation, code structure, reasoning patterns, and the rhythms of dialogue — because all of those are needed somewhere to reduce prediction error. The objective is simple; the competence it demands is not. This is the central surprise of the last decade: a single, dumb-sounding objective, scaled far enough, produces general capability.
token / embeddingoutput probabilities
Prompt
"The cat sat"
raw text
→
Tokenize
Thecatsat
subword IDs
→
Embed
one vector / token
→
Transformer block ×N
attention + FFN
context mixing
→
Unembed + softmax
→ vocabulary
scores → probabilities
→
Next token
on
down
there
pick & append
How a GPT processes a prompt: text becomes tokens, each token becomes a vector, the block stack mixes context across positions, and a final projection turns the last position into a probability distribution over the whole vocabulary. Vectors are drawn as stacks of colored cells — the same motif recurs below.
1.1 The objective, formally
Math track — causal language modeling
Factor the probability of a sequence by the chain rule of probability:
$$P(x_1,\dots,x_n)=\prod_{t=1}^{n}P(x_t\mid x_{\lt t})$$
The model maximizes \(\sum_t \log P_\theta(x_t\mid x_{\lt t})\) — equivalently, it minimizes cross-entropy per position. The causal mask is what makes \(P_\theta(x_t\mid x_{\lt t})\) depend only on the past, so the model can never "cheat" by looking at the answer. This is the entire training loss of a GPT.
1.1b Attention, the engine inside each block
Every block mixes information across positions with attention — the mechanism built from scratch in Part 1 of this series. The one-line intuition: each word looks at the other words and rebuilds itself as a weighted blend of the ones that matter. It works like a search engine.
Which earlier word does "it" refer to? Attention lets "it" look back and weight every prior word by relevance.
Theanimalcrossedtheroadbecauseitwastired
Query — what I'm looking forlike the text you type into a search box
Key — what each word advertiseslike the titles of pages in the results
Value — what gets pulled backlike the page contents you actually read
Each query is matched against every key (dot product), softmax turns matches into weights, and the output is those weights times the values. Below: the full weight grid for a 4-word sentence, with the causal mask hiding the future.
The
cat
sat
down
The
1.0
cat
.4
.6
sat
.2
.5
.3
down
.1
.3
.2
.4
Rows are query words ("what am I?"), columns are key words ("who's available to attend to"). Darker = more attention. The hatched upper triangle is the causal mask: a token can only look at itself and the past, never the future — which is exactly what makes next-token prediction honest. Each row sums to 1.
1.2 The generation loop
Once trained, generation is a short loop: get the probability distribution for the next token, pick one, append it, and repeat.
logits, loss = self(idx) # scores for next token, shape (B, T, vocab)
logits = logits[:, -1, :] / temperature # take the last position; temperature sharpens/flattens
probs = F.softmax(logits, dim=-1) # -> a distribution over the vocabulary
idx_next = torch.multinomial(probs, 1) # sample one token
idx = torch.cat((idx, idx_next), dim=1) # append it, then repeat
Decoding strategies — the knobs behind "creativity"
Temperature \(T\) rescales the logits (\(z/T\)) before softmax: as \(T\!\to\!0\) generation becomes greedy and deterministic; higher \(T\) makes it more random. Top-k and top-p (nucleus) sampling restrict the choice to the most probable tokens so the model doesn't wander into gibberish. Crucially these are inference-time choices — the weights never change — yet they are what make a model feel focused or wild. A product decision as much as a technical one.
The cat sat▋
on.42
down.19
there.08
↓ sample "on", append, run again
The cat sat on▋
the.51
a.17
my.06
↓ sample "the", append, run again
The cat sat on the▋
mat.39
floor.21
roof.07
Generation is a loop, not a single shot. The model outputs a distribution over the next token; one token is sampled (temperature and top-p decide how adventurously), appended to the context, and the whole thing is fed back in. "The cat sat" → "on" → "the" → "mat", one token at a time. This is what autoregressive means.
One of the most important discoveries about these models is that they get better in a predictable way as you make them bigger. "Better" here means a lower test loss — the model's average surprise at the next token on text it hasn't seen. Lower loss = a sharper predictor. What researchers found is that this loss drops as a smooth power law as you scale up three things:
What each symbol means
Symbol
Reads as
In plain words
\(\mathcal{L}\)
the loss
how wrong the model is on average (lower is better)
\(N\)
parameters
the model's size — the number of weights it has to learn with
\(D\)
data
how many training tokens it reads (its "reading volume")
\(C\)
compute
total arithmetic spent training — roughly \(C \approx 6\,N\,D\)
\(\alpha\)
the exponent
how fast loss falls as you scale (the slope on a log-log plot)
\(N_c\)
a constant
a fitted scale factor; sets where the curve sits
Holding data and compute plentiful, loss versus model size follows:
Read it as: bigger \(N\) ⇒ the fraction shrinks ⇒ loss falls — but only down to an irreducible floor \(\mathcal{L}_\infty\) (the noise inherent in language that no model can predict away).
Test loss vs. model size. Each 10× in size buys a steady drop in loss — a straight line on a log scale — flattening toward a floor.
The magic is the smoothness: the curve has no cliffs or surprises, so a lab can train a few small models, fit the line, and predict the loss of a model 100× bigger before spending a dollar on it. That predictability is what turned scaling into an engineering plan rather than a gamble.
The Chinchilla correction — balance size and data
A power law in size alone tempts you to just build a giant model. But compute is finite, and compute is split between size and data (\(C \approx 6ND\)). Chinchilla asked: for a fixed compute budget, what split of size vs. data gives the lowest loss? The answer — scale them together, at roughly 20 training tokens per parameter. Many famous early models were far too big for how little they were trained on.
✗ Over-sized, under-fed
huge model, too little data
size \(N\)
data \(D\)
Wastes parameters it never taught. Higher loss for the cost.
✓ Chinchilla-balanced
size and data scaled together (≈20 tokens / param)
size \(N\)
data \(D\)
Same compute, better balance. Lower loss.
For a fixed compute budget you're spending the same total either way — the only question is how to divide it between a bigger model and more reading. Chinchilla showed the field had been over-spending on size and starving models of data.
Emergent abilities — the surprising part
Loss falls smoothly, but some capabilities don't appear gradually — they switch on past a certain scale. In-context learning (picking up a task from examples in the prompt), multi-step arithmetic, and chain-of-thought reasoning are largely absent in small models and then emerge in larger ones, even though no one trained for them specifically. Smooth curve, lumpy skills.
Strategic lens — scaling laws turned research into planning
Scaling laws are the single most consequential idea for anyone funding AI. They convert model-building from alchemy into a forecastable engineering discipline: a lab can predict a model's loss before committing tens of millions of dollars to the run. This is why the field pursued ever-larger training runs with such confidence — the return was projectable. The strategic corollary, post-Chinchilla, is that data quality and quantity — not raw parameter count — now gate progress. If you are evaluating an AI strategy and it emphasizes parameter count over data and evaluation quality, that is a red flag.
Part 2 · The architectural upgrades
What changed since 2017 — and the intuition behind each change
Here is the layer stack of a modern decoder-only LLM. Read it top to bottom, the way a token flows through it. Every label in bold is an upgrade over the vanilla Transformer — and the rest of this section explains each one in plain language before the math.
2.1 RoPE — teaching attention about relative position
The intuition first
The original Transformer added a fixed sinusoidal "position vector" to each token's embedding — a decent hack, but it bolts position on from outside and struggles to generalize to sequences longer than those seen in training. Rotary Position Embedding (RoPE) takes a more elegant route: instead of adding position, it rotates each query and key vector by an angle proportional to its position. When two rotated vectors are then compared with a dot product, the result depends only on how far apart the two tokens are — their relative distance — not their absolute indices. Position becomes a property of the comparison, not a tag glued onto the input.
Math track
For a token at position \(m\), RoPE applies a rotation matrix \(R_m\) to its query/key: \(\tilde{\mathbf q}_m = R_m \mathbf q_m\). Because rotations compose, the attention score between positions \(m\) and \(n\) becomes \(\mathbf q_m^\top R_{n-m}\,\mathbf k_n\) — a function of the offset \(n-m\) alone. That relative structure is exactly what lets RoPE-based models be stretched to longer contexts after training (via interpolation of the rotation frequencies).
Strategic lens
RoPE is the quiet enabler behind the "100K / 200K / 1M token context window" race. Long context is not a single trick — it is RoPE (position that extrapolates) plus FlashAttention (memory that fits) plus GQA (a cache that stays small). When a vendor advertises a giant context window, the real questions to ask are about effective use of that context (does accuracy hold in the middle?) and the per-token cost of filling it — not the headline number.
2.2 RMSNorm & pre-norm — cheaper, more stable depth
The intuition
Normalization keeps the numbers flowing through a deep network from drifting to extreme scales. The original used LayerNorm (re-center to mean zero, then re-scale) placed after each sublayer. Frontier models use RMSNorm (skip the re-centering; just divide by the root-mean-square) placed before each sublayer. Dropping the mean subtraction is cheaper and, empirically, loses nothing. Moving the norm before the sublayer ("pre-norm") gives gradients a cleaner path and is what lets stacks of 100+ layers train without blowing up.
Math track
$$\text{RMSNorm}(x)=\frac{x}{\sqrt{\tfrac{1}{d}\sum_i x_i^2 + \epsilon}}\;\odot\;\gamma$$
Compare LayerNorm, which also subtracts the mean \(\mu\) and divides by the standard deviation. RMSNorm keeps only the scale term. Pre-norm means we compute \(x + \text{Sublayer}(\text{Norm}(x))\) rather than \(\text{Norm}(x + \text{Sublayer}(x))\), preserving a clean residual identity path from input to output.
2.3 SwiGLU — a smarter feed-forward network
The intuition
The feed-forward network inside each block is where a large share of a model's parameters — and, per interpretability research, much of its stored factual knowledge — lives. The 2017 version was a plain expand-ReLU-contract sandwich. SwiGLU replaces it with a gated version: one linear branch produces candidate values, a second branch produces a smooth gate that decides how much of each value passes through. The network can now modulate information rather than just hard-clipping negatives to zero. For the same parameter budget, it reliably gives better quality.
Math track
$$\text{SwiGLU}(x) = \big(\text{Swish}(xW_1)\;\otimes\;xV\big)\,W_2,\qquad \text{Swish}(z)=z\cdot\sigma(z)$$
The elementwise product \(\otimes\) is the gate: \(\text{Swish}(xW_1)\) softly opens or closes each channel of \(xV\). It is the same "gating" idea that made LSTMs work, ported into the feed-forward block.
2.4 Grouped-Query Attention — shrinking the memory bill at inference
The intuition
When a model generates text, it caches the Keys and Values of every previous token so it doesn't recompute them each step (the "KV-cache"). With standard multi-head attention, every head keeps its own K and V — and for long contexts that cache becomes the dominant memory cost. Grouped-Query Attention (GQA) keeps many separate query heads (for expressiveness) but lets them share a much smaller number of Key/Value heads. The model gives up very little quality and shrinks the KV-cache dramatically.
Strategic lens
GQA is a pure economics upgrade. It barely moves benchmark quality; what it moves is the cost and speed of serving long-context requests at scale. This is a recurring pattern worth internalizing as a leader: a large fraction of "frontier progress" is not about making the model smarter, but about making a given level of intelligence cheaper to deliver. The unit economics of inference — tokens per dollar, tokens per second — are frequently what decides whether an AI product is viable, and GQA, KV-caching, and quantization are where those battles are fought.
2.5 FlashAttention — same math, hardware-aware
The intuition
Recall from Part 1 that attention builds an \(n\times n\) score matrix for \(n\) tokens — quadratic, and the memory blows up fast. FlashAttention changes none of the math. It reorganizes the computation so the giant score matrix is never fully written to slow memory; instead it is computed in tiles that stay in the GPU's fast on-chip memory, combining partial softmax results as it goes. Identical output, a fraction of the memory traffic, and much faster.
Strategic lens
FlashAttention is the clearest example of a theme that decides who wins in AI: the algorithm and the hardware are not separable. The breakthrough was not a new idea about attention — it was writing attention in a way that respects the memory hierarchy of the specific chips it runs on. Organizations that treat systems/hardware expertise as a first-class discipline (not an afterthought to modeling) capture outsized efficiency gains. This is why frontier labs employ kernel and systems engineers alongside researchers.
2.6 Mixture-of-Experts — decoupling knowledge from compute
The intuition
In a dense model, every token passes through every parameter — so making the model "know more" (more parameters) makes every token more expensive. Mixture-of-Experts (MoE) breaks that link. It replaces the single feed-forward block with many parallel expert blocks and a small router that sends each token to just a few of them. The model can hold an enormous total number of parameters (vast capacity) while each token only activates a small slice (modest per-token compute). You get the knowledge of a huge model at the running cost of a much smaller one.
Math track
A router produces scores over \(E\) experts; the top-\(k\) (often \(k=2\)) are selected and their outputs combined, weighted by the (softmaxed) router scores:
$$\text{MoE}(x)=\sum_{i\in\text{TopK}} g_i(x)\,\text{Expert}_i(x),\qquad g(x)=\text{softmax}(x W_{\text{router}})$$
The practical challenge is load balancing — keeping the router from overusing a few favorite experts — which is handled with auxiliary balancing losses.
Strategic lens
MoE is why "how many parameters does it have?" has become a nearly meaningless question. A trillion-parameter MoE model may activate only tens of billions per token, so its serving cost resembles a much smaller dense model — but its infrastructure complexity (routing, expert placement across GPUs, balancing) is far higher. For a leader, MoE reframes the trade: it buys capability-per-dollar-of-inference at the price of engineering and operational complexity. Gemini and GPT-4-class systems are widely believed to use it precisely because at their scale the inference savings justify that complexity.
2.7 Multimodality — one embedding space for many senses
The intuition
The Transformer never cared that its inputs were words — it operates on vectors. So to make a model "see," you encode an image into patches, turn each patch into a vector in the same embedding space as text tokens, and feed them into the same stack. The model attends across words and image-patches indiscriminately. The same trick extends to audio. Nothing about the core architecture changes; you have simply taught more kinds of input to speak the model's vector language.
Strategic lens
Multimodality is where a lot of near-term product value sits — reading documents, screenshots, charts, and diagrams is what turns a chat toy into a workflow tool. The strategic point is that multimodality is largely a data and encoder investment layered onto an architecture you already have, not a from-scratch rebuild. That lowers the barrier to entry for capable multimodal products and raises the premium on proprietary, high-quality multimodal data.
The key reassurance — read this twice
Every upgrade in this section is an optimization or extension of the mechanism you already understand — not a new paradigm. Masked self-attention, residual connections, normalization, a feed-forward network, softmax over a vocabulary: that skeleton is unchanged from 2017. A frontier model is the same GPT you can build from scratch, made wide, deep, efficient, and multimodal. If you hold onto one idea from this series, hold onto that one — it is what lets you reason about new models as they appear instead of being surprised by each one.
Part 3 · From predictor to assistant
Alignment — the training pipeline that makes a model helpful
A freshly pretrained model is a spectacularly knowledgeable next-token predictor — and a poor assistant. Ask it a question and it might continue with three more questions, because that is a plausible continuation of text. Turning raw capability into a helpful, honest, harmless assistant takes three further stages.
1 · Pretraining
next-token prediction on trillions of tokens → raw capability & world knowledge
▼
2 · Supervised Fine-Tuning (SFT)
train on curated (instruction → good answer) pairs → learns to respond, not just continue
▼
3 · Alignment: RLHF / DPO / Constitutional AI
optimize toward human or AI preferences for helpful, honest, harmless behavior
Math track — RLHF and the modern shortcut (DPO)
RLHF trains a reward model \(r_\phi\) on human preference pairs using the Bradley–Terry model \(P(a\succ b)=\sigma\big(r(a)-r(b)\big)\), then optimizes the LLM (with PPO) to score highly under \(r_\phi\), penalized by a KL-divergence term that keeps it from drifting away from the base model into gibberish. DPO (Direct Preference Optimization) shows you can skip the separate reward model and reinforcement-learning loop entirely, optimizing a closed-form classification loss directly on the preference pairs — simpler, more stable, and now widely used. Constitutional AI (the approach behind Claude) replaces much of the human labeling with a written set of principles the model uses to critique and revise its own answers (RLAIF).
Strategic lens — knowledge vs. behavior is the whole game
This is the most important distinction in the post for anyone thinking about AI strategy: pretraining decides what the model knows; alignment decides how it behaves. Claude, Gemini, and GPT differ far less in raw architecture than in their data mixture, alignment philosophy, and product engineering (tools, safety, latency, refusal behavior). The "personality," trustworthiness, and safety posture of an assistant are shaped after the intelligence is pretrained — which means they are where a company's values and judgment actually get encoded into the product. If you are building an AI organization, alignment and evaluation are not compliance overhead bolted on at the end; they are where your differentiation and your risk both live.
Part 4 · Capabilities on top
What emerges once the machine is built
The architecture and training pipeline produce a base of capability; several further behaviors are layered on through prompting, light training, or system design rather than new architecture.
In-context learning lets a model learn a task purely from examples in the prompt, with no weight change — an emergent property of scale, believed to rest on "induction heads" that copy and extend patterns. Chain-of-thought and reasoning models produce intermediate steps before answering; the newest systems spend more inference-time compute on hard problems, effectively "thinking longer." Tool use and agents let the model emit structured calls to search the web, run code, or hit APIs — the loop that turns a text predictor into something that acts. Retrieval-augmented generation (RAG) pulls relevant documents into the context window so answers are grounded in fresh or private data, fighting hallucination and staleness. And long context — RoPE scaling plus FlashAttention plus GQA — pushes windows to hundreds of thousands or millions of tokens, letting whole codebases or books sit in a single prompt.
Strategic lens — where value moves next
Notice that this entire layer is built around the model, not inside it. Tool use, RAG, agent loops, and evaluation harnesses are software engineering and product problems. As base models commoditize, a growing share of durable value shifts to this layer: proprietary data pipelines, retrieval quality, agent orchestration, and rigorous evaluation. An AI organization that owns only prompts is fragile; one that owns data, retrieval, tooling, and evaluation is defensible.
Trace a single response all the way back
Your text is tokenized (BPE) → embedded → passed through dozens of blocks of RoPE-positioned, masked, grouped-query multi-head attention + SwiGLU (or MoE) feed-forward, each with a residual connection and RMSNorm → a final norm and unembedding → softmax over a 100K+ token vocabulary → a token is sampled with temperature and top-p → appended, and the loop repeats. The weights themselves were set by AdamW gradient descent over trillions of tokens, then shaped by RLHF / Constitutional AI. Every single stage is something you can now explain from first principles.
Synthesis
The whole picture, and what it means to lead here
Here is the complete arc of this series in one line:
If you can explain each arrow in a sentence, you understand modern AI at the level that lets you reason about it independently — evaluate a new model, question a vendor's claims, or set technical direction — rather than react to headlines. That is the difference between someone who uses AI and someone equipped to lead its development.
The three ideas a technical leader should carry out of this
1. Frontier progress is mostly efficiency, not new paradigms. GQA, FlashAttention, MoE, and quantization make a given level of intelligence cheaper to deliver — and inference economics often decide whether a product lives or dies.
2. Behavior is engineered separately from knowledge. Pretraining sets what a model knows; alignment and product design set how it behaves. Your organization's judgment lives in the second half.
3. Durable value is migrating up the stack. As base models commoditize, defensibility comes from data, retrieval, tooling, evaluation, and the agent systems built around the model.
A path to go deeper
If you want to move from understanding to fluency, the highest-leverage practice is to build the pieces yourself and read the primary sources behind each section above: Attention Is All You Need (the architecture); the GPT and scaling-laws papers and Chinchilla (why scale works and how to allocate it); RoPE (position); FlashAttention (systems-aware attention); DPO and Anthropic's Constitutional AI (alignment). Implement a small GPT, then modernize it one upgrade at a time — swap additive positional encoding for RoPE, LayerNorm for RMSNorm, the ReLU feed-forward for SwiGLU — and measure what each change does. Nothing builds intuition like watching the loss curve respond to an idea you understand.
Frontier details (RoPE, GQA, MoE, RLHF/DPO, Constitutional AI) reflect publicly discussed approaches. The exact architectures of proprietary models (Claude, Gemini, GPT-4) are not officially disclosed and are described here at the level of well-established public understanding.
Appendix
References & further reading
The explanations here synthesize the primary research literature with a few exceptional public explainers. Grouped below by type; everything is freely available online.
Brown et al. (2020). Language Models are Few-Shot Learners (GPT-3). Scale and in-context learning. https://arxiv.org/abs/2005.14165
Kaplan et al. (2020). Scaling Laws for Neural Language Models. The power-law loss curves in section 1.3. https://arxiv.org/abs/2001.08361
Hoffmann et al. (2022). Training Compute-Optimal Large Language Models (Chinchilla). The ~20 tokens-per-parameter balance. https://arxiv.org/abs/2203.15556
Shazeer et al. (2017). Outrageously Large Neural Networks (Sparsely-Gated MoE). Mixture-of-Experts. https://arxiv.org/abs/1701.06538
Ouyang et al. (2022). Training language models to follow instructions with human feedback (InstructGPT). RLHF / SFT pipeline. https://arxiv.org/abs/2203.02155
Touvron et al. (2023). LLaMA: Open and Efficient Foundation Language Models. A concrete decoder-only stack using RoPE, RMSNorm, SwiGLU. https://arxiv.org/abs/2302.13971
Brendan Bycroft. LLM Visualization. 3D walkthrough of nanoGPT; inspiration for the vertical token flow. https://bbycroft.net/llm
Financial Times (2023). Generative AI exists because of the transformer. Editorial scrollytelling of tokenization and next-token prediction. https://ig.ft.com/generative-ai/
Frontier-model details reflect publicly discussed methods; the exact architectures of proprietary systems (Claude, Gemini, GPT-4) are not officially disclosed and are described at the level of well-established public understanding.