The Transformer, Deep Dive
Vaswani's 2017 architecture, block by block — with the math inside and the intuition beside.
In 2017, a single paper — Attention Is All You Need — replaced the recurrent networks that had dominated sequence modeling with an architecture built entirely on attention. Every large language model today is a descendant of the design in that paper.
This is Part 2 of the series. Part 1 built the foundations, ending on the attention mechanism itself. Here we assemble the full Transformer exactly as Vaswani et al. drew it: input embeddings and positional encoding, scaled dot-product attention, multi-head and self-attention, Add & Norm with a position-wise feed-forward network to form an encoder block, and then the decoder — masked self-attention plus cross-attention — that most treatments skip over.
We walk it stage by stage, top to bottom, with the math inside each block and a plain-language "why" beside it. By the end you'll be able to trace a sentence from raw text to output probabilities and explain every box on the way.
★ Deep-Dive Submodule
The Transformer, completed — from a architecture
"
This deep dive walks through the architecture in order, adding the heavy math and intuition, then completes the decoder in the same visual style so you have the full encoder–decoder Transformer end to end.
★ Stage A
Input → Tokenize → One-hot
raw input sentence
sentence = ('Hello', 'world', '!', <pad>, …, <pad>)
$$\text{sentence}=(8667,\ 1362,\ 106,\ 0,\ 0,\ \dots,\ 0)$$
• Tokenizer assigns a number to every word based on its index in a vocabulary set.
every word would be encoded in binary as follows:
$$\mathbf{w}_j=\begin{bmatrix}0 & 0 & \cdots & 1 & \cdots & 0\end{bmatrix}^\top$$
Hence a sentence would be a matrix of size vocab_size × vocab_size.
The math & the why (added depth)
A one-hot vector \(\mathbf{w}_j\in\{0,1\}^{V}\) has a single 1 at the word's vocabulary index. Two problems these notes implicitly set up: it's enormous (length \(V\), tens of thousands) and it's semantically blind — every pair of words is equidistant, so "cat" is no closer to "dog" than to "car". That is exactly why the next stage (embeddings) exists. Modern LLMs additionally use subword tokenization (BPE), splitting rare words into reusable pieces so a fixed ~50K–100K vocabulary covers any string.
★ Stage B
Word embeddings + positional embeddings
Define a matrix \(W_e\) that multiplies each one-hot word \(\mathbf{w}_j\) to generate a smaller vector \(\mathbf{e}_j\):
$$\mathbf{e}_j = W_e\,\mathbf{w}_j$$
This is a purely linear operation (no bias, no activation). So we can fold it into the network and back-propagate errors to update the embedding weights.
kikaben.com/word-embedding-lookup
Added to the word embeddings. To find a word's positional information we use: (1) its position pos in the sentence, (2) the embedding dimension index \(i\), (3) an angle \(\theta\) in radians, (4) sine and cosine:
$$\theta(pos,i)=\frac{pos}{10000^{\,2i/d_{\text{model}}}}$$ $$PE_{(pos,\,2i)}=\sin\theta,\qquad PE_{(pos,\,2i+1)}=\cos\theta$$Without it, an attention-only model might read these as identical: Tom bit a dog / A dog bit Tom.
Ref: kikaben.com/transformers-positional-encoding
With \(W_e\in\mathbb{R}^{d_{\text{model}}\times V}\) and \(\mathbf{w}_j\) one-hot, the product simply selects the \(j\)-th column of \(W_e\) — a fast lookup:
$$W_e\,\mathbf{w}_j = \sum_{v} (W_e)_{:,v}\,(\mathbf{w}_j)_v = (W_e)_{:,j}.$$ It compresses the huge sparse one-hot into a dense vector where similar words land nearby; the famous geometry \(\ \text{king}-\text{man}+\text{woman}\approx\text{queen}\ \) emerges purely from training.★ Stage C
Scaled dot-product attention — your derivation, top to bottom
A dot product between two vectors is \(\ \mathbf{a}\cdot\mathbf{b}=\|\mathbf{a}\|\,\|\mathbf{b}\|\cos\theta\ \), where \(\theta\) is the angle between them — maximum at \(\theta=0^\circ\), minimum at \(\theta=180^\circ\).
For a translation problem the output must contain word vectors. Call the first output word vector \(\mathbf{y}_1\). We project it with a matrix \(W^Q\) to get a query:
$$\mathbf{q}_1 = \mathbf{y}_1 W^Q$$We project each input word \(\mathbf{x}_1\) with \(W^K\) to get a key:
$$\mathbf{k}_1 = \mathbf{x}_1 W^K$$Keep in mind \(\mathbf{q}_1\) and \(\mathbf{k}_1\) have the same dimension. Their dot product gives a score — how strongly the query matches the key:
$$s_{11} = \mathbf{q}_1\cdot \mathbf{k}_1$$Generalize over all input words \(X=[\mathbf{x}_1,\dots,\mathbf{x}_n]\): extract keys \(K\) with the same \(W^K\) and dot \(\mathbf{q}_1\) against all of them:
$$\mathbf{q}_1 K^\top = \begin{bmatrix} s_{11} & s_{12} & \cdots & s_{1n}\end{bmatrix}$$Apply softmax to turn scores into attention weights:
$$\text{softmax}(\mathbf{q}_1 K^\top)=\begin{bmatrix} w_{11} & w_{12} & \cdots & w_{1n}\end{bmatrix}$$Extract values \(V\) with \(W^V\) and take the weighted sum:
$$\text{softmax}(\mathbf{q}_1 K^\top)\,V=\sum_{i} w_{1i}\,\mathbf{v}_i$$Stacking all output queries into \(Q\) gives the matrix form:
$$\text{Attention}(Q,K,V)=\text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V,\qquad Q=YW^Q,\ K=XW^K,\ V=XW^V$$Here \(d_k\) is the dimension of the query/key vectors. When \(d_k\) is large the dot products grow large in magnitude — problematic because softmax's exponential then saturates. Ref: kikaben.com/transformers-self-attention
This is a complete, correct derivation — from a single query–key dot product \(s_{11}\), to a row of scores \(\mathbf{q}_1K^\top\), to softmax weights, to the weighted value sum, to the matrix form. It matches the by-hand example above.
★ Stage D
Multi-head & self-attention
A word can have a different meaning or function depending on context, so we use multiple queries per word rather than one. Vaswani uses \(h=8\) parallel attention calculations — each is a head:
$$\text{head}_i=\text{Attention}(Q_i,K_i,V_i),\quad i=1\dots h$$ $$Q_i=YW_i^Q,\quad K_i=XW_i^K,\quad V_i=XW_i^V$$Concatenate all heads and apply one more matrix \(W^O\):
$$\text{MultiHead}(Y,X)=\text{Concat}(\text{head}_1,\dots,\text{head}_h)\,W^O$$A word's meaning depends on the word itself and its neighbours. For example, "second":
Give me a second, please. I came second in the exam.
But there is only one embedding for "second". So we must treat the word with its context — extract word contexts and relationships within a sentence.
In self-attention, \(Q,K,V\) all originate from the same sentence, so we use \(\text{MultiHead}(X,X)\) to extract each word's context.
★ Stage E
Add & Norm + position-wise FFN = one encoder block
A residual connection adds the sublayer's input back to its output, then normalizes:
$$\text{AddNorm}(x)=\text{LayerNorm}\big(x+\text{Sublayer}(x)\big)$$The residual preserves position/identity information; layer-norm keeps the mean and standard deviation of the vector elements stable, so training stays fast and stable.
The dimension of \(x\) increases from 512 to 2048 by \(W_1\) and reduces from 2048 back to 512 by \(W_2\). Weights are shared across all positions in the layer. ReLU discards negatives, so information is lost — but expanding the dimensionality before ReLU makes it more likely to preserve information. So we add non-linearity without losing much, thanks to the intermediate expansion.
That completes the encoder. The next stage builds the decoder.
★ Stage F — completing the picture
The Decoder, built step by step
" Here is the decoder, drawn top-to-bottom in a consistent flowchart language — colored blocks, formulas inside, a "why?" reasoning bubble alongside. The decoder's job: generate the output sequence one token at a time, attending both to what it has produced so far and to the encoder's understanding of the input.
A decoder block has three sublayers (the encoder had two). Reading top-down:
this is where the input sentence is consulted
Recreated decoder block in a clean visual grammar. Stack ×6, and the encoder feeds every block's cross-attention (the K,V arrows).
① Output embeddings, "shifted right"
During training the decoder is fed the target sequence shifted right by one (prepended with a start token), so at each position it must predict the next token from only the previous ones. Same embedding + positional-encoding machinery as the encoder (Stage B).
② Masked self-attention — the one new idea
tril = torch.tril(torch.ones(T, T))
attention_mat = attention_mat.masked_fill(tril == 0, float('-inf')) # block the future
attention_mat = F.softmax(attention_mat, dim=-1) # future → weight 0
③ Cross-attention — where encoder meets decoder
This sublayer is the bridge these notes pointed at. It's ordinary attention, but with a twist in where Q, K, V come from:
The Query comes from the decoder (what am I trying to generate now?), while Keys and Values come from the encoder's final output (the fully-understood input sentence). So when translating, each output word looks back at the most relevant input words — this is where "chat" aligns to "cat"↔"chat". This single sublayer is the entire reason the architecture is called encoder–decoder.
Linear + Softmax head
The top decoder block's output vector per position is projected by a linear layer to vocabulary size \(V\), then softmaxed into \(P(\text{next token}\mid\text{context})\). Training minimizes cross-entropy against the true next token — the same softmax+CE from Part 1, now over the whole vocabulary.
Decoder block = [masked self-attention] → [cross-attention to encoder] → [FFN], each with Add&Norm (3 sublayers).
The two new ingredients are the causal mask and cross-attention. That's the whole decoder.
★ Stage G
The full architecture, training, and the three families
Putting both halves together — the complete Transformer as Vaswani et al. (2017) drew it, now as a flow, each tower reading top-to-bottom:
The encoder's final output feeds the K,V of every decoder cross-attention (the ⟵ arrow). Both towers are 6 identical blocks in the original paper.
Training: teacher forcing & parallelism
The encoder runs once over the whole source. The decoder is trained with teacher forcing: feed the true (shifted) target and, thanks to the causal mask, predict all positions in parallel in a single pass — a massive speedup over an RNN's step-by-step training. At inference it's autoregressive: generate one token, append, feed back, repeat.
The three architectural families (this is the payoff)
| Family | Which half | Attention | Best at | Examples |
|---|---|---|---|---|
| Encoder-only | Encoder | Bidirectional (unmasked) | Understanding: classification, embeddings, retrieval | BERT, embedding models |
| Decoder-only | Decoder | Causal (masked) | Generation: text, chat, code | GPT, Claude, Gemini, LLaMA |
| Encoder–Decoder | Both | Both + cross-attn | Sequence-to-sequence: translation, summarization | T5, original Transformer, an encoder–decoder translator |
Appendix
References & further reading
The explanations here synthesize the primary research literature with a few exceptional public explainers. Grouped below by type; everything is freely available online.
Research papers & primary sources
- Attention Is All You Need. The original Transformer; the architecture recreated block by block here. https://arxiv.org/abs/1706.03762
- Neural Machine Translation by Jointly Learning to Align and Translate. The attention mechanism that predated the Transformer. https://arxiv.org/abs/1409.0473
- Layer Normalization. The Add & Norm sublayer. https://arxiv.org/abs/1607.06450
- Deep Residual Learning for Image Recognition. Residual connections used around each sublayer. https://arxiv.org/abs/1512.03385
- BERT: Pre-training of Deep Bidirectional Transformers. The encoder-only family. https://arxiv.org/abs/1810.04805
Visualizations & explainers
- The Illustrated Transformer. Colored vector-boxes and step-by-step attention arithmetic. https://jalammar.github.io/illustrated-transformer/
- Attention in transformers, step by step. Attention as a weighted-lookup grid. https://www.3blue1brown.com/lessons/attention
- Transformer Explainer. Live in-browser GPT-2; the Q/K/V search-engine analogy. https://poloclub.github.io/transformer-explainer/
- Transformer Positional Encoding / Self-Attention. Worked walkthroughs of positional encoding and self-attention. https://kikaben.com/transformers-positional-encoding/
Further reading
- The Annotated Transformer. The paper reimplemented line-by-line in PyTorch. https://nlp.seas.harvard.edu/annotated-transformer/
- Let's build GPT: from scratch, in code. Implements a decoder Transformer end to end. https://www.youtube.com/watch?v=kCc8FmEb1nY
Frontier-model details reflect publicly discussed methods; the exact architectures of proprietary systems (Claude, Gemini, GPT-4) are not officially disclosed and are described at the level of well-established public understanding.