Deep Learning, In Depth · Part 2 of 3

The Transformer, Deep Dive

Vaswani's 2017 architecture, block by block — with the math inside and the intuition beside.

By PrithvirajPart 2 of 3~35 min read

In 2017, a single paper — Attention Is All You Need — replaced the recurrent networks that had dominated sequence modeling with an architecture built entirely on attention. Every large language model today is a descendant of the design in that paper.

This is Part 2 of the series. Part 1 built the foundations, ending on the attention mechanism itself. Here we assemble the full Transformer exactly as Vaswani et al. drew it: input embeddings and positional encoding, scaled dot-product attention, multi-head and self-attention, Add & Norm with a position-wise feed-forward network to form an encoder block, and then the decoder — masked self-attention plus cross-attention — that most treatments skip over.

We walk it stage by stage, top to bottom, with the math inside each block and a plain-language "why" beside it. By the end you'll be able to trace a sentence from raw text to output probabilities and explain every box on the way.

The whole architecture in one breath
A Transformer is a stack of identical encoder blocks (each = self-attention + feed-forward, both wrapped in residual + normalization) that turn input tokens into context-rich vectors, and a stack of identical decoder blocks (each = masked self-attention + cross-attention + feed-forward) that generate the output one token at a time. GPT keeps only the decoder; BERT keeps only the encoder; the original translator uses both.

★ Deep-Dive Submodule

The Transformer, completed — from a architecture

"

This deep dive walks through the architecture in order, adding the heavy math and intuition, then completes the decoder in the same visual style so you have the full encoder–decoder Transformer end to end.

The whole architecture in one breath
A Transformer is a stack of identical encoder blocks (each = self-attention + FFN, both wrapped in residual + LayerNorm) that turn input tokens into context-rich vectors, and a stack of identical decoder blocks (each = masked self-attention + cross-attention + FFN) that generate the output one token at a time. GPT keeps only the decoder; BERT keeps only the encoder; the original translator uses both.

★ Stage A

Input → Tokenize → One-hot

"Hello world !"
raw input sentence
▼
Tokenize

sentence = ('Hello', 'world', '!', <pad>, …, <pad>)

$$\text{sentence}=(8667,\ 1362,\ 106,\ 0,\ 0,\ \dots,\ 0)$$

Why Tokenize? • maximum length of our sentence is 200. Hence we add trailing padding to the smaller sentences.
• Tokenizer assigns a number to every word based on its index in a vocabulary set.
▼
One hot Encoding

every word would be encoded in binary as follows:

$$\mathbf{w}_j=\begin{bmatrix}0 & 0 & \cdots & 1 & \cdots & 0\end{bmatrix}^\top$$

Hence a sentence would be a matrix of size vocab_size × vocab_size.

Why we need One hot encoding? Naturally every word would have been represented using a one hot vector with the size of the vocabulary set.

The math & the why (added depth)

A one-hot vector \(\mathbf{w}_j\in\{0,1\}^{V}\) has a single 1 at the word's vocabulary index. Two problems these notes implicitly set up: it's enormous (length \(V\), tens of thousands) and it's semantically blind — every pair of words is equidistant, so "cat" is no closer to "dog" than to "car". That is exactly why the next stage (embeddings) exists. Modern LLMs additionally use subword tokenization (BPE), splitting rare words into reusable pieces so a fixed ~50K–100K vocabulary covers any string.

Real-world implication
Tokenization silently governs everything downstream: context length, per-token pricing, and multilingual/code quality all depend on how efficiently the tokenizer encodes text. "Why does the model miscount letters in a word?" — because it sees tokens, not characters.

★ Stage B

Word embeddings + positional embeddings

word Embeddings

Define a matrix \(W_e\) that multiplies each one-hot word \(\mathbf{w}_j\) to generate a smaller vector \(\mathbf{e}_j\):

$$\mathbf{e}_j = W_e\,\mathbf{w}_j$$

This is a purely linear operation (no bias, no activation). So we can fold it into the network and back-propagate errors to update the embedding weights.

Why we need a word embedding in NLP? To reduce the dimension of a single one hot vector, we define a 2D embedding matrix of size num_embedding for each word. Hence there would be vocab_size × num_embedding matrices.
kikaben.com/word-embedding-lookup
▼
Positional Embeddings

Added to the word embeddings. To find a word's positional information we use: (1) its position pos in the sentence, (2) the embedding dimension index \(i\), (3) an angle \(\theta\) in radians, (4) sine and cosine:

$$\theta(pos,i)=\frac{pos}{10000^{\,2i/d_{\text{model}}}}$$ $$PE_{(pos,\,2i)}=\sin\theta,\qquad PE_{(pos,\,2i+1)}=\cos\theta$$
Why we need positional embeddings? Transformers have no recurrence and no convolution, so to use word order we must inject each token's position.
Without it, an attention-only model might read these as identical: Tom bit a dog / A dog bit Tom.
Ref: kikaben.com/transformers-positional-encoding
if word-embedding and positional vectors have similar magnitudes, how does the model distinguish them? The word embedding layer has learnable parameters, so optimization can adjust their magnitudes if necessary. Amazingly, the model learns to use word embeddings and positional encodings without mixing them up.
Math track — the embedding lookup

With \(W_e\in\mathbb{R}^{d_{\text{model}}\times V}\) and \(\mathbf{w}_j\) one-hot, the product simply selects the \(j\)-th column of \(W_e\) — a fast lookup:

$$W_e\,\mathbf{w}_j = \sum_{v} (W_e)_{:,v}\,(\mathbf{w}_j)_v = (W_e)_{:,j}.$$ It compresses the huge sparse one-hot into a dense vector where similar words land nearby; the famous geometry \(\ \text{king}-\text{man}+\text{woman}\approx\text{queen}\ \) emerges purely from training.
Math track — why sinusoids work
For any fixed offset \(k\), \(PE_{pos+k}\) is a linear function of \(PE_{pos}\) (a rotation), so the model can learn to attend by relative distance. These vectors are simply added to the word embeddings: \(\ \mathbf{h}_j=\mathbf{e}_j+PE_{pos(j)}\). Frontier models later replaced this with RoPE (Part 6).

★ Stage C

Scaled dot-product attention — your derivation, top to bottom

Scaled Dot Product attention

A dot product between two vectors is \(\ \mathbf{a}\cdot\mathbf{b}=\|\mathbf{a}\|\,\|\mathbf{b}\|\cos\theta\ \), where \(\theta\) is the angle between them — maximum at \(\theta=0^\circ\), minimum at \(\theta=180^\circ\).

For a translation problem the output must contain word vectors. Call the first output word vector \(\mathbf{y}_1\). We project it with a matrix \(W^Q\) to get a query:

$$\mathbf{q}_1 = \mathbf{y}_1 W^Q$$

We project each input word \(\mathbf{x}_1\) with \(W^K\) to get a key:

$$\mathbf{k}_1 = \mathbf{x}_1 W^K$$

Keep in mind \(\mathbf{q}_1\) and \(\mathbf{k}_1\) have the same dimension. Their dot product gives a score — how strongly the query matches the key:

$$s_{11} = \mathbf{q}_1\cdot \mathbf{k}_1$$

Generalize over all input words \(X=[\mathbf{x}_1,\dots,\mathbf{x}_n]\): extract keys \(K\) with the same \(W^K\) and dot \(\mathbf{q}_1\) against all of them:

$$\mathbf{q}_1 K^\top = \begin{bmatrix} s_{11} & s_{12} & \cdots & s_{1n}\end{bmatrix}$$

Apply softmax to turn scores into attention weights:

$$\text{softmax}(\mathbf{q}_1 K^\top)=\begin{bmatrix} w_{11} & w_{12} & \cdots & w_{1n}\end{bmatrix}$$

Extract values \(V\) with \(W^V\) and take the weighted sum:

$$\text{softmax}(\mathbf{q}_1 K^\top)\,V=\sum_{i} w_{1i}\,\mathbf{v}_i$$

Stacking all output queries into \(Q\) gives the matrix form:

$$\text{Attention}(Q,K,V)=\text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V,\qquad Q=YW^Q,\ K=XW^K,\ V=XW^V$$

Here \(d_k\) is the dimension of the query/key vectors. When \(d_k\) is large the dot products grow large in magnitude — problematic because softmax's exponential then saturates. Ref: kikaben.com/transformers-self-attention

This is a complete, correct derivation — from a single query–key dot product \(s_{11}\), to a row of scores \(\mathbf{q}_1K^\top\), to softmax weights, to the weighted value sum, to the matrix form. It matches the by-hand example above.

★ Stage D

Multi-head & self-attention

Multi-headed attention

A word can have a different meaning or function depending on context, so we use multiple queries per word rather than one. Vaswani uses \(h=8\) parallel attention calculations — each is a head:

$$\text{head}_i=\text{Attention}(Q_i,K_i,V_i),\quad i=1\dots h$$ $$Q_i=YW_i^Q,\quad K_i=XW_i^K,\quad V_i=XW_i^V$$

Concatenate all heads and apply one more matrix \(W^O\):

$$\text{MultiHead}(Y,X)=\text{Concat}(\text{head}_1,\dots,\text{head}_h)\,W^O$$
▼
Self attention

A word's meaning depends on the word itself and its neighbours. For example, "second":

Give me a second, please.    I came second in the exam.

But there is only one embedding for "second". So we must treat the word with its context — extract word contexts and relationships within a sentence.

In self-attention, \(Q,K,V\) all originate from the same sentence, so we use \(\text{MultiHead}(X,X)\) to extract each word's context.

Math track — why heads, sized
With \(h\) heads each of dimension \(d_{\text{model}}/h\) (e.g. \(512/8=64\)), total compute equals one full-width attention, but each head can specialize — one tracks syntax, another coreference, another position. Self-attention means \(Q,K,V\) all come from the same sequence.
Real-world implication
Interpretability researchers read these heads directly — "induction heads" that copy earlier patterns are believed to underlie in-context learning. Multi-head structure is also what Grouped-Query Attention (Part 6) trims to make inference cheaper.

★ Stage E

Add & Norm + position-wise FFN = one encoder block

Add and Norm

A residual connection adds the sublayer's input back to its output, then normalizes:

$$\text{AddNorm}(x)=\text{LayerNorm}\big(x+\text{Sublayer}(x)\big)$$

The residual preserves position/identity information; layer-norm keeps the mean and standard deviation of the vector elements stable, so training stays fast and stable.

Why we need Layer Normalization it prevents the mean and standard deviation of embedding-vector elements from moving around, which makes training unstable and slow. Unlike batch normalization, layer normalization works at each embedding vector (not at the batch level). Hinton's team introduced it because applying batch norm to recurrent networks was impractical.
▼
Position wise Feed Forward Network
$$\text{FFN}(x)=\text{ReLU}(xW_1+b_1)\,W_2+b_2$$

The dimension of \(x\) increases from 512 to 2048 by \(W_1\) and reduces from 2048 back to 512 by \(W_2\). Weights are shared across all positions in the layer. ReLU discards negatives, so information is lost — but expanding the dimensionality before ReLU makes it more likely to preserve information. So we add non-linearity without losing much, thanks to the intermediate expansion.

▼
The transformer uses six stacked encoder blocks. The outputs from the last encoder block become the input features for the decoder.
Math track — the encoder block, assembled
$$\begin{aligned} \mathbf{z}&=\text{LayerNorm}\big(\mathbf{x}+\text{MultiHead}(\mathbf{x},\mathbf{x})\big)\\[4pt] \mathbf{h}&=\text{LayerNorm}\big(\mathbf{z}+\text{FFN}(\mathbf{z})\big) \end{aligned}$$ Two sublayers, each wrapped in residual + LayerNorm. The residual gives gradients an identity path (Part 2.3) so deep stacks train; LayerNorm stabilizes per-token scale. The FFN holds most of the parameters and (per recent interpretability work) much of the model's stored factual knowledge.

That completes the encoder. The next stage builds the decoder.

★ Stage F — completing the picture

The Decoder, built step by step

" Here is the decoder, drawn top-to-bottom in a consistent flowchart language — colored blocks, formulas inside, a "why?" reasoning bubble alongside. The decoder's job: generate the output sequence one token at a time, attending both to what it has produced so far and to the encoder's understanding of the input.

A decoder block has three sublayers (the encoder had two). Reading top-down:

↓ decoder input (previously generated tokens)
① Output embeddings + positional encoding
the tokens generated so far (shifted right)
▼
② Masked Multi-Head Self-Attention
\(Q,K,V\) from decoder; future positions masked to \(-\infty\)
▼
Add & Norm
residual + LayerNorm
▼
③ Cross-Attention (Encoder–Decoder)
\(Q\) from decoder, \(K,V\) from encoder output
this is where the input sentence is consulted
▼
Add & Norm
residual + LayerNorm
▼
Position-wise Feed-Forward
\(\text{ReLU}(xW_1+b_1)W_2+b_2\) — same as encoder
▼
Add & Norm
\(\text{LayerNorm}(x+\text{FFN}(x))\)
▼
Linear + Softmax
project to vocab size \(V\), softmax → \(P(\text{next token})\)
↓ final output probabilities

Recreated decoder block in a clean visual grammar. Stack ×6, and the encoder feeds every block's cross-attention (the K,V arrows).

① Output embeddings, "shifted right"

During training the decoder is fed the target sequence shifted right by one (prepended with a start token), so at each position it must predict the next token from only the previous ones. Same embedding + positional-encoding machinery as the encoder (Stage B).

② Masked self-attention — the one new idea

Math track — causal mask
To predict token \(t\) honestly, the decoder must not peek at tokens \(>t\). We add a mask \(M\) before softmax, with \(M_{ij}=0\) for \(j\le i\) and \(-\infty\) for \(j>i\): $$\text{MaskedAttn}(Q,K,V)=\text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}+M\right)V$$ The \(-\infty\) entries become 0 after softmax, so each position attends only to itself and the past. This is the exact trick used in a from-scratch GPT:
tril = torch.tril(torch.ones(T, T))
attention_mat = attention_mat.masked_fill(tril == 0, float('-inf'))  # block the future
attention_mat = F.softmax(attention_mat, dim=-1)                     # future → weight 0

③ Cross-attention — where encoder meets decoder

This sublayer is the bridge these notes pointed at. It's ordinary attention, but with a twist in where Q, K, V come from:

$$\text{CrossAttn}=\text{softmax}\!\left(\frac{Q_{\text{dec}}\,K_{\text{enc}}^\top}{\sqrt{d_k}}\right)V_{\text{enc}}$$

The Query comes from the decoder (what am I trying to generate now?), while Keys and Values come from the encoder's final output (the fully-understood input sentence). So when translating, each output word looks back at the most relevant input words — this is where "chat" aligns to "cat"↔"chat". This single sublayer is the entire reason the architecture is called encoder–decoder.

Linear + Softmax head

The top decoder block's output vector per position is projected by a linear layer to vocabulary size \(V\), then softmaxed into \(P(\text{next token}\mid\text{context})\). Training minimizes cross-entropy against the true next token — the same softmax+CE from Part 1, now over the whole vocabulary.

★ Stage G

The full architecture, training, and the three families

Putting both halves together — the complete Transformer as Vaswani et al. (2017) drew it, now as a flow, each tower reading top-to-bottom:

ENCODER ×6  ↓ source sentence
Input embed + positional
▼
Multi-Head Self-Attention
(unmasked — sees whole input)
▼
Add & Norm
FFN
↓ encoder output → feeds decoder K,V
DECODER ×6  ↓ output (shifted right)
Output embed + positional
▼
Masked Self-Attention
▼
Cross-Attention
K,V ⟵ encoder output
▼
Add & Norm
FFN
▼
Linear + Softmax
↓ output probabilities

The encoder's final output feeds the K,V of every decoder cross-attention (the ⟵ arrow). Both towers are 6 identical blocks in the original paper.

Training: teacher forcing & parallelism

The encoder runs once over the whole source. The decoder is trained with teacher forcing: feed the true (shifted) target and, thanks to the causal mask, predict all positions in parallel in a single pass — a massive speedup over an RNN's step-by-step training. At inference it's autoregressive: generate one token, append, feed back, repeat.

The three architectural families (this is the payoff)

FamilyWhich halfAttentionBest atExamples
Encoder-onlyEncoderBidirectional (unmasked)Understanding: classification, embeddings, retrievalBERT, embedding models
Decoder-onlyDecoderCausal (masked)Generation: text, chat, codeGPT, Claude, Gemini, LLaMA
Encoder–DecoderBothBoth + cross-attnSequence-to-sequence: translation, summarizationT5, original Transformer, an encoder–decoder translator
Real-world implication — why frontier LLMs are decoder-only
Almost every frontier chat model (GPT-4, Claude, Gemini, LLaMA) is decoder-only. Reasons: a single stack is simpler and scales cleanly; the causal LM objective provides a training signal at every position over unlimited raw text; and cross-attention's separate encoder is unnecessary when you can just concatenate everything (instructions, documents, history) into one context and let masked self-attention handle it. frontier LLMs keep only the decoder tower.
Check: What two sublayers does a decoder block have that an encoder block doesn't (or that differ)? (1) self-attention is masked/causal; (2) an extra cross-attention sublayer taking K,V from the encoder.)

Appendix

References & further reading

The explanations here synthesize the primary research literature with a few exceptional public explainers. Grouped below by type; everything is freely available online.

Research papers & primary sources

  1. Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser & Polosukhin (2017). Attention Is All You Need. The original Transformer; the architecture recreated block by block here. https://arxiv.org/abs/1706.03762
  2. Bahdanau, Cho & Bengio (2014). Neural Machine Translation by Jointly Learning to Align and Translate. The attention mechanism that predated the Transformer. https://arxiv.org/abs/1409.0473
  3. Ba, Kiros & Hinton (2016). Layer Normalization. The Add & Norm sublayer. https://arxiv.org/abs/1607.06450
  4. He, Zhang, Ren & Sun (2015). Deep Residual Learning for Image Recognition. Residual connections used around each sublayer. https://arxiv.org/abs/1512.03385
  5. Devlin, Chang, Lee & Toutanova (2018). BERT: Pre-training of Deep Bidirectional Transformers. The encoder-only family. https://arxiv.org/abs/1810.04805

Visualizations & explainers

  1. Jay Alammar. The Illustrated Transformer. Colored vector-boxes and step-by-step attention arithmetic. https://jalammar.github.io/illustrated-transformer/
  2. 3Blue1Brown (Grant Sanderson). Attention in transformers, step by step. Attention as a weighted-lookup grid. https://www.3blue1brown.com/lessons/attention
  3. Wang et al., Georgia Tech Polo Club. Transformer Explainer. Live in-browser GPT-2; the Q/K/V search-engine analogy. https://poloclub.github.io/transformer-explainer/
  4. KiKaBeN. Transformer Positional Encoding / Self-Attention. Worked walkthroughs of positional encoding and self-attention. https://kikaben.com/transformers-positional-encoding/

Further reading

  1. Harvard NLP (Rush et al.). The Annotated Transformer. The paper reimplemented line-by-line in PyTorch. https://nlp.seas.harvard.edu/annotated-transformer/
  2. Andrej Karpathy. Let's build GPT: from scratch, in code. Implements a decoder Transformer end to end. https://www.youtube.com/watch?v=kCc8FmEb1nY

Frontier-model details reflect publicly discussed methods; the exact architectures of proprietary systems (Claude, Gemini, GPT-4) are not officially disclosed and are described at the level of well-established public understanding.