Foundations of Deep Learning
The ideas behind every LLM — from a single neuron to attention.
Modern AI can feel like an impenetrable stack of jargon — attention, transformers, RLHF, mixture-of-experts. But almost all of it is built from one humble operation repeated at scale: multiply by a matrix, bend the result with a nonlinearity, and adjust the numbers so the output gets a little less wrong.
This is the first post in a three-part series that rebuilds deep learning from the ground up — not as a list of buzzwords, but as a small number of ideas you can hold in your head and reason from. In this post we cover the foundations: the neuron, why depth works, how networks learn, and the two classic architectures (CNNs and RNN/LSTMs) that set the stage for everything that came after. Part 2 dissects the Transformer in full. Part 3 shows how today's frontier models are built on exactly these bones.
Everything here is derived, not asserted. Where it helps, there are worked numerical examples you can check with a calculator, and short code snippets that map one-to-one onto the math.
Part 1 · Foundations
The neuron, linear algebra, and why depth works
Everything in deep learning is built from one operation repeated at scale: an affine transform followed by a nonlinearity. Master this one line and its calculus, and every later architecture is a rearrangement of it.
1.1 A single neuron
A neuron takes an input vector \(\mathbf{x}\in\mathbb{R}^d\), weights \(\mathbf{w}\in\mathbb{R}^d\), a bias \(b\), and produces a scalar activation:
The dot product \(\mathbf{w}^\top\mathbf{x}=\|\mathbf{w}\|\,\|\mathbf{x}\|\cos\theta\) measures alignment between input and weights. A neuron is a learned template matcher: it fires strongly when the input points in the direction of \(\mathbf{w}\). Keep this geometric reading — it returns in attention as the query–key dot product.
1.2 A layer is a matrix multiply
Stack \(m\) neurons and their weight vectors become the rows of a matrix \(W\in\mathbb{R}^{m\times d}\). A whole layer is then one matrix–vector product plus a broadcast bias:
Process a batch of \(N\) inputs at once by stacking them into \(X\in\mathbb{R}^{N\times d}\): \(Z=XW^\top+\mathbf{b}\). This is why GPUs matter — the entire forward pass of a layer is a single dense matmul, the operation GPUs are built to do. Deep learning is, computationally, a pipeline of matrix multiplications separated by cheap nonlinearities.
1.3 Activation functions — and why they're non-negotiable
Compose two linear layers with no nonlinearity: \(W_2(W_1\mathbf{x})=(W_2W_1)\mathbf{x}=W'\mathbf{x}\). Still linear. Depth collapses to a single matrix. The nonlinearity \(\varphi\) is what breaks this collapse and gives the network its expressive power. The common choices:
| Activation | Formula | Why / when |
|---|---|---|
| Sigmoid | \(\sigma(z)=\frac{1}{1+e^{-z}}\) | Squashes to (0,1). Saturates → vanishing gradients. It is used in the MNIST net. |
| Tanh | \(\tanh(z)\) | Zero-centered (−1,1). Used inside LSTM gates. |
| ReLU | \(\max(0,z)\) | Cheap, non-saturating for \(z>0\). The workhorse of CNNs. Gradient is 0 or 1. |
| GELU | \(z\,\Phi(z)\) | Smooth ReLU; the default inside Transformers/GPT. |
| SwiGLU | \((xW)\otimes\text{Swish}(xV)\) | Gated variant used in LLaMA/PaLM-class LLMs — Part 6. |
1.4 Why depth: the Universal Approximation Theorem, honestly stated
A network with a single hidden layer of sufficient width can approximate any continuous function on a compact set to arbitrary accuracy. That sounds like depth is unnecessary — but the catch is width: shallow networks may need exponentially many neurons to represent functions that a deep network captures with polynomially many. Depth buys parameter efficiency and hierarchical features — edges → shapes → objects, or characters → words → meaning. This hierarchy is the entire reason "deep" beats "wide".
1.5 From scores to probabilities: softmax + cross-entropy
Classifiers output raw logits \(\mathbf{z}\); softmax turns them into a probability distribution, and cross-entropy scores it:
Softmax is a smooth, differentiable argmax: it exponentiates (making big scores dominate) then normalizes (so the outputs sum to 1). Cross-entropy is the negative log-likelihood of the true class — minimizing it maximizes the probability the model assigns to the right answer.
Part 2 · Foundations
Loss surfaces, backpropagation, and optimization
Training = searching a very high-dimensional loss surface \(\mathcal{L}(\theta)\) for a low point. We can't see the surface, but at any point we can compute the gradient \(\nabla_\theta \mathcal{L}\) — the direction of steepest ascent — and step the opposite way. The engine that computes that gradient efficiently is backpropagation.
2.1 Gradient descent, precisely
with learning rate \(\eta\). In practice we don't use the full dataset per step (too slow) — we estimate the gradient on a mini-batch. That noisy estimate is Stochastic Gradient Descent (SGD). The noise is a feature: it helps escape sharp minima and saddle points.
for k in gradients.keys():
weights[k] = weights[k] - self.learning_rate * gradients[k]
2.2 Backpropagation = the chain rule, organized
Consider your two-layer MNIST network. The forward pass is a chain of functions:
To update \(W_1\) we need \(\partial\mathcal{L}/\partial W_1\). The chain rule multiplies the local derivatives back along this path. Backprop computes them in one backward sweep, reusing intermediate results (this reuse is why it's \(O(\text{network size})\), not exponential).
2.3 Vanishing & exploding gradients
Backprop multiplies many Jacobians. If their singular values are consistently <1, the gradient shrinks toward zero over depth (vanishing); if >1, it blows up (exploding). This single fact motivates a huge amount of modern design: ReLU/GELU (gradient 1 for active units), careful initialization, normalization layers, and residual connections (which give gradients a direct \(+1\) path — the reason 100+ layer Transformers train at all).
2.4 Optimizers beyond vanilla SGD
| Optimizer | Update idea | Why it helps |
|---|---|---|
| SGD + Momentum | \(v\leftarrow\beta v+\nabla\mathcal{L};\ \theta\leftarrow\theta-\eta v\) | Accumulates velocity → smooths noise, accelerates down ravines. |
| RMSProp | Scale step by running RMS of gradients | Per-parameter adaptive rate; handles different curvatures. |
| Adam / AdamW | Momentum + RMSProp + bias correction | The default for Transformers. AdamW decouples weight decay for cleaner regularization. |
2.5 Regularization, normalization, initialization
- Weight decay / L2 (\(+\tfrac{\lambda}{2}\|\theta\|^2\)) and dropout (randomly zero activations) fight overfitting. You experimented with the reg strength in A1's
reg changes. - Normalization keeps activations well-scaled. BatchNorm normalizes across the batch (CNNs); LayerNorm normalizes across features per token (Transformers): \(\text{LN}(x)=\gamma\,\frac{x-\mu}{\sqrt{\sigma^2+\epsilon}}+\beta\). Your Transformer notes correctly flag LayerNorm as the choice because it's batch-size independent — critical for variable-length text.
- Initialization (Xavier/He) sets the initial weight variance so signal neither vanishes nor explodes at layer 1.
Part 3 · Foundations
Structured architectures: CNNs, RNNs, LSTMs
A fully-connected layer ignores structure in the input. Two structures dominate real data — space (images) and sequence (text, audio, time series) — and each gets a specialized layer that bakes in the right inductive bias.
3.1 Convolutions — weight sharing over space
A convolution slides a small learnable kernel \(K\) across the image, computing at each location a dot product:
Two inductive biases make this powerful: parameter sharing (the same kernel is reused at every location → drastically fewer weights than a dense layer) and translation equivariance (shift the input, the feature map shifts the same way). Stacking convolutions builds a hierarchy: early filters detect edges, deeper ones detect textures, parts, then whole objects. Pooling downsamples for translation invariance and a growing receptive field.
3.2 Recurrent networks — sharing weights over time
An RNN reads a sequence one step at a time, carrying a hidden state that summarizes the past:
The same weights \(W\) are reused at every timestep — weight sharing over time. Elegant, but training uses backpropagation through time, which unrolls the loop and multiplies the same Jacobian \(t\) times. If its largest eigenvalue is <1 the gradient vanishes; >1 it explodes. In practice a vanilla RNN forgets context beyond ~10–20 steps.
3.3 LSTMs — a protected memory highway
The LSTM adds a cell state \(\mathbf{c}_t\) — a conveyor belt gradients can traverse with minimal decay — controlled by three learned gates:
The key line is \(\mathbf{c}_t=\mathbf{f}_t\odot\mathbf{c}_{t-1}+\dots\): when the forget gate is ≈1, the cell state passes through nearly unchanged, so the gradient \(\partial\mathbf{c}_t/\partial\mathbf{c}_{t-1}\approx\mathbf{f}_t\) doesn't vanish. Gates decide what to erase, write, and read. This is why LSTMs powered translation, speech, and time-series forecasters for years.
Part 4 · Foundations
Attention, derived from scratch — with worked examples
Attention is the single most important mechanism in modern AI, so we'll build it slowly and with actual numbers. The one-sentence goal: rewrite each word's vector as a mix of the words most relevant to it. Let's make that concrete before any formula.
4.1 The dot product = a relevance meter
How do we measure "relevance" between two word vectors? The dot product. Recall \(\mathbf{a}\cdot\mathbf{b}=\|\mathbf{a}\|\,\|\mathbf{b}\|\cos\theta\): it's large and positive when two vectors point the same way, zero when perpendicular, negative when opposed. So a dot product is a similarity score we get for free.
Say three toy 2-D word vectors:
$$\mathbf{cat}=\begin{bmatrix}2\\[2pt]1\end{bmatrix},\quad \mathbf{dog}=\begin{bmatrix}2\\[2pt]0.8\end{bmatrix},\quad \mathbf{car}=\begin{bmatrix}-1\\[2pt]2\end{bmatrix}$$ $$\begin{aligned} \mathbf{cat}\cdot\mathbf{dog} &= (2)(2)+(1)(0.8) = 4.8 &&\text{high (both animals)}\\ \mathbf{cat}\cdot\mathbf{car} &= (2)(-1)+(1)(2) = 0.0 &&\text{unrelated}\\ \mathbf{dog}\cdot\mathbf{car} &= (2)(-1)+(0.8)(2) = -0.4 &&\text{unrelated} \end{aligned}$$The number literally ranks how related the words are. Attention is built on this meter.
4.2 Query, Key, Value — with a search-engine analogy
We don't compare raw word vectors directly. Instead each word produces three different views of itself through three learned matrices — like a search engine:
| Vector | Made by | Search-engine analogy |
|---|---|---|
| Query \(\mathbf q=W^Q\mathbf x\) | \(W^Q\) | the search box text — what this word is looking for |
| Key \(\mathbf k=W^K\mathbf x\) | \(W^K\) | the page title/tags — what each word advertises about itself |
| Value \(\mathbf v=W^V\mathbf x\) | \(W^V\) | the page content — what you actually get back if matched |
A word's Query is matched against every word's Key (dot products). The matches become weights. You then retrieve a weighted blend of everyone's Values. Why three separate matrices? Because "what I'm looking for" (query), "what I offer" (key), and "what I contribute" (value) are genuinely different roles — letting the model learn them independently is far more expressive than comparing raw vectors.
4.3 A fully worked self-attention pass (3 words, by hand)
Let's push real numbers all the way through for the sentence fragment "cat sat mat". We use 2-dimensional vectors so you can verify every step with a calculator.
For simplicity take \(W^Q=W^K=W^V=I\) (identity), so here \(Q=K=V=X\). In a real model these are learned and different; identity just lets us see the mechanics.
Each row is how much that word attends to every word. For example \(\ \text{score}(\text{cat},\text{mat})=\begin{bmatrix}1&0\end{bmatrix}\!\cdot\!\begin{bmatrix}1&1\end{bmatrix}=1\) and \(\ \text{score}(\text{mat},\text{mat})=\begin{bmatrix}1&1\end{bmatrix}\!\cdot\!\begin{bmatrix}1&1\end{bmatrix}=2.\)
Divide every entry by \(\sqrt{2}\approx1.414\). The "mat" row becomes:
$$\begin{bmatrix}1 & 1 & 2\end{bmatrix}\ \xrightarrow{\ \div\sqrt2\ }\ \begin{bmatrix}0.707 & 0.707 & 1.414\end{bmatrix}$$Take the scaled "mat" row \(\begin{bmatrix}0.707 & 0.707 & 1.414\end{bmatrix}\):
$$\text{softmax}(z)_i=\frac{e^{z_i}}{\sum_j e^{z_j}},\qquad \frac{\begin{bmatrix}e^{0.707} & e^{0.707} & e^{1.414}\end{bmatrix}}{e^{0.707}+e^{0.707}+e^{1.414}} =\frac{\begin{bmatrix}2.028 & 2.028 & 4.113\end{bmatrix}}{8.169}$$ $$=\begin{bmatrix}0.248 & 0.248 & 0.503\end{bmatrix}$$So "mat" pays \(\approx25\%\) attention to cat, \(25\%\) to sat, and \(50\%\) to itself.
The new "mat" vector is a context-aware blend — pulled toward cat and sat, but still mostly itself. Every word gets rewritten this way, in parallel, by the single matrix formula below.
Read it left to right exactly as we computed: \(QK^\top\) = all-pairs scores → divide by \(\sqrt{d_k}\) → softmax into weights → multiply by \(V\) to blend the values. That's the entire step you just did by hand, now as one line the GPU runs on the whole sentence at once.
4.4 Why the \(\sqrt{d_k}\) scaling? (the problem, then the fix)
Scores are sums of \(d_k\) products. With small \(d_k\) they stay modest; but real models use \(d_k=64\) or more, so scores get large. Watch what large scores do to softmax:
$$\begin{bmatrix}1 & 2 & 3\end{bmatrix}\ \xrightarrow{\text{softmax}}\ \begin{bmatrix}0.09 & 0.24 & 0.67\end{bmatrix}\quad(\text{soft, informative})$$ $$\begin{bmatrix}10 & 20 & 30\end{bmatrix}\ \xrightarrow{\text{softmax}}\ \begin{bmatrix}0.000 & 0.000 & 1.000\end{bmatrix}\quad(\text{razor-sharp, saturated})$$Once softmax saturates to nearly one-hot, its gradient is almost zero — the model can't learn to adjust the weights. Training stalls.
If each \(q_i,k_i\) has mean 0 and variance 1, then
$$\operatorname{Var}\!\left(\mathbf q\cdot\mathbf k\right)=\operatorname{Var}\!\left(\sum_{i=1}^{d_k}q_i k_i\right)=d_k \quad\Longrightarrow\quad \text{std}=\sqrt{d_k}.$$Dividing by \(\sqrt{d_k}\) rescales the scores back to variance 1 regardless of dimension, keeping softmax in its responsive, gradient-rich range. That single \(\tfrac{1}{\sqrt{d_k}}\) is why attention trains stably at large \(d_k\). This is a foundational point about scaled attention.
4.5 The same idea, in code
def attention(self, q, k, v):
hidden_dim = self.dim_k
head = torch.matmul(q, k.transpose(-2, -1)) # Step 2: QKᵀ, all-pairs scores
head = head / math.sqrt(hidden_dim) # Step 3: scale by √d_k
attn = self.softmax(head) # Step 4: softmax → weights
v_matrix = torch.matmul(attn, v) # Step 5: weighted sum of Values
return v_matrix, attn
The five hand-computed steps map one-to-one onto these four lines. This is exactly what we just did with pen and paper.
Appendix
References & further reading
The explanations here synthesize the primary research literature with a few exceptional public explainers. Grouped below by type; everything is freely available online.
Research papers & primary sources
- Learning representations by back-propagating errors. The backpropagation algorithm. https://www.nature.com/articles/323533a0
- Gradient-based learning applied to document recognition. Convolutional networks (LeNet). https://ieeexplore.ieee.org/document/726791
- Long Short-Term Memory. The LSTM and its gated cell state. https://www.bioinf.jku.at/publications/older/2604.pdf
- Deep Residual Learning for Image Recognition. Residual connections. https://arxiv.org/abs/1512.03385
- Adam: A Method for Stochastic Optimization. https://arxiv.org/abs/1412.6980
- Attention Is All You Need. Introduces scaled dot-product attention (covered in depth in Part 2). https://arxiv.org/abs/1706.03762
- Multilayer feedforward networks are universal approximators. The Universal Approximation Theorem. https://www.sciencedirect.com/science/article/abs/pii/0893608089900208
Visualizations & explainers
- The Illustrated Transformer. The vector-as-boxes visual vocabulary. https://jalammar.github.io/illustrated-transformer/
- Neural Networks (chapters 1–4). Geometric intuition for networks, gradient descent, and backprop. https://www.3blue1brown.com/topics/neural-networks
Further reading
- Dive into Deep Learning (d2l.ai). Free interactive textbook with runnable code. https://d2l.ai/
- Neural Networks: Zero to Hero. Builds micrograd, makemore, and a GPT from scratch. https://karpathy.ai/zero-to-hero.html
- Deep Learning. The standard graduate textbook. https://www.deeplearningbook.org/
Frontier-model details reflect publicly discussed methods; the exact architectures of proprietary systems (Claude, Gemini, GPT-4) are not officially disclosed and are described at the level of well-established public understanding.