Deep Learning, In Depth · Part 1 of 3

Foundations of Deep Learning

The ideas behind every LLM — from a single neuron to attention.

By PrithvirajPart 1 of 3~30 min read

Modern AI can feel like an impenetrable stack of jargon — attention, transformers, RLHF, mixture-of-experts. But almost all of it is built from one humble operation repeated at scale: multiply by a matrix, bend the result with a nonlinearity, and adjust the numbers so the output gets a little less wrong.

This is the first post in a three-part series that rebuilds deep learning from the ground up — not as a list of buzzwords, but as a small number of ideas you can hold in your head and reason from. In this post we cover the foundations: the neuron, why depth works, how networks learn, and the two classic architectures (CNNs and RNN/LSTMs) that set the stage for everything that came after. Part 2 dissects the Transformer in full. Part 3 shows how today's frontier models are built on exactly these bones.

Everything here is derived, not asserted. Where it helps, there are worked numerical examples you can check with a calculator, and short code snippets that map one-to-one onto the math.

How to read this series
Each idea appears three ways: the intuition (a picture or analogy), the math track (the precise statement), and a real-world implication (why it matters for models you actually use). Skim the intuitions on a first pass; come back for the math when you want the full picture.

Part 1 · Foundations

The neuron, linear algebra, and why depth works

Everything in deep learning is built from one operation repeated at scale: an affine transform followed by a nonlinearity. Master this one line and its calculus, and every later architecture is a rearrangement of it.

1.1 A single neuron

A neuron takes an input vector \(\mathbf{x}\in\mathbb{R}^d\), weights \(\mathbf{w}\in\mathbb{R}^d\), a bias \(b\), and produces a scalar activation:

$$a=\varphi\!\left(\sum_{i=1}^{d} w_i x_i + b\right)=\varphi(\mathbf{w}^\top\mathbf{x}+b)$$

The dot product \(\mathbf{w}^\top\mathbf{x}=\|\mathbf{w}\|\,\|\mathbf{x}\|\cos\theta\) measures alignment between input and weights. A neuron is a learned template matcher: it fires strongly when the input points in the direction of \(\mathbf{w}\). Keep this geometric reading — it returns in attention as the query–key dot product.

1.2 A layer is a matrix multiply

Stack \(m\) neurons and their weight vectors become the rows of a matrix \(W\in\mathbb{R}^{m\times d}\). A whole layer is then one matrix–vector product plus a broadcast bias:

$$\mathbf{z}=W\mathbf{x}+\mathbf{b},\qquad \mathbf{a}=\varphi(\mathbf{z})$$

Process a batch of \(N\) inputs at once by stacking them into \(X\in\mathbb{R}^{N\times d}\): \(Z=XW^\top+\mathbf{b}\). This is why GPUs matter — the entire forward pass of a layer is a single dense matmul, the operation GPUs are built to do. Deep learning is, computationally, a pipeline of matrix multiplications separated by cheap nonlinearities.

Real-world implication
A frontier model's cost is dominated by these matmuls. "175B parameters" essentially counts the entries of all the \(W\) matrices; each token's forward pass multiplies through every one of them. This is why inference is expensive and why hardware (H100s, TPUs) and low-precision arithmetic (fp16/bf16/fp8) are strategic.

1.3 Activation functions — and why they're non-negotiable

Compose two linear layers with no nonlinearity: \(W_2(W_1\mathbf{x})=(W_2W_1)\mathbf{x}=W'\mathbf{x}\). Still linear. Depth collapses to a single matrix. The nonlinearity \(\varphi\) is what breaks this collapse and gives the network its expressive power. The common choices:

ActivationFormulaWhy / when
Sigmoid\(\sigma(z)=\frac{1}{1+e^{-z}}\)Squashes to (0,1). Saturates → vanishing gradients. It is used in the MNIST net.
Tanh\(\tanh(z)\)Zero-centered (−1,1). Used inside LSTM gates.
ReLU\(\max(0,z)\)Cheap, non-saturating for \(z>0\). The workhorse of CNNs. Gradient is 0 or 1.
GELU\(z\,\Phi(z)\)Smooth ReLU; the default inside Transformers/GPT.
SwiGLU\((xW)\otimes\text{Swish}(xV)\)Gated variant used in LLaMA/PaLM-class LLMs — Part 6.
Math track — derivative of sigmoid
A fact you'll reuse: \(\sigma'(z)=\sigma(z)\,(1-\sigma(z))\). Notice it's maximal (0.25) at \(z=0\) and →0 as \(|z|\) grows. Chain many of these together and gradients shrink geometrically — the vanishing gradient problem that motivated ReLU, LSTMs, and residual connections.

1.4 Why depth: the Universal Approximation Theorem, honestly stated

A network with a single hidden layer of sufficient width can approximate any continuous function on a compact set to arbitrary accuracy. That sounds like depth is unnecessary — but the catch is width: shallow networks may need exponentially many neurons to represent functions that a deep network captures with polynomially many. Depth buys parameter efficiency and hierarchical features — edges → shapes → objects, or characters → words → meaning. This hierarchy is the entire reason "deep" beats "wide".

1.5 From scores to probabilities: softmax + cross-entropy

Classifiers output raw logits \(\mathbf{z}\); softmax turns them into a probability distribution, and cross-entropy scores it:

$$p_i=\frac{e^{z_i}}{\sum_j e^{z_j}},\qquad \mathcal{L}_{\text{CE}}=-\sum_i y_i\log p_i$$

Softmax is a smooth, differentiable argmax: it exponentiates (making big scores dominate) then normalizes (so the outputs sum to 1). Cross-entropy is the negative log-likelihood of the true class — minimizing it maximizes the probability the model assigns to the right answer.

Intuition — the same trio, four times
Softmax + cross-entropy is not just a classifier tail. It reappears as the attention weighting (softmax over similarity scores), the LLM output head (softmax over the vocabulary), and the sampling step in generation. If you deeply understand this one page, you understand a surprising fraction of an LLM.

Part 2 · Foundations

Loss surfaces, backpropagation, and optimization

Training = searching a very high-dimensional loss surface \(\mathcal{L}(\theta)\) for a low point. We can't see the surface, but at any point we can compute the gradient \(\nabla_\theta \mathcal{L}\) — the direction of steepest ascent — and step the opposite way. The engine that computes that gradient efficiently is backpropagation.

2.1 Gradient descent, precisely

$$\theta \leftarrow \theta - \eta\,\nabla_\theta \mathcal{L}(\theta)$$

with learning rate \(\eta\). In practice we don't use the full dataset per step (too slow) — we estimate the gradient on a mini-batch. That noisy estimate is Stochastic Gradient Descent (SGD). The noise is a feature: it helps escape sharp minima and saddle points.

for k in gradients.keys():
    weights[k] = weights[k] - self.learning_rate * gradients[k]

2.2 Backpropagation = the chain rule, organized

Consider your two-layer MNIST network. The forward pass is a chain of functions:

$$\mathbf{x}\xrightarrow{W_1,b_1}\mathbf{z}_1\xrightarrow{\sigma}\mathbf{s}\xrightarrow{W_2,b_2}\mathbf{z}_2\xrightarrow{\text{softmax}}\mathbf{p}\xrightarrow{\text{CE}}\mathcal{L}$$

To update \(W_1\) we need \(\partial\mathcal{L}/\partial W_1\). The chain rule multiplies the local derivatives back along this path. Backprop computes them in one backward sweep, reusing intermediate results (this reuse is why it's \(O(\text{network size})\), not exponential).

Math track — the beautiful softmax+CE gradient (your derivation)
The single most important gradient in classification. Here is the derivation: with \(p=\text{softmax}(z_2)\) and one-hot target \(y\), $$\frac{\partial \mathcal{L}_{\text{CE}}}{\partial \mathbf{z}_2}=\mathbf{p}-\mathbf{y}.$$ The messy \(-y/p\) from the CE derivative and the softmax Jacobian cancel exactly, leaving "prediction minus target". Then by the chain rule \(\dfrac{\partial\mathcal{L}}{\partial W_2}=(\mathbf{p}-\mathbf{y})\,\mathbf{s}^\top\). This clean form is why softmax and cross-entropy are always paired.

2.3 Vanishing & exploding gradients

Backprop multiplies many Jacobians. If their singular values are consistently <1, the gradient shrinks toward zero over depth (vanishing); if >1, it blows up (exploding). This single fact motivates a huge amount of modern design: ReLU/GELU (gradient 1 for active units), careful initialization, normalization layers, and residual connections (which give gradients a direct \(+1\) path — the reason 100+ layer Transformers train at all).

2.4 Optimizers beyond vanilla SGD

OptimizerUpdate ideaWhy it helps
SGD + Momentum\(v\leftarrow\beta v+\nabla\mathcal{L};\ \theta\leftarrow\theta-\eta v\)Accumulates velocity → smooths noise, accelerates down ravines.
RMSPropScale step by running RMS of gradientsPer-parameter adaptive rate; handles different curvatures.
Adam / AdamWMomentum + RMSProp + bias correctionThe default for Transformers. AdamW decouples weight decay for cleaner regularization.
Math track — Adam
$$m_t=\beta_1 m_{t-1}+(1-\beta_1)g_t,\quad v_t=\beta_2 v_{t-1}+(1-\beta_2)g_t^2$$ $$\hat m_t=\tfrac{m_t}{1-\beta_1^t},\ \hat v_t=\tfrac{v_t}{1-\beta_2^t},\qquad \theta\leftarrow\theta-\eta\,\tfrac{\hat m_t}{\sqrt{\hat v_t}+\epsilon}$$ \(m_t\) is a smoothed gradient (momentum); \(v_t\) tracks its variance so each parameter gets its own effective step size. Nearly every LLM is trained with AdamW.

2.5 Regularization, normalization, initialization

  • Weight decay / L2 (\(+\tfrac{\lambda}{2}\|\theta\|^2\)) and dropout (randomly zero activations) fight overfitting. You experimented with the reg strength in A1's reg changes.
  • Normalization keeps activations well-scaled. BatchNorm normalizes across the batch (CNNs); LayerNorm normalizes across features per token (Transformers): \(\text{LN}(x)=\gamma\,\frac{x-\mu}{\sqrt{\sigma^2+\epsilon}}+\beta\). Your Transformer notes correctly flag LayerNorm as the choice because it's batch-size independent — critical for variable-length text.
  • Initialization (Xavier/He) sets the initial weight variance so signal neither vanishes nor explodes at layer 1.

Part 3 · Foundations

Structured architectures: CNNs, RNNs, LSTMs

A fully-connected layer ignores structure in the input. Two structures dominate real data — space (images) and sequence (text, audio, time series) — and each gets a specialized layer that bakes in the right inductive bias.

3.1 Convolutions — weight sharing over space

A convolution slides a small learnable kernel \(K\) across the image, computing at each location a dot product:

$$(I*K)_{ij}=\sum_{m}\sum_{n} I_{i+m,\,j+n}\,K_{m,n}$$

Two inductive biases make this powerful: parameter sharing (the same kernel is reused at every location → drastically fewer weights than a dense layer) and translation equivariance (shift the input, the feature map shifts the same way). Stacking convolutions builds a hierarchy: early filters detect edges, deeper ones detect textures, parts, then whole objects. Pooling downsamples for translation invariance and a growing receptive field.

Real-world implication
The same "share weights across positions" idea is exactly what attention does across sequence positions. And Vision Transformers now cut images into patches and feed them to a Transformer — so the CNN's spatial prior and the Transformer's global attention have partly merged in frontier multimodal models.

3.2 Recurrent networks — sharing weights over time

An RNN reads a sequence one step at a time, carrying a hidden state that summarizes the past:

$$\mathbf{h}_t=\tanh\!\big(W_{hh}\mathbf{h}_{t-1}+W_{xh}\mathbf{x}_t+\mathbf{b}\big),\qquad \mathbf{y}_t=W_{hy}\mathbf{h}_t$$

The same weights \(W\) are reused at every timestep — weight sharing over time. Elegant, but training uses backpropagation through time, which unrolls the loop and multiplies the same Jacobian \(t\) times. If its largest eigenvalue is <1 the gradient vanishes; >1 it explodes. In practice a vanilla RNN forgets context beyond ~10–20 steps.

3.3 LSTMs — a protected memory highway

The LSTM adds a cell state \(\mathbf{c}_t\) — a conveyor belt gradients can traverse with minimal decay — controlled by three learned gates:

$$\begin{aligned}\mathbf{f}_t&=\sigma(W_f[\mathbf{h}_{t-1},\mathbf{x}_t]+b_f)\quad(\text{forget})\\ \mathbf{i}_t&=\sigma(W_i[\cdots]+b_i)\quad(\text{input})\\ \mathbf{o}_t&=\sigma(W_o[\cdots]+b_o)\quad(\text{output})\\ \tilde{\mathbf{c}}_t&=\tanh(W_c[\cdots]+b_c)\\ \mathbf{c}_t&=\mathbf{f}_t\odot\mathbf{c}_{t-1}+\mathbf{i}_t\odot\tilde{\mathbf{c}}_t\\ \mathbf{h}_t&=\mathbf{o}_t\odot\tanh(\mathbf{c}_t)\end{aligned}$$

The key line is \(\mathbf{c}_t=\mathbf{f}_t\odot\mathbf{c}_{t-1}+\dots\): when the forget gate is ≈1, the cell state passes through nearly unchanged, so the gradient \(\partial\mathbf{c}_t/\partial\mathbf{c}_{t-1}\approx\mathbf{f}_t\) doesn't vanish. Gates decide what to erase, write, and read. This is why LSTMs powered translation, speech, and time-series forecasters for years.

Check: Why does the LSTM cell-state update use addition (\(c_{t-1}+\dots\)) rather than a matrix multiply like the RNN? (Addition gives gradients a near-identity path backward → no vanishing. The multiplicative RNN recurrence is exactly what kills long-range gradients.)

Part 4 · Foundations

Attention, derived from scratch — with worked examples

Attention is the single most important mechanism in modern AI, so we'll build it slowly and with actual numbers. The one-sentence goal: rewrite each word's vector as a mix of the words most relevant to it. Let's make that concrete before any formula.

The intuition, with a sentence
Take "The animal didn't cross the street because it was tired." What does "it" refer to — the animal or the street? You resolve this by letting "it" look at the other words and decide "animal" is most relevant. Attention does exactly this: for every word, it computes a relevance score to every other word, then builds a new representation of that word by averaging the others, weighted by relevance. "it" ends up as mostly-"animal", a little "tired". That re-mixing is the whole idea.

4.1 The dot product = a relevance meter

How do we measure "relevance" between two word vectors? The dot product. Recall \(\mathbf{a}\cdot\mathbf{b}=\|\mathbf{a}\|\,\|\mathbf{b}\|\cos\theta\): it's large and positive when two vectors point the same way, zero when perpendicular, negative when opposed. So a dot product is a similarity score we get for free.

Worked example — dot product as similarity

Say three toy 2-D word vectors:

$$\mathbf{cat}=\begin{bmatrix}2\\[2pt]1\end{bmatrix},\quad \mathbf{dog}=\begin{bmatrix}2\\[2pt]0.8\end{bmatrix},\quad \mathbf{car}=\begin{bmatrix}-1\\[2pt]2\end{bmatrix}$$ $$\begin{aligned} \mathbf{cat}\cdot\mathbf{dog} &= (2)(2)+(1)(0.8) = 4.8 &&\text{high (both animals)}\\ \mathbf{cat}\cdot\mathbf{car} &= (2)(-1)+(1)(2) = 0.0 &&\text{unrelated}\\ \mathbf{dog}\cdot\mathbf{car} &= (2)(-1)+(0.8)(2) = -0.4 &&\text{unrelated} \end{aligned}$$

The number literally ranks how related the words are. Attention is built on this meter.

4.2 Query, Key, Value — with a search-engine analogy

We don't compare raw word vectors directly. Instead each word produces three different views of itself through three learned matrices — like a search engine:

VectorMade bySearch-engine analogy
Query \(\mathbf q=W^Q\mathbf x\)\(W^Q\)the search box text — what this word is looking for
Key \(\mathbf k=W^K\mathbf x\)\(W^K\)the page title/tags — what each word advertises about itself
Value \(\mathbf v=W^V\mathbf x\)\(W^V\)the page content — what you actually get back if matched

A word's Query is matched against every word's Key (dot products). The matches become weights. You then retrieve a weighted blend of everyone's Values. Why three separate matrices? Because "what I'm looking for" (query), "what I offer" (key), and "what I contribute" (value) are genuinely different roles — letting the model learn them independently is far more expressive than comparing raw vectors.

4.3 A fully worked self-attention pass (3 words, by hand)

Let's push real numbers all the way through for the sentence fragment "cat sat mat". We use 2-dimensional vectors so you can verify every step with a calculator.

Step 0 — the word embeddings (input \(X\))
$$X=\begin{bmatrix} 1 & 0\\ 0 & 1\\ 1 & 1\end{bmatrix}\ \begin{matrix}\leftarrow\ \text{cat}\\ \leftarrow\ \text{sat}\\ \leftarrow\ \text{mat}\end{matrix}\qquad(3\times2)$$
Step 1 — project to \(Q,K,V\)

For simplicity take \(W^Q=W^K=W^V=I\) (identity), so here \(Q=K=V=X\). In a real model these are learned and different; identity just lets us see the mechanics.

Step 2 — scores \(=QK^\top\) (every query \(\cdot\) every key)
$$QK^\top=\begin{bmatrix} 1 & 0\\ 0 & 1\\ 1 & 1\end{bmatrix}\!\begin{bmatrix} 1 & 0 & 1\\ 0 & 1 & 1\end{bmatrix}= \begin{array}{c} \begin{matrix}\ \ \scriptstyle k=\text{cat} & \scriptstyle k=\text{sat} & \scriptstyle k=\text{mat}\end{matrix}\\ \begin{bmatrix} 1 & 0 & 1\\ 0 & 1 & 1\\ 1 & 1 & 2\end{bmatrix} \end{array} \begin{matrix}\scriptstyle q=\text{cat}\\ \scriptstyle q=\text{sat}\\ \scriptstyle q=\text{mat}\end{matrix}$$

Each row is how much that word attends to every word. For example \(\ \text{score}(\text{cat},\text{mat})=\begin{bmatrix}1&0\end{bmatrix}\!\cdot\!\begin{bmatrix}1&1\end{bmatrix}=1\) and \(\ \text{score}(\text{mat},\text{mat})=\begin{bmatrix}1&1\end{bmatrix}\!\cdot\!\begin{bmatrix}1&1\end{bmatrix}=2.\)

Step 3 — scale by \(\sqrt{d_k}\) (here \(d_k=2\))

Divide every entry by \(\sqrt{2}\approx1.414\). The "mat" row becomes:

$$\begin{bmatrix}1 & 1 & 2\end{bmatrix}\ \xrightarrow{\ \div\sqrt2\ }\ \begin{bmatrix}0.707 & 0.707 & 1.414\end{bmatrix}$$
Step 4 — softmax each row \(\rightarrow\) attention weights (rows sum to 1)

Take the scaled "mat" row \(\begin{bmatrix}0.707 & 0.707 & 1.414\end{bmatrix}\):

$$\text{softmax}(z)_i=\frac{e^{z_i}}{\sum_j e^{z_j}},\qquad \frac{\begin{bmatrix}e^{0.707} & e^{0.707} & e^{1.414}\end{bmatrix}}{e^{0.707}+e^{0.707}+e^{1.414}} =\frac{\begin{bmatrix}2.028 & 2.028 & 4.113\end{bmatrix}}{8.169}$$ $$=\begin{bmatrix}0.248 & 0.248 & 0.503\end{bmatrix}$$

So "mat" pays \(\approx25\%\) attention to cat, \(25\%\) to sat, and \(50\%\) to itself.

Step 5 — output \(=\) weights \(\cdot\,V\) (weighted blend of value vectors)
$$\text{new }\mathbf{mat}=0.248\begin{bmatrix}1\\0\end{bmatrix}+0.248\begin{bmatrix}0\\1\end{bmatrix}+0.503\begin{bmatrix}1\\1\end{bmatrix} =\begin{bmatrix}0.248+0.503\\ 0.248+0.503\end{bmatrix}=\begin{bmatrix}0.751\\ 0.751\end{bmatrix}$$

The new "mat" vector is a context-aware blend — pulled toward cat and sat, but still mostly itself. Every word gets rewritten this way, in parallel, by the single matrix formula below.

$$\text{Attention}(Q,K,V)=\text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V$$

Read it left to right exactly as we computed: \(QK^\top\) = all-pairs scores → divide by \(\sqrt{d_k}\) → softmax into weights → multiply by \(V\) to blend the values. That's the entire step you just did by hand, now as one line the GPU runs on the whole sentence at once.

4.4 Why the \(\sqrt{d_k}\) scaling? (the problem, then the fix)

The problem — with numbers

Scores are sums of \(d_k\) products. With small \(d_k\) they stay modest; but real models use \(d_k=64\) or more, so scores get large. Watch what large scores do to softmax:

$$\begin{bmatrix}1 & 2 & 3\end{bmatrix}\ \xrightarrow{\text{softmax}}\ \begin{bmatrix}0.09 & 0.24 & 0.67\end{bmatrix}\quad(\text{soft, informative})$$ $$\begin{bmatrix}10 & 20 & 30\end{bmatrix}\ \xrightarrow{\text{softmax}}\ \begin{bmatrix}0.000 & 0.000 & 1.000\end{bmatrix}\quad(\text{razor-sharp, saturated})$$

Once softmax saturates to nearly one-hot, its gradient is almost zero — the model can't learn to adjust the weights. Training stalls.

The fix — variance argument

If each \(q_i,k_i\) has mean 0 and variance 1, then

$$\operatorname{Var}\!\left(\mathbf q\cdot\mathbf k\right)=\operatorname{Var}\!\left(\sum_{i=1}^{d_k}q_i k_i\right)=d_k \quad\Longrightarrow\quad \text{std}=\sqrt{d_k}.$$

Dividing by \(\sqrt{d_k}\) rescales the scores back to variance 1 regardless of dimension, keeping softmax in its responsive, gradient-rich range. That single \(\tfrac{1}{\sqrt{d_k}}\) is why attention trains stably at large \(d_k\). This is a foundational point about scaled attention.

4.5 The same idea, in code

def attention(self, q, k, v):
    hidden_dim = self.dim_k
    head = torch.matmul(q, k.transpose(-2, -1))   # Step 2: QKᵀ, all-pairs scores
    head = head / math.sqrt(hidden_dim)           # Step 3: scale by √d_k
    attn = self.softmax(head)                      # Step 4: softmax → weights
    v_matrix = torch.matmul(attn, v)               # Step 5: weighted sum of Values
    return v_matrix, attn

The five hand-computed steps map one-to-one onto these four lines. This is exactly what we just did with pen and paper.

Real-world implication — the O(n²) wall
\(QK^\top\) is an \(n\times n\) score matrix for \(n\) tokens. Double the context and you quadruple this matrix. That quadratic cost is the central constraint of modern LLMs — the reason 100K–1M-token context windows are hard and expensive, and the reason FlashAttention, sparse/linear attention, and KV-caching (Part 6) exist. Every "longer context" milestone is a fight against this \(n^2\).
Check: In Step 5, why do we multiply weights by \(V\) rather than by the original embeddings \(X\)? (Because Values are a learned "what I contribute if selected" view — separating retrieval content from the key used for matching lets the model optimize each role independently.)

Appendix

References & further reading

The explanations here synthesize the primary research literature with a few exceptional public explainers. Grouped below by type; everything is freely available online.

Research papers & primary sources

  1. Rumelhart, Hinton & Williams (1986). Learning representations by back-propagating errors. The backpropagation algorithm. https://www.nature.com/articles/323533a0
  2. LeCun, Bottou, Bengio & Haffner (1998). Gradient-based learning applied to document recognition. Convolutional networks (LeNet). https://ieeexplore.ieee.org/document/726791
  3. Hochreiter & Schmidhuber (1997). Long Short-Term Memory. The LSTM and its gated cell state. https://www.bioinf.jku.at/publications/older/2604.pdf
  4. He, Zhang, Ren & Sun (2015). Deep Residual Learning for Image Recognition. Residual connections. https://arxiv.org/abs/1512.03385
  5. Kingma & Ba (2014). Adam: A Method for Stochastic Optimization. https://arxiv.org/abs/1412.6980
  6. Vaswani et al. (2017). Attention Is All You Need. Introduces scaled dot-product attention (covered in depth in Part 2). https://arxiv.org/abs/1706.03762
  7. Hornik, Stinchcombe & White (1989). Multilayer feedforward networks are universal approximators. The Universal Approximation Theorem. https://www.sciencedirect.com/science/article/abs/pii/0893608089900208

Visualizations & explainers

  1. Jay Alammar. The Illustrated Transformer. The vector-as-boxes visual vocabulary. https://jalammar.github.io/illustrated-transformer/
  2. 3Blue1Brown (Grant Sanderson). Neural Networks (chapters 1–4). Geometric intuition for networks, gradient descent, and backprop. https://www.3blue1brown.com/topics/neural-networks

Further reading

  1. Zhang, Lipton, Li & Smola. Dive into Deep Learning (d2l.ai). Free interactive textbook with runnable code. https://d2l.ai/
  2. Andrej Karpathy. Neural Networks: Zero to Hero. Builds micrograd, makemore, and a GPT from scratch. https://karpathy.ai/zero-to-hero.html
  3. Goodfellow, Bengio & Courville (2016). Deep Learning. The standard graduate textbook. https://www.deeplearningbook.org/

Frontier-model details reflect publicly discussed methods; the exact architectures of proprietary systems (Claude, Gemini, GPT-4) are not officially disclosed and are described at the level of well-established public understanding.