...except when it is. The subtle confusion that led to AdamW.
The problem with SGD: vanilla stochastic gradient descent uses a single, fixed learning rate for every parameter. On loss surfaces with different curvature in different directions (extremely common in neural networks), this causes oscillation along steep axes while making glacially slow progress along flat ones — the classic "zigzag" problem.
Adam (Kingma & Ba, 2014) solves this by maintaining per-parameter adaptive learning rates using two running averages of gradient history: the first moment (mean of gradients, like momentum) and the second moment (mean of squared gradients, measuring variance). Dividing by √(second moment) normalizes each parameter's step size by its gradient scale.
Moving averages (per parameter):
m
t
= β₁ · m
t-1
+ (1 − β₁) · g
t
← first moment (momentum)
v
t
= β₂ · v
t-1
+ (1 − β₂) · g
t
²
← second moment (variance)
Bias correction + update:
m̂
t
= m
t
/ (1 − β₁
t
)
← compensate zero-init bias
v̂
t
= v
t
/ (1 − β₂
t
)
w ← w − lr · m̂
t
/ (√v̂
t
+ ε)
← the adaptive step
An exponential moving average of past gradients. Instead of following the current gradient's noisy direction, the optimizer accumulates a smoothed direction (β₁ = 0.9 typical). This dampens oscillations and sustains progress along consistent gradient directions — like a ball rolling with inertia.
An exponential moving average of squared gradients (β₂ = 0.999 typical). Dividing by √v̂ gives each parameter its own effective learning rate: parameters with historically large gradients get smaller steps, parameters with small gradients get larger steps. This is what eliminates the zigzag.
The elliptical contours matter: the loss surface has 3× steeper curvature along w₂ than w₁. SGD's single learning rate causes it to overshoot on the steep axis, creating the characteristic zigzag. Adam's per-parameter 1/√v̂ scaling automatically compensates — it takes shorter steps where gradients are large and longer steps where they're small, producing a much more direct path.
Why this matters for the rest of this tutorial: Adam's per-parameter adaptive scaling is exactly what creates the divergence between weight decay and L2 regularization. When a penalty term flows through this adaptive machinery, it gets transformed differently than when applied directly to the weights.
Neural networks are powerful function approximators — sometimes too powerful. Given enough parameters, a model can memorize the training data perfectly, including its noise and idiosyncrasies. This is overfitting : the model performs well on training data but fails to generalize to new, unseen data.
Regularization is any technique that constrains or penalizes the model to prevent overfitting and encourage generalization. The core idea: simpler models (smaller weights, smoother functions) tend to generalize better. Regularization nudges the optimizer toward these simpler solutions.
Common forms include dropout (randomly zeroing activations), data augmentation (expanding training distribution), early stopping (halting before overfitting), and — the focus of this tutorial — weight-based penalties that discourage the model from developing excessively large parameter values.
The name comes from the L2 norm (also called the Euclidean norm), which measures the "length" of a vector. For a weight vector w , the Lp norms are defined as:
L2 regularization adds the squared L2 norm of the weights to the loss function: λ · ‖w‖₂² . We use the squared norm (not the raw norm) because its gradient is clean and differentiable everywhere — it yields 2λw , a simple linear push toward zero proportional to each weight's magnitude.
Penalty: λ · ‖w‖₁ = λ · Σ|wᵢ|
Produces
sparse
solutions — many weights become exactly zero. Useful for feature selection. The gradient is
±λ (constant magnitude), which pushes small weights all the way to zero.
Penalty: λ · ‖w‖₂² = λ · Σwᵢ²
Produces
small but non-zero
weights — a smooth shrinkage toward zero. The gradient is 2λw (proportional to weight), so
large weights get penalized more heavily. This is the one commonly confused with weight
decay.
Both L2 regularization and weight decay push weights toward zero. With standard SGD, they produce mathematically identical updates (as we'll see in the next sections). This led the field to treat them as synonyms for years. But they are conceptually different mechanisms — and with adaptive optimizers like Adam, they produce genuinely different behavior. Understanding why requires looking at each one carefully.
Dropout is a strikingly simple regularization technique: during each training forward pass, randomly set each neuron's output to zero with probability p (typically 0.1–0.5). The dropped neurons change every pass, so the network can never rely on any single neuron or co-adaptation between specific neurons.
The key insight: this is equivalent to training an exponentially large ensemble of "thinned" sub-networks that share weights. At inference, all neurons are active but outputs are scaled by (1 − p) to match the expected training-time activation magnitude. The result: better generalization through redundant, distributed representations.
In transformers, dropout was historically applied in several locations: after the attention softmax (attention dropout), after each sub-layer before the residual connection, and within the feedforward MLP blocks. The original "Attention Is All You Need" paper used p = 0.1 throughout.
A network with n units can produce 2 n possible thinned sub-networks via dropout. Each training batch effectively trains a different sub-network. Inference with scaled weights approximates the geometric mean of all these sub-networks — an efficient ensemble without the cost of training separate models.
Training:
For each neuron, sample r ~ Bernoulli(1 − p). Output becomes: ŷ = r · y. On average, a fraction
p of neurons are zeroed each pass. The surviving neurons' gradients are amplified by 1/(1 − p)
to maintain expected magnitude (inverted dropout).
Inference:
No dropout applied — all neurons active. With inverted dropout (the modern default), no scaling
adjustment is needed at inference time since the training-time scaling already compensated.
The original idea from Hanson & Pratt (1988): every update step, multiply each weight by a factor slightly less than 1 . That's it. No loss function modification — just shrink the weights.
A different approach: add a penalty term directly to the loss function that penalizes large weights. The gradient of this penalty then naturally pushes weights toward zero during optimization.
The regularization gradient 2λ·w gets mixed into the total gradient — this matters for adaptive optimizers!
Weight decay
acts
directly on the weights
— it's a multiplicative shrinkage that doesn't touch the gradient at all.
L2 regularization
modifies
the gradient
by adding a term to the loss function — the regularization signal flows through the same
gradient pipeline as the task loss.
For vanilla SGD, these produce identical update rules (just rescale λ). But for any optimizer
that
transforms
the gradient before applying it — like Adam's per-parameter adaptive learning rates — they
diverge.
Weight decay:
w ← w − lr·∇L(w) −
lr·λ·w
L2 regularization:
w ← w − lr·(∇L(w) +
2λ·w
) = w − lr·∇L(w) −
lr·2λ·w
Set
λ_L2
=
λ_WD
/2 and they produce
the exact same parameter update
. This algebraic equivalence led the field to treat them as synonyms for decades.
Why they match: SGD applies the raw gradient directly — there is no transformation between the gradient computation and the weight update. Whether you add the penalty to the loss (L2) or subtract it from the weights (decay), the final Δw is identical. The solid and dashed lines overlap completely.
Adam divides the gradient by √(v̂) — a running estimate of gradient variance. This is per-parameter. When L2 regularization injects the penalty into the gradient , it gets divided by √(v̂) too. The regularization strength becomes inversely proportional to gradient magnitude .
The 2λw penalty gets scaled by 1/√(v̂). Large-gradient params: regularization gets weakened . Small-gradient params: regularization gets amplified . The decay becomes non-uniform across parameters.
λw applied after Adam's adaptive update — never passes through the 1/√(v̂) normalization. Every parameter decays by the same proportion , regardless of gradient scale.
Why they diverge: The elliptical contours represent different curvature per axis (common in real networks). Adam normalizes gradients per-parameter to handle exactly this. But when L2's penalty is inside the gradient, it gets normalized too — so the regularization strength becomes entangled with the loss landscape geometry. True weight decay acts on the raw weights, keeping the regularization clean and geometry-independent. The dashed bracket shows the gap between the two endpoints.
The landmark paper that identified the problem. They showed that virtually every implementation of "weight decay" in deep learning was actually doing L2 regularization — feeding the penalty into the gradient where Adam's adaptive scaling would distort it.
The fix was remarkably simple: decouple the weight decay from the gradient-based update . Apply Adam's adaptive update first, then subtract the decay term separately.
Loshchilov & Hutter showed that proper decoupling improved generalization across the board.
AdamW is now the default optimizer in most transformer training — from BERT to GPT to most
large-scale training since 2019.
The lesson: a "trivial" algebraic equivalence in one setting (SGD) became a meaningful
performance difference in another (Adam). The field spent years applying L2 regularization
thinking it was weight decay, and this confusion measurably held back training quality until
Loshchilov and Hutter formalized the distinction.
The regularization landscape looks very different for modern large language models compared to the deep learning of the 2010s. The scale of data and the single-epoch training regime have fundamentally changed which techniques matter.
AdamW with decoupled weight decay remains the default optimizer for virtually all frontier LLM pretraining. Typical weight decay values: 0.01–0.1. One notable refinement: recent work excludes embeddings and normalization layers from weight decay, as decaying these can hurt training stability.
GPT-3, PaLM, LLaMA, Chinchilla, and Gopher all train with dropout = 0 . LLaMA's config defaults attention_dropout to 0.0. Reason: modern pretraining sees each token only once (single-epoch), so there is minimal overfitting risk. Dropout just slows convergence with no regularization benefit. Still sometimes used in fine-tuning on small datasets.
Since AdamW became the standard, L2 regularization (penalty in the loss) has been effectively superseded by proper decoupled weight decay. No major frontier model uses L2 regularization directly — the Loshchilov & Hutter paper resolved this definitively.
Data scale as implicit regularization:
When training on trillions of unique tokens in a single epoch, the model never sees the same
data twice. This eliminates the primary overfitting mechanism that dropout and L2 were designed
to combat.
Weight decay via AdamW:
The one classical regularizer that survived — and for good reason. It prevents weight magnitude
explosion without distorting the adaptive gradient signal, and it helps with training stability
independent of overfitting concerns.
Architectural choices:
Modern regularization is increasingly built into the architecture itself — RMSNorm, QK-norm for
attention stability, logit soft-capping (Gemma-style), and careful initialization schemes. These
provide training stability without the generalization-vs-convergence tradeoffs of classical
regularizers.
Gradient clipping:
Not a regularizer in the classical sense, but gradient norm clipping (typically at 1.0) is
universally applied in LLM training to prevent loss spikes and training instability.
While pretraining has moved past dropout and L2, fine-tuning on small datasets still benefits from classical regularization. LoRA fine-tuning commonly uses dropout (0.05–0.1) on the adapter weights, stronger weight decay (0.01–0.1), and early stopping. The data scarcity that triggers overfitting in fine-tuning is precisely the regime these techniques were designed for.