Weight Decay ≠ L2 Regularization

...except when it is. The subtle confusion that led to AdamW.

Adam: Adaptive Moment Estimation

The problem with SGD: vanilla stochastic gradient descent uses a single, fixed learning rate for every parameter. On loss surfaces with different curvature in different directions (extremely common in neural networks), this causes oscillation along steep axes while making glacially slow progress along flat ones — the classic "zigzag" problem.

Adam (Kingma & Ba, 2014) solves this by maintaining per-parameter adaptive learning rates using two running averages of gradient history: the first moment (mean of gradients, like momentum) and the second moment (mean of squared gradients, measuring variance). Dividing by √(second moment) normalizes each parameter's step size by its gradient scale.

The Adam Update Rule

Moving averages (per parameter):

m t = β₁ · m t-1 + (1 − β₁) · g t ← first moment (momentum)
v t = β₂ · v t-1 + (1 − β₂) · g t ² ← second moment (variance)

Bias correction + update:

m̂ t = m t / (1 − β₁ t ) ← compensate zero-init bias
v̂ t = v t / (1 − β₂ t )
w ← w − lr · m̂ t / (√v̂ t + ε) ← the adaptive step

First Moment m: Momentum

An exponential moving average of past gradients. Instead of following the current gradient's noisy direction, the optimizer accumulates a smoothed direction (β₁ = 0.9 typical). This dampens oscillations and sustains progress along consistent gradient directions — like a ball rolling with inertia.

Second Moment v: Adaptive Scaling

An exponential moving average of squared gradients (β₂ = 0.999 typical). Dividing by √v̂ gives each parameter its own effective learning rate: parameters with historically large gradients get smaller steps, parameters with small gradients get larger steps. This is what eliminates the zigzag.

The elliptical contours matter: the loss surface has 3× steeper curvature along w₂ than w₁. SGD's single learning rate causes it to overshoot on the steep axis, creating the characteristic zigzag. Adam's per-parameter 1/√v̂ scaling automatically compensates — it takes shorter steps where gradients are large and longer steps where they're small, producing a much more direct path.

Why this matters for the rest of this tutorial: Adam's per-parameter adaptive scaling is exactly what creates the divergence between weight decay and L2 regularization. When a penalty term flows through this adaptive machinery, it gets transformed differently than when applied directly to the weights.

What is Regularization?

Neural networks are powerful function approximators — sometimes too powerful. Given enough parameters, a model can memorize the training data perfectly, including its noise and idiosyncrasies. This is overfitting : the model performs well on training data but fails to generalize to new, unseen data.

Regularization is any technique that constrains or penalizes the model to prevent overfitting and encourage generalization. The core idea: simpler models (smaller weights, smoother functions) tend to generalize better. Regularization nudges the optimizer toward these simpler solutions.

Common forms include dropout (randomly zeroing activations), data augmentation (expanding training distribution), early stopping (halting before overfitting), and — the focus of this tutorial — weight-based penalties that discourage the model from developing excessively large parameter values.

Why "L2" Regularization? — Norms and Naming

The name comes from the L2 norm (also called the Euclidean norm), which measures the "length" of a vector. For a weight vector w , the Lp norms are defined as:

General Lp norm: ‖w‖ₚ = ( |w₁|ᵖ + |w₂|ᵖ + ... + |wₙ|ᵖ ) 1/p

L1 norm (p=1): ‖w‖₁ = |w₁| + |w₂| + ... + |wₙ| — sum of absolute values (promotes sparsity)

L2 norm (p=2): ‖w‖₂ = √( w₁² + w₂² + ... + wₙ² ) — Euclidean distance from origin

L2 regularization adds the squared L2 norm of the weights to the loss function: λ · ‖w‖₂² . We use the squared norm (not the raw norm) because its gradient is clean and differentiable everywhere — it yields 2λw , a simple linear push toward zero proportional to each weight's magnitude.

L1 Regularization (Lasso)

Penalty: λ · ‖w‖₁ = λ · Σ|wᵢ|

Produces sparse solutions — many weights become exactly zero. Useful for feature selection. The gradient is ±λ (constant magnitude), which pushes small weights all the way to zero.

L2 Regularization (Ridge)

Penalty: λ · ‖w‖₂² = λ · Σwᵢ²

Produces small but non-zero weights — a smooth shrinkage toward zero. The gradient is 2λw (proportional to weight), so large weights get penalized more heavily. This is the one commonly confused with weight decay.

The Source of Confusion

Both L2 regularization and weight decay push weights toward zero. With standard SGD, they produce mathematically identical updates (as we'll see in the next sections). This led the field to treat them as synonyms for years. But they are conceptually different mechanisms — and with adaptive optimizers like Adam, they produce genuinely different behavior. Understanding why requires looking at each one carefully.

Dropout (Srivastava et al., 2014)

Dropout is a strikingly simple regularization technique: during each training forward pass, randomly set each neuron's output to zero with probability p (typically 0.1–0.5). The dropped neurons change every pass, so the network can never rely on any single neuron or co-adaptation between specific neurons.

The key insight: this is equivalent to training an exponentially large ensemble of "thinned" sub-networks that share weights. At inference, all neurons are active but outputs are scaled by (1 − p) to match the expected training-time activation magnitude. The result: better generalization through redundant, distributed representations.

Where Dropout is Applied

In transformers, dropout was historically applied in several locations: after the attention softmax (attention dropout), after each sub-layer before the residual connection, and within the feedforward MLP blocks. The original "Attention Is All You Need" paper used p = 0.1 throughout.

The Ensemble Interpretation

A network with n units can produce 2 n possible thinned sub-networks via dropout. Each training batch effectively trains a different sub-network. Inference with scaled weights approximates the geometric mean of all these sub-networks — an efficient ensemble without the cost of training separate models.

The Mechanics

Training: For each neuron, sample r ~ Bernoulli(1 − p). Output becomes: ŷ = r · y. On average, a fraction p of neurons are zeroed each pass. The surviving neurons' gradients are amplified by 1/(1 − p) to maintain expected magnitude (inverted dropout).

Inference: No dropout applied — all neurons active. With inverted dropout (the modern default), no scaling adjustment is needed at inference time since the training-time scaling already compensated.

⚖️ Weight Decay

The original idea from Hanson & Pratt (1988): every update step, multiply each weight by a factor slightly less than 1 . That's it. No loss function modification — just shrink the weights.

w ← w · (1 − λ) − lr · ∇L( w )
= w − lr·∇L(w) − λ·w ← the decay term

📐 L2 Regularization

A different approach: add a penalty term directly to the loss function that penalizes large weights. The gradient of this penalty then naturally pushes weights toward zero during optimization.

L̃ (w) = L(w) + λ·‖w‖²

∇ L̃ = ∇L(w) + 2λ·w
w ← w − lr · (∇L(w) + 2λ·w )
= w − lr·∇L(w) − lr·2λ·w ← penalty gradient

The regularization gradient 2λ·w gets mixed into the total gradient — this matters for adaptive optimizers!

🔑 The Key Difference

Weight decay acts directly on the weights — it's a multiplicative shrinkage that doesn't touch the gradient at all.
L2 regularization modifies the gradient by adding a term to the loss function — the regularization signal flows through the same gradient pipeline as the task loss.

For vanilla SGD, these produce identical update rules (just rescale λ). But for any optimizer that transforms the gradient before applying it — like Adam's per-parameter adaptive learning rates — they diverge.

With SGD, weight decay and L2 are equivalent

Weight decay: w ← w − lr·∇L(w) − lr·λ·w
L2 regularization: w ← w − lr·(∇L(w) + 2λ·w ) = w − lr·∇L(w) − lr·2λ·w

Set λ_L2 = λ_WD /2 and they produce the exact same parameter update . This algebraic equivalence led the field to treat them as synonyms for decades.

Why they match: SGD applies the raw gradient directly — there is no transformation between the gradient computation and the weight update. Whether you add the penalty to the loss (L2) or subtract it from the weights (decay), the final Δw is identical. The solid and dashed lines overlap completely.

Adam's adaptive scaling breaks the equivalence

Adam divides the gradient by √(v̂) — a running estimate of gradient variance. This is per-parameter. When L2 regularization injects the penalty into the gradient , it gets divided by √(v̂) too. The regularization strength becomes inversely proportional to gradient magnitude .

Adam + L2 (penalty in gradient)

The 2λw penalty gets scaled by 1/√(v̂). Large-gradient params: regularization gets weakened . Small-gradient params: regularization gets amplified . The decay becomes non-uniform across parameters.

Adam + true weight decay

λw applied after Adam's adaptive update — never passes through the 1/√(v̂) normalization. Every parameter decays by the same proportion , regardless of gradient scale.

Why they diverge: The elliptical contours represent different curvature per axis (common in real networks). Adam normalizes gradients per-parameter to handle exactly this. But when L2's penalty is inside the gradient, it gets normalized too — so the regularization strength becomes entangled with the loss landscape geometry. True weight decay acts on the raw weights, keeping the regularization clean and geometry-independent. The dashed bracket shows the gap between the two endpoints.

Loshchilov & Hutter (2019): "Decoupled Weight Decay Regularization"

The landmark paper that identified the problem. They showed that virtually every implementation of "weight decay" in deep learning was actually doing L2 regularization — feeding the penalty into the gradient where Adam's adaptive scaling would distort it.

The fix was remarkably simple: decouple the weight decay from the gradient-based update . Apply Adam's adaptive update first, then subtract the decay term separately.

❌ Adam (with L2 — the common mistake)

g ← ∇L(w) + 2λw
m ← β₁m + (1-β₁)g
v ← β₂v + (1-β₂)g²
w ← w − lr · m̂/√(v̂+ε)
↑ regularization signal gets distorted by adaptive scaling

✅ AdamW (decoupled — the fix)

g ← ∇L(w) (no L2 term)
m ← β₁m + (1-β₁)g
v ← β₂v + (1-β₂)g²
w ← w − lr · m̂/√(v̂+ε) − lr·λ·w
↑ decay applied after adaptive update, clean separation

Why This Matters

Loshchilov & Hutter showed that proper decoupling improved generalization across the board. AdamW is now the default optimizer in most transformer training — from BERT to GPT to most large-scale training since 2019.

The lesson: a "trivial" algebraic equivalence in one setting (SGD) became a meaningful performance difference in another (Adam). The field spent years applying L2 regularization thinking it was weight decay, and this confusion measurably held back training quality until Loshchilov and Hutter formalized the distinction.

Regularization in Frontier LLMs (2024–2025)

The regularization landscape looks very different for modern large language models compared to the deep learning of the 2010s. The scale of data and the single-epoch training regime have fundamentally changed which techniques matter.

STILL USED

Weight Decay (AdamW)

AdamW with decoupled weight decay remains the default optimizer for virtually all frontier LLM pretraining. Typical weight decay values: 0.01–0.1. One notable refinement: recent work excludes embeddings and normalization layers from weight decay, as decaying these can hurt training stability.

NOT USED IN PRETRAINING

Dropout

GPT-3, PaLM, LLaMA, Chinchilla, and Gopher all train with dropout = 0 . LLaMA's config defaults attention_dropout to 0.0. Reason: modern pretraining sees each token only once (single-epoch), so there is minimal overfitting risk. Dropout just slows convergence with no regularization benefit. Still sometimes used in fine-tuning on small datasets.

REPLACED

L2 Regularization

Since AdamW became the standard, L2 regularization (penalty in the loss) has been effectively superseded by proper decoupled weight decay. No major frontier model uses L2 regularization directly — the Loshchilov & Hutter paper resolved this definitively.

What Regularizes Modern LLMs Instead?

Data scale as implicit regularization: When training on trillions of unique tokens in a single epoch, the model never sees the same data twice. This eliminates the primary overfitting mechanism that dropout and L2 were designed to combat.

Weight decay via AdamW: The one classical regularizer that survived — and for good reason. It prevents weight magnitude explosion without distorting the adaptive gradient signal, and it helps with training stability independent of overfitting concerns.

Architectural choices: Modern regularization is increasingly built into the architecture itself — RMSNorm, QK-norm for attention stability, logit soft-capping (Gemma-style), and careful initialization schemes. These provide training stability without the generalization-vs-convergence tradeoffs of classical regularizers.

Gradient clipping: Not a regularizer in the classical sense, but gradient norm clipping (typically at 1.0) is universally applied in LLM training to prevent loss spikes and training instability.

The Exception: Fine-Tuning

While pretraining has moved past dropout and L2, fine-tuning on small datasets still benefits from classical regularization. LoRA fine-tuning commonly uses dropout (0.05–0.1) on the adapter weights, stronger weight decay (0.01–0.1), and early stopping. The data scarcity that triggers overfitting in fine-tuning is precisely the regime these techniques were designed for.