conceptStatistical Machine Learning~1 min readUpdated 2026-06-07#machine-learning#regularization#overfitting#l1#l2

Regularization: L1, L2 & how they differ

Regularization is the main dial for the bias–variance tradeoff: add a penalty for complexity so the model can't fit noise. You trade a little training accuracy for better generalization.

The idea

Instead of minimizing just the loss, minimize loss + λ × penalty(weights). Large weights mean a more flexible, wigglier function; penalizing them keeps the model simpler. λ (lambda) controls the strength — a key hyperparameter you tune by cross-validation.

L1 vs L2

Penalty Effect on weights Use when
L2 (Ridge) sum of squares shrinks all weights toward zero, smoothly default; correlated features
L1 (Lasso) sum of absolute values drives some weights to exactly zero you want automatic feature selection / sparsity
Elastic Net mix of both shrinks and selects many correlated features

The key intuition: L1 produces sparse models (built-in feature selection), because its penalty geometry has corners that push weights to zero. L2 keeps all features but small, which is more stable when features are correlated.

Same idea, other names

Regularization is everywhere; the form changes:

  • Early stopping — stop training when validation loss rises (limits effective capacity).
  • Dropout — randomly zero activations during training (a deep learning regularizer).
  • Weight decay — L2 by another name, baked into optimizers like AdamW.
  • More data / augmentation — the strongest "regularizer" of all.

Pitfall

Scale your features before L1/L2 — the penalty treats all weights equally, so an unscaled large-range feature gets unfairly penalized (or spared). And too-strong λ underfits: the symptom is high error on both train and validation.

Connects to: bias–variance · tuning λ · feature selection