Maximum Likelihood Estimation and Maximum A Posteriori Estimation

MLE and MAP are two core methods of statistical inference.The essence of training neural networks ≈ maximum likelihood estimation.


Concept Analysis

Suppose you have a set of observed data, and you believe these data come from a known type of probability distribution (e.g., normal distribution), but you don't know the parameters of this distribution (mean μ and standard deviation σ).

MLE and MAP answer the same question: given the data, what is the most reasonable parameter value?

The difference is that MLE only looks at the data itself, while MAP also references your prior knowledge about the parameters.

Likelihood vs Probability: Two Ways to View the Same Thing

Probability P(data|parameters): parameters are fixed, asking "Given these parameters, how likely is it to observe these data?"

Likelihood L(parameters|data): data are fixed, asking "Which parameter value is most likely to produce these data?"

The value of the likelihood function itself is not normalized (the sum does not need to be 1); its significance lies incomparing the relative plausibility of different parameter values。

MLE: Which parameters are most likely to generate these data?

\[ \hat{\theta}_{\text{MLE}} = \arg\max_\theta P(D | \theta) \]

Looking only at the data, choose the parameters that most likely produce the observed data.

Why take the log?

  • Product becomes sum: numerically more stable (prevents underflow)
  • Differentiation is easier: logarithm turns multiplication into addition
  • Does not change extrema: log is a monotonically increasing function

MAP: Incorporating Prior Knowledge

\[ \hat{\theta}_{\text{MAP}} = \arg\max_\theta P(\theta | D) = \arg\max_\theta P(D | \theta) \cdot P(\theta) \]

MAP = MLE + prior P(θ). When the prior is a uniform distribution, MAP = MLE.

MLE only looks at data, MAP also looks at the prior. When data is large, the influence of the prior is diluted, and MAP approaches MLE.

When data is small, the prior acts as "regularization" — L2 regularization ≡ MAP with a Gaussian prior.


Real-life Example

Estimating average daily sales

Online store daily sales for 10 consecutive days: [23, 25, 22, 24, 26, 23, 25, 24, 22, 24].

Assume sales follow a normal distribution. The MLE estimate of the mean = the average of these numbers ≈ 23.8.

Intuitively you would do this too — MLE just turns intuition into mathematics.


Hands-on Practice with Python

Example

import numpy as np

np.random.seed(42)
data = np.random.normal(5.0, 2.0, 100)  # True μ=5, σ=2

# MLE: μ = sample mean
mu_mle = np.mean(data)
print(f"MLE μ = {mu_mle:.3f} (true=5.0)")

# MAP: prior μ ~ N(0, 1), shrinks toward the prior
prior_mu, prior_var = 0.0, 1.0
n = len(data)
mu_map = (prior_mu/prior_var + n*mu_mle/np.std(data)**2) / (1/prior_var + n/np.std(data)**2)
print(f"MAP μ = {mu_map:.3f} (compromise between MLE and prior 0)")

Running output:

MLE μ = 4.968 (真实=5.0)
MAP μ = 4.776 (在MLE和先验0之间折中)

Application Scenarios in AI

Loss Function Design = MLE

Minimizing cross-entropy loss ≈ maximum likelihood estimation under a multinomial distribution assumption. Minimizing MSE ≈ MLE under a normal error distribution assumption. The loss function is not chosen arbitrarily — each loss function corresponds to an implicit probability distribution assumption.

L2 Regularization = MAP with Gaussian Prior

Adding \( \lambda\|w\|^2 \) to the loss is equivalent to assuming the prior of weights w is \( N(0, 1/\lambda) \). MAP compromises between data likelihood and prior — when data is scarce, the prior dominates (weights tend to 0); when data is abundant, the likelihood dominates.

L1 Regularization = MAP with Laplace Prior

Adding \( \lambda\|w\|_1 \) to the loss is equivalent to assuming the prior of weights w is \( Laplace(0, 1/\lambda) \) distribution. The Laplace distribution is sharp at 0 — MAP tends to produce sparse solutions (many weights exactly 0), enabling automatic feature selection.

Bayesian Neural Networks

The weights of an ordinary neural network are deterministic point estimates (MLE/MAP). A Bayesian neural network assigns a probability distribution (posterior) to each weight, and the output is also a probability distribution — thus providing natural uncertainty estimation. During prediction, you can sample the weights multiple times and look at the variance of the outputs to judge the model's confidence in its own predictions.


Other extensions