Continuous Random Variables and Common Distributions

A continuous random variable can take any value on the real number line.The normal (Gaussian) distribution is the most important distribution in AI, bar none.


Concept Explanation

PDF: Probability Density Function

The PDF is not probability itself; the area is the probability.The PDF value f(x) at a particular point has no meaning by itself — the probability that a continuous variable equals a specific value is 0. But the integral of the PDF over an interval = the probability that X falls within that interval.

Normal Distribution N(μ, σ²)

\[ f(x) = \frac{1}{\sigma\sqrt{2\pi}} \exp\left(-\frac{(x-\mu)^2}{2\sigma^2}\right) \]

Determined by the mean μ (location) and standard deviation σ (width). About 68% of data falls within μ±σ, and about 95% within μ±2σ.

Central Limit Theorem

The superposition of many independent factors → the result approaches a normal distribution. This explains why the normal distribution is everywhere in nature and AI.

Neural network weight initialization, Batch Normalization, VAE latent space, and noising/denoising in diffusion models — all rely on the normal distribution.


Real-life Examples

Human height: 170.1, 170.12, 170.128 cm — height can take any value within a continuous range and overall follows a normal distribution.

Most people's heights are close to the average; extremely tall and extremely short people are a minority — this is the bell curve.


Python Hands-on Practice

Example

import numpy as np

samples = np.random.normal(0, 1, 10000)
print(f"N(0,1) sampling: mean={samples.mean():.3f}, std={samples.std():.3f}")

# 68-95-99.7 rule
for k in [1, 2, 3]:
    within = np.mean(np.abs(samples) < k)
    theory = {1:0.6827, 2:0.9545, 3:0.9973}[k]
    print(f"μ±{k}σ: {within:.4f} (theoretical {theory:.4f})")

# Central limit theorem: sum of 12 dice ≈ normal
dice = np.random.randint(1,7,(100000,12)).sum(axis=1)
print(f"\nSum of 12 dice: mean={dice.mean():.1f} (theoretical 42), close to normal)

Run output:

N(0,1) 采样: 均值=-0.002, 标准差=1.001
μ±1σ: 0.6827 (理论 0.6827)
μ±2σ: 0.9545 (理论 0.9545)
μ±3σ: 0.9973 (理论 0.9973)

12个骰子和: 均值=42.0(理论42), 接近正态

Normal Distribution Histogram


Application Scenarios in AI

Weight Initialization

Neural network weights are usually randomly initialized from a normal distribution or uniform distribution. Kaiming initialization (for ReLU networks): \( W \sim N(0, \sqrt{2/n_{in}}) \); Xavier initialization: \( W \sim U(-\sqrt{6/(n_{in}+n_{out})}, \sqrt{6/(n_{in}+n_{out})}) \). The carefully chosen variance ensures signals propagate stably in forward and backward pass, neither exploding nor vanishing.

Batch Normalization

BN assumes that the activations of each layer approximate a normal distribution. It subtracts the mean and divides by the standard deviation to standardize them to N(0,1), then scales and shifts via learnable \(\gamma\) and \(\beta\). This accelerates training, allows larger learning rates, and reduces sensitivity to initialization. BN is one of the most important training techniques in the history of deep learning.

VAE Latent Space Assumption

The variational autoencoder assumes that the latent variable z follows a standard normal distribution N(0, I). The encoder outputs μ and log(σ²), and the decoder samples z from N(μ, σ²) to reconstruct the input. The KL divergence loss constrains N(μ, σ²) to be close to the prior N(0, I), preventing the encoder from collapsing to a Dirac distribution.

Noising and Denoising in Diffusion Models

The forward process of diffusion models (DDPM) gradually adds Gaussian noise \epsilon ∼ N(0, I) to the data until it becomes pure noise. The reverse process trains a neural network to predict the noise added at each step, and gradually denoises from pure noise to generate images. Stable Diffusion and DALL·E are both based on this principle.


Other extensions