Visualization of the Central Limit Theorem

Repeatedly sample from a uniform distribution that is completely non-normal, calculate the mean each time, and observe how the distribution of means approaches a bell shape.

After completing this case, you will understand:Why the Gaussian distribution is ubiquitous in AI — it is the inevitable result of the superposition of many independent factors.


Life Introduction

Why is the class average height always bell-shaped?

Measure the heights of 50 people in the class. An individual's height is influenced by hundreds of factors: genetics, nutrition, exercise, sleep... Each factor has a large or small impact, positive or negative.

When many independent small factors are stacked together, the final distribution is bell-shaped — this is a mathematical theorem, not a coincidence.The central limit theorem states: the sum (or mean) of a large number of independent random variables tends toward a normal distribution.


Intuitive Understanding

We sample from a uniform distribution U(0,1) — this distribution is completely flat and has nothing to do with a bell shape. But when we take multiple samples (e.g., 30) each time to compute one mean, and repeat this operation 5,000 times — the histogram of these 5,000 means begins to look bell-shaped. The larger n is, the closer it approaches a perfect normal distribution.


Mathematical Definition

\[ \frac{\bar{X}_n - \mu}{\sigma / \sqrt{n}} \xrightarrow{d} \mathcal{N}(0, 1) \]

Where \(\bar{X}_n = \frac{1}{n}\sum X_i\). The standard deviation of the mean shrinks at a rate of \(1/\sqrt{n}\).


Python Hands-on Practice

Example

import numpy as np
np.random.seed(2)

def sample_raw(size):
    return np.random.uniform(0, 1, size=size)

def clt_experiment(n, repeat=5000):
    means = [sample_raw(n).mean() for _ in range(repeat)]
    return np.array(means)

sample_sizes = [1, 2, 5, 30]
results = {n: clt_experiment(n) for n in sample_sizes}

print("EXAMPLE Central Limit Theorem Verification:\n")
print("Original distribution: Uniform U(0,1) (completely flat)")
print("Theory: standard deviation of mean = 1/sqrt(12n)\n")
print(f"{'n':<6} {'Mean SD':<12} {'Theory':<12} {'Match'}")
print("-" * 42)
for n, means in results.items():
    theo = 1 / np.sqrt(12 * n)
    ok = "Yes" if abs(means.std() - theo) < 0.01 else "No"
    print(f"{n:<6} {means.std():<12.4f} {theo:<12.4f} {ok}")

# Normality test for n=30
m30 = results[30]
skew = np.mean(((m30-m30.mean())/m30.std())**3)
kurt = np.mean(((m30-m30.mean())/m30.std())**4) - 3
print(f"\nEXAMPLE n=30: skewness={skew:.3f} (should be ~0), excess kurtosis={kurt:.3f} (should be ~0)")
EXAMPLE 中心极限定理验证:

原始分布: 均匀分布 U(0,1)(完全平的)
理论: 均值的标准差 = 1/sqrt(12n)

n      均值SD       理论值       匹配
------------------------------------------
1      0.2891       0.2887       Yes
2      0.2043       0.2041       Yes
5      0.1292       0.1291       Yes
30     0.0526       0.0527       Yes

EXAMPLE n=30 时:偏度=-0.023 (应~0), 超峰度=-0.005 (应~0)

Application Scenarios in AI

ScenarioConnection to CLT
Batch NormalizationAssume that the activation values within each mini-batch are approximately normally distributed, then subtract the mean and divide by the standard deviation.
Weight initializationXavier/He initialization samples initial weights from a normal distribution
Error modelingIn regression problems, it is assumed that errors follow a normal distribution—errors are the result of the superposition of multiple unmodeled factors.
Other extensions