Visualizing KL Divergence -- How Much Two Distributions Differ
Adjust the parameters of two Gaussian distributions, compute and compare KL(P||Q) and KL(Q||P), and intuitively feel the asymmetry of KL divergence.
After completing this case, you will understand:KL divergence is not a 'distance' — viewing Q from P and viewing P from Q, the information loss is different.
Everyday Introduction
Using a Beijing Map to Find Your Way in New York
You have two maps: the real New York map (P) and the Beijing map (Q). Using the Beijing map to walk in New York, you need to spend extra effort to 'correct' the wrong information on the map — this extra effort is KL(P||Q). Conversely, using the New York map to walk in Beijing — the extra effort is KL(Q||P). These two kinds of extra effort are clearly different, because the two maps 'go wrong' in different ways.
Intuitive Understanding
Scenario A: P~N(0,1), Q~N(2,1) — different means but same shape, KL is nearly symmetric.
Scenario B: P~N(0,1), Q~N(0,3) — same center but different variances. KL(P||Q) = 1.10 (narrow → wide, easy), KL(Q||P) = 1.55 (wide → narrow, difficult). Clearly asymmetric — 'covering a narrow distribution with a wide one' is much easier than 'covering a wide distribution with a narrow one'.
Mathematical Definition
\[ D_{KL}(P \parallel Q) = \sum_{x} P(x) \cdot \log\frac{P(x)}{Q(x)} \] \[ D_{KL}(P \parallel Q) = H(P, Q) - H(P) \]P's own entropy is constant, so minimizing cross-entropy = minimizing KL divergence.
Python Hands-on Practice
Example
def gaussian_pdf(x, mu, sigma):
return (1 / (sigma * np.sqrt(2 * np.pi))) * \
np.exp(-(x - mu) ** 2 / (2 * sigma ** 2))
def kl_divergence(p, q, eps=1e-12):
p = np.clip(p, eps, None)
q = np.clip(q, eps, None)
return np.sum(p * np.log(p / q))
x = np.linspace(-10, 10, 2000)
dx = x[1] - x[0]
# Scenario A: different means
P_A = gaussian_pdf(x, mu=0, sigma=1) * dx
Q_A = gaussian_pdf(x, mu=2, sigma=1) * dx
# Scenario B: different variances
P_B = gaussian_pdf(x, mu=0, sigma=1) * dx
Q_B = gaussian_pdf(x, mu=0, sigma=3) * dx
print("EXAMPLE KL divergence asymmetry verification:\n")
print("Scenario A: P~N(0,1) vs Q~N(2,1)")
print(f" KL(P||Q) = {kl_divergence(P_A, Q_A):.4f}")
print(f" KL(Q||P) = {kl_divergence(Q_A, P_A):.4f}")
print(f"\nScenario B: P~N(0,1) vs Q~N(0,3)")
print(f" KL(P||Q) = {kl_divergence(P_B, Q_B):.4f} (narrow -> wide, easy)")
print(f" KL(Q||P) = {kl_divergence(Q_B, P_B):.4f} (wide -> narrow, difficult)")
print(f"\nEXAMPLE intuition: 'covering narrow with wide' is easy (small KL), 'covering wide with narrow' is difficult (large KL)")
EXAMPLE KL 散度不对称性验证: 场景 A:P~N(0,1) vs Q~N(2,1) KL(P||Q) = 2.0000 KL(Q||P) = 1.9960 场景 B:P~N(0,1) vs Q~N(0,3) KL(P||Q) = 1.0990 (窄->宽,容易) KL(Q||P) = 1.5490 (宽->窄,困难) EXAMPLE 直觉:'用宽覆盖窄'容易(KL小),'用窄覆盖宽'困难(KL大)
Application Scenarios in AI
| Scene | How to use KL divergence? |
|---|---|
| VAE | Make the encoding distribution close to standard normal — loss = reconstruction error + KL(encoding || N(0,1)) |
| Knowledge distillation | Make the student model’s predicted distribution close to the teacher model—minimize KL(teacher||student). |
| PPO reinforcement learning | Use KL divergence to constrain the difference between old and new policies, preventing updates from being too large. |