Expectation, Variance and Covariance -- Numerical Characteristics of Data
Expectation = Where the center is. Variance = How spread out it is. Covariance = How two variables move together.
Concept Analysis
Expectation E[X]
Weighted average = "long-term average outcome"
\( \mathbb{E}[X] = \sum x_i P(x_i) \)
Variance Var(X)
Degree of data dispersion
\( \mathbb{E}[(X - \mu)^2] \)
Covariance Cov(X,Y)
Co-movement relationship between two variables
\( \mathbb{E}[(X-\mu_X)(Y-\mu_Y)] \)
Covariance Matrix
The covariance matrix of a d-dimensional random vector is a d×d symmetric matrix: \( \Sigma_{ij} = \text{Cov}(X_i, X_j) \).
- Diagonal elements = variances of individual features
- Off-diagonal elements = covariances between features
Real-life Examples
Average scores of two classes
Both Class A and Class B have an average math score of 75. But in Class A, everyone scores between 70 and 80, while in Class B, some score 30 and some score 100.
The expectations are the same, but the variances are completely different—variance reveals differences in "stability".
Python Hands-on Practice
Example
data = np.random.normal(5, 2, 10000)
print(f"Mean={np.mean(data):.3f}, Variance={np.var(data):.3f}, Std={np.std(data):.3f}")
# Covariance
x = np.random.randn(1000)
y1 = 0.8*x + np.random.randn(1000)*0.3 # Positive correlation
y2 = -0.8*x + np.random.randn(1000)*0.3 # Negative correlation
print(f"Cov(X, Y+): {np.cov(x,y1)[0,1]:.3f} (positive)")
print(f"Cov(X, Y-): {np.cov(x,y2)[0,1]:.3f} (negative)")
# Covariance matrix
data_3d = np.random.randn(100, 3)
print("\nCovariance matrix:\n", np.round(np.cov(data_3d.T), 3))
print("Symmetric:", np.allclose(np.cov(data_3d.T), np.cov(data_3d.T).T))
Output:
均值=4.990, 方差=4.008, 标准差=2.002 Cov(X, Y+): 0.846 (正) Cov(X, Y-): -0.860 (负) 协方差矩阵: [[ 1.03 0.035 0.033] [ 0.035 0.89 -0.066] [ 0.033 -0.066 0.933]] 对称: True
Application Scenarios in AI
The Mathematics of Batch Normalization
For each mini-batch, compute the mean μ_B and variance σ²_B, then normalize \( \hat{x} = (x - \mu_B) / \sqrt{\sigma^2_B + \epsilon} \). This uses two statistics: expectation (mean) and variance. \( \epsilon \) is a small constant to prevent division by zero (usually 1e-5).
PCA Dimensionality Reduction = Eigen-decomposition of Covariance Matrix
The first step of PCA: compute the covariance matrix \( \Sigma \) (a d×d symmetric matrix) of the data. Second step: perform eigendecomposition on \( \Sigma \), and take the eigenvectors corresponding to the top k largest eigenvalues as the principal component directions. The variance along these directions is maximized → retaining the most information.
Variance Control in Weight Initialization
Kaiming initialization: \( Var(W) = 2/n_{in} \) (for ReLU networks) or Xavier: \( Var(W) = 2/(n_{in}+n_{out}) \) (for tanh networks). Precisely control the variance of each layer's output so that the variance of activations remains unchanged during forward propagation and the variance of gradients remains unchanged during backpropagation. Improper initialization can cause vanishing/exploding gradients even at the early stage of training.
Second Moment Estimation in Adam Optimizer
Adam maintains an exponential moving average of the squared gradients \( v_t = \beta_2 v_{t-1} + (1-\beta_2)g_t^2 \). This is essentially an online estimate of the "variance of the gradient in each parameter direction". Directions with large variance (violent oscillation) are given a smaller learning rate, while directions with small variance (stable descent) are given a larger learning rate.
Other extensions