Special Matrices -- Identity Matrix, Diagonal Matrix, Symmetric Matrix
Some matrix types recur frequently in AI; they have special structures and properties. Understanding them helps you quickly grasp the behavior of optimizers, the effects of regularization, and the meaning of data transformations.
Identity Matrix I
All entries on the main diagonal are 1, all others are 0
Multiplying by any matrix = unchanged
Equivalent to '1' in numbers
Diagonal Matrix
Only the main diagonal has values
Each dimension is scaled independently
Equivalent to an audio equalizer
Symmetric Matrix
Transpose = itself
Symmetric about the main diagonal
All eigenvalues are real
Identity Matrix: The '1' in Mathematics
\[ \mathbf{I}_n = \begin{bmatrix} 1 & 0 & \cdots & 0 \\ 0 & 1 & \cdots & 0 \\ \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & \cdots & 1 \end{bmatrix} \]For any matrix A with compatible shape: \( \mathbf{I}A = A \), \( A\mathbf{I} = A \).
Diagonal Matrix: Independently Scaling Each Dimension
\[ \mathbf{D} = \text{diag}(d_1, d_2, \dots, d_n) = \begin{bmatrix} d_1 & 0 & \cdots & 0 \\ 0 & d_2 & \cdots & 0 \\ \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & \cdots & d_n \end{bmatrix} \]Multiplying a diagonal matrix by a vector = independently applying a different scaling factor to each component of the vector.
Symmetric Matrix: Still Itself After Transposition
\( \mathbf{A}^T = \mathbf{A} \), i.e., \( a_{ij} = a_{ji} \).
A symmetric matrix has an extremely important property:All eigenvalues are real, and the eigenvectors are mutually orthogonal.
This makes symmetric matrices extremely useful in optimization and statistics—covariance matrices and Hessian matrices are both symmetric.
Everyday Examples
Identity Matrix: The Operation That Does Nothing
You open an image in PS, make no edits, and save it directly. Input = output.
The identity matrix is the mathematical expression of this 'doing nothing'.
Diagonal matrix: audio equalizer
Each slider on the equalizer independently controls the volume of one frequency band—high frequencies +3dB, low frequencies -2dB, mid frequencies unchanged.
Each band operates independently without interference and is scaled separately. This is the behavior of a diagonal matrix.
Symmetric matrix: reflection in the mirror
Standing before a mirror, your left hand corresponds to your right hand's position in the reflection. Up and down remain unchanged; left and right are swapped.
A symmetric matrix is similar—symmetric about the main diagonal, where values remain unchanged after swapping positions across it.
Mathematical Definition
| Matrix Type | Definition Conditions | Mark |
|---|---|---|
| Identity Matrix | \( I_{ii}=1, I_{ij}=0 \; (i \neq j) \) | \( \mathbf{I}_n \) |
| Diagonal Matrix | \( D_{ij}=0 \) when \( i \neq j \) | diag(d₁,...,dₙ) |
| Symmetric Matrix | \( A_{ij} = A_{ji} \) for all i,j | \( \mathbf{A} = \mathbf{A}^T \) |
Python Hands-on Practice
Example
# Identity matrix
I = np.eye(3)
A = np.array([[4,7,2],[3,5,1],[6,8,9]])
print("I @ A == A:", np.allclose(I @ A, A))
# Diagonal matrix: independent scaling
D = np.diag([2, 3, 5])
v = np.array([1, 1, 1])
print("D @ [1,1,1] =", D @ v) # [2, 3, 5]
# Symmetric matrix
S = np.array([[1,4,5],[4,2,6],[5,6,3]])
print("S is symmetric:", np.allclose(S.T, S))
eigenvals = np.linalg.eigvals(S)
print("All eigenvalues are real:", np.all(np.isreal(eigenvals)))
# Covariance matrix is naturally symmetric
data = np.random.randn(100, 5)
cov = np.cov(data.T)
print("Covariance matrix is symmetric:", np.allclose(cov, cov.T))
I @ A == A: True D @ [1,1,1] = [2 3 5] S 对称: True 特征值全实数: True 协方差矩阵对称: True
Application Scenarios in AI
Adam Optimizer = Adaptive Scaling via Diagonal Matrix
Adam maintains an exponential moving average v_t of squared gradients for each parameter, and updates using \( \theta_{t+1} = \theta_t - \eta \cdot m_t / \sqrt{v_t} \). Mathematically, \( 1/\sqrt{v_t} \) scales each parameter component independently — this is a diagonal matrix multiplied by the gradient vector. Active gradient components have their learning rate reduced, while inactive gradient components have it amplified.
Covariance Matrix for PCA
The first step of PCA dimensionality reduction: compute the covariance matrix \( \Sigma = X^TX/(n-1) \) after centering the data. This is a symmetric matrix (\Sigma_{ij} = \Sigma_{ji}), and its eigenvalues are all real. The eigenvectors corresponding to the largest k eigenvalues are the principal component directions. The PCA class in sklearn does this under the hood.
Identity Initialization for Recurrent Neural Networks
The recurrent weight matrices of RNNs and LSTMs are often initialized as identity matrices. This ensures that early in training, hidden state information can be passed through "unchanged," preventing gradients from decaying too quickly across time steps. This technique is called IRNN (Identity RNN).
Xavier/Glorot Initialization
Xavier initialization meticulously chooses the variance of weights so that the variance of each layer's output remains consistent: \( Var(W) = 2/(n_{in} + n_{out}) \). This is equivalent to the constraint of the diagonal matrix (variance) from a normal distribution with that variance.
Other extensions