Special Matrices -- Identity Matrix, Diagonal Matrix, Symmetric Matrix

Some matrix types recur frequently in AI; they have special structures and properties. Understanding them helps you quickly grasp the behavior of optimizers, the effects of regularization, and the meaning of data transformations.

Identity Matrix I

All entries on the main diagonal are 1, all others are 0

Multiplying by any matrix = unchanged

Equivalent to '1' in numbers

Diagonal Matrix

Only the main diagonal has values

Each dimension is scaled independently

Equivalent to an audio equalizer

Symmetric Matrix

Transpose = itself

Symmetric about the main diagonal

All eigenvalues are real

Identity Matrix: The '1' in Mathematics

\[ \mathbf{I}_n = \begin{bmatrix} 1 & 0 & \cdots & 0 \\ 0 & 1 & \cdots & 0 \\ \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & \cdots & 1 \end{bmatrix} \]

For any matrix A with compatible shape: \( \mathbf{I}A = A \), \( A\mathbf{I} = A \).

Diagonal Matrix: Independently Scaling Each Dimension

\[ \mathbf{D} = \text{diag}(d_1, d_2, \dots, d_n) = \begin{bmatrix} d_1 & 0 & \cdots & 0 \\ 0 & d_2 & \cdots & 0 \\ \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & \cdots & d_n \end{bmatrix} \]

Multiplying a diagonal matrix by a vector = independently applying a different scaling factor to each component of the vector.

Symmetric Matrix: Still Itself After Transposition

\( \mathbf{A}^T = \mathbf{A} \), i.e., \( a_{ij} = a_{ji} \).

A symmetric matrix has an extremely important property:All eigenvalues are real, and the eigenvectors are mutually orthogonal.

This makes symmetric matrices extremely useful in optimization and statistics—covariance matrices and Hessian matrices are both symmetric.


Everyday Examples

Identity Matrix: The Operation That Does Nothing

You open an image in PS, make no edits, and save it directly. Input = output.

The identity matrix is the mathematical expression of this 'doing nothing'.

Diagonal matrix: audio equalizer

Each slider on the equalizer independently controls the volume of one frequency band—high frequencies +3dB, low frequencies -2dB, mid frequencies unchanged.

Each band operates independently without interference and is scaled separately. This is the behavior of a diagonal matrix.

Symmetric matrix: reflection in the mirror

Standing before a mirror, your left hand corresponds to your right hand's position in the reflection. Up and down remain unchanged; left and right are swapped.

A symmetric matrix is similar—symmetric about the main diagonal, where values remain unchanged after swapping positions across it.


Mathematical Definition

Matrix TypeDefinition ConditionsMark
Identity Matrix\( I_{ii}=1, I_{ij}=0 \; (i \neq j) \)\( \mathbf{I}_n \)
Diagonal Matrix\( D_{ij}=0 \) when \( i \neq j \)diag(d₁,...,dₙ)
Symmetric Matrix\( A_{ij} = A_{ji} \) for all i,j\( \mathbf{A} = \mathbf{A}^T \)

Python Hands-on Practice

Example

import numpy as np

# Identity matrix
I = np.eye(3)
A = np.array([[4,7,2],[3,5,1],[6,8,9]])
print("I @ A == A:", np.allclose(I @ A, A))

# Diagonal matrix: independent scaling
D = np.diag([2, 3, 5])
v = np.array([1, 1, 1])
print("D @ [1,1,1] =", D @ v)  # [2, 3, 5]

# Symmetric matrix
S = np.array([[1,4,5],[4,2,6],[5,6,3]])
print("S is symmetric:", np.allclose(S.T, S))
eigenvals = np.linalg.eigvals(S)
print("All eigenvalues are real:", np.all(np.isreal(eigenvals)))

# Covariance matrix is naturally symmetric
data = np.random.randn(100, 5)
cov = np.cov(data.T)
print("Covariance matrix is symmetric:", np.allclose(cov, cov.T))
I @ A == A: True
D @ [1,1,1] = [2 3 5]
S 对称: True
特征值全实数: True
协方差矩阵对称: True

Application Scenarios in AI

Adam Optimizer = Adaptive Scaling via Diagonal Matrix

Adam maintains an exponential moving average v_t of squared gradients for each parameter, and updates using \( \theta_{t+1} = \theta_t - \eta \cdot m_t / \sqrt{v_t} \). Mathematically, \( 1/\sqrt{v_t} \) scales each parameter component independently — this is a diagonal matrix multiplied by the gradient vector. Active gradient components have their learning rate reduced, while inactive gradient components have it amplified.

Covariance Matrix for PCA

The first step of PCA dimensionality reduction: compute the covariance matrix \( \Sigma = X^TX/(n-1) \) after centering the data. This is a symmetric matrix (\Sigma_{ij} = \Sigma_{ji}), and its eigenvalues are all real. The eigenvectors corresponding to the largest k eigenvalues are the principal component directions. The PCA class in sklearn does this under the hood.

Identity Initialization for Recurrent Neural Networks

The recurrent weight matrices of RNNs and LSTMs are often initialized as identity matrices. This ensures that early in training, hidden state information can be passed through "unchanged," preventing gradients from decaying too quickly across time steps. This technique is called IRNN (Identity RNN).

Xavier/Glorot Initialization

Xavier initialization meticulously chooses the variance of weights so that the variance of each layer's output remains consistent: \( Var(W) = 2/(n_{in} + n_{out}) \). This is equivalent to the constraint of the diagonal matrix (variance) from a normal distribution with that variance.


Other extensions