Matrix Addition and Matrix Multiplication
After understanding vector operations, level up to matrices. Matrix multiplication is the key to understanding forward propagation in neural networks — the essence of a fully connected layer is matrix multiplication.
Matrix Addition: Adding Corresponding Positions
For two matrices with the same shape, add the elements at corresponding positions — as intuitive as vector addition.
Matrix Multiplication: Dot Product of Rows and Columns
Matrix multiplication is not as simple as "multiplying corresponding positions." Why?
The deeper meaning of a matrix is a "linear transformation" — it maps one vector to another vector.
The geometric meaning of multiplying two matrices AB is: first apply transformation B, then apply transformation A. Therefore, the operation rule for AB must reflect the mathematical structure of "composite transformations," not simple element-wise multiplication.
For example: A represents "horizontal stretching by 2x," B represents "counterclockwise rotation by 90°." AB means "rotate first, then stretch" — this is completely different from BA (stretch first, then rotate), so matrix multiplication does not satisfy the commutative law.
The operation rule is:
i-th row →
j-th column ↓
(i, j)
The dot product of the i-th row of the left matrix and the j-th column of the right matrix → the value at position (i, j) of the result matrix.
Dimension Matching Rules
\( (m \times \color{#e74c3c}{n}) \;\times\; (\color{#e74c3c}{n} \times p) = (m \times p) \)
The number of columns of Amust equalthe number of rows of B. The two red-highlighted n's must be equal.
The shape of the result = (number of rows of A, number of columns of B)
Matrix multiplicationdoes not satisfy the commutative law: AB ≠ BA (in general).
The dimensions may even be mismatched, making BA undefined — this is where beginners make mistakes most easily.
The essence of matrix multiplication is the composition of linear transformations
Multiplying by matrix A → complete one transformation; then multiplying by matrix B → complete the second transformation.
Composing two transformations = one (BA) transformation. This is why matrix multiplication is defined this way.
Everyday Examples
Convenience store revenue calculation
Sales matrix for three days (rows = dates, columns = products):
| Cola | Chips | Instant noodles | |
|---|---|---|---|
| Monday | 10 | 5 | 3 |
| Tuesday | 8 | 7 | 4 |
| Wednesday | 12 | 4 | 6 |
Price matrix (rows = products, columns = stores):
| Store 1 | Store 2 | |
|---|---|---|
| Cola | 3 | 3.5 |
| Chips | 5 | 5.5 |
| Instant noodles | 4 | 4.0 |
Sales matrix (3×3) × price matrix (3×2) = revenue matrix (3×2).
Each cell of the result = total revenue from a certain day at a certain store.
Mathematical Definition
Matrix Addition
\[ (\mathbf{A} + \mathbf{B})_{ij} = a_{ij} + b_{ij} \]Prerequisite: A and B must have exactly the same shape.
Matrix Multiplication
Let A be m×n, B be n×p, then C = AB is m×p:
\[ c_{ij} = \sum_{k=1}^{n} a_{ik} b_{kj} \]Python Hands-on Practice
Examples
A_example = np.array([[10, 5, 3], [8, 7, 4], [12, 4, 6]])
price = np.array([[3, 3.5], [5, 5.5], [4, 4.0]])
# Matrix addition (same shape)
bonus = np.array([[1, 0, 1], [0, 1, 0], [1, 1, 0]])
total = A_example + bonus
print("A + bonus:\n", total)
# Matrix multiplication: (3×3) @ (3×2) = (3×2)
revenue = A_example @ price
print("\nRevenue matrix A @ price:\n", revenue)
print("Shape:", revenue.shape)
# Manually verify position (0,0)
manual = A_example[0,0]*price[0,0] + A_example[0,1]*price[1,0] + A_example[0,2]*price[2,0]
print(f"\nManual verification [0,0]: {manual} == {revenue[0,0]}")
# Trap of dimension mismatch
print(f"\nA shape: {A_example.shape}") # (3, 3)
print(f"price.T shape: {price.T.shape}") # (2, 3)
# A @ price.T → (3,3) @ (2,3) mismatch!
# Correct: price.T @ A → (2,3) @ (3,3) = (2,3)
print(f"price.T @ A shape: {(price.T @ A_example).shape}") # (2, 3)
收入矩阵 A @ price: [[67. 72. ] [75. 79.5] [80. 88. ]] 形状: (3, 2) A 形状: (3, 3) price.T 形状: (2, 3) price.T @ A 形状: (2, 3)
Interactive Matrix Multiplication Visualization
Below demonstrates 2×2 matrix A multiplied by 2×1 vector v, observe how the linear transformation maps points on a circle to an ellipse:
Application Scenarios in AI
Fully Connected Layer = Matrix Multiplication
nn.Linear(d_in, d_out)The essence: weight matrix W (d_out × d_in) multiplied by input vector x (d_in dimensions), yields output (d_out dimensions).
When processing batch data, the input X is a (batch, d_in) matrix, and we compute \( Y = XW^T \) or \( Y = WX^T \). This is a single matrix multiplication completing the forward computation for the entire batch. GPUs have specialized hardware acceleration for matrix multiplication (Tensor Core).
QK^T in the Attention Mechanism
In Transformer, Q and K are both (seq_len, d_k) matrices. \( QK^T \) is seq_len×d_k multiplied by d_k×seq_len, yielding a (seq_len, seq_len) attention score matrix.
In large models at the GPT-4 level, seq_len can reach 128K, and the computational cost of this matrix multiplication accounts for a large portion of the entire inference cost. The FlashAttention algorithm specifically optimizes the memory access pattern of this matrix multiplication.
Matrix Multiplication Implementation of Convolution (im2col)
The convolution operation can be expanded into matrix multiplication: unfold the input image into a large matrix using sliding windows (im2col), then flatten the convolution kernel into a matrix, and multiply the two. This allows convolution to leverage highly optimized GEMM (General Matrix Multiplication) libraries for acceleration.
Matrix Decomposition in LoRA Fine-tuning
LoRA adds the product of two small matrices A and B alongside the pretrained weights: \( W' = W + AB \). Here AB is matrix multiplication—A is d×r, B is r×d, and the product is a d×d low-rank matrix. A single matrix multiplication achieves parameter-efficient fine-tuning.
Other extensions