Matrix Transpose and Inverse Matrix

Transpose appears frequently in attention mechanisms, and inverse matrices help us understand "invertible transformations" — such as coordinate transformations in PCA.

Transpose: swap rows and columns

A (2×3)
123
456
AT (3×2)
14
25
36

The value originally at position (i, j) goes to position (j, i) after transposition.The shape changes from m×n to n×m.

Inverse matrix: undo the transformation

Matrix A represents a linear transformation. \( A^{-1} \) represents "undoing the effect of A".

Mathematical expression: \( AA^{-1} = A^{-1}A = I \) (I is the identity matrix, equivalent to the number 1).

Onlysquare and full-rankmatrices have an inverse matrix.

Non-square matrices or square matrices that are not full rank are "singular matrices" and have no inverse — equivalent to an irreversible operation.

Key properties

  • Transpose of a product reverses order: \( (\mathbf{AB})^T = \mathbf{B}^T\mathbf{A}^T \)
  • Inverse of a product also reverses order: \( (\mathbf{AB})^{-1} = \mathbf{B}^{-1}\mathbf{A}^{-1} \)

Real-life examples

Transpose: swapping rows and columns in an Excel spreadsheet

If you transpose a sales table with "names in rows, months in columns", it becomes "names in columns, months in rows".

The data hasn't changed, only the arrangement has — that's transposition.

Inverse matrix: Ctrl+Z undo

You perform a rotation + scaling operation (matrix A) and want to restore it (A^{-1}).

The composition of the two operations = nothing has changed (AA^{-1} = I, identity transformation).


Mathematical definitions

Transpose

\[ (\mathbf{A}^T)_{ij} = \mathbf{A}_{ji} \]

Inverse matrix

For an n×n square matrix A, if there exists B such that \( \mathbf{AB} = \mathbf{BA} = \mathbf{I}_n \), then B = A^{-1}.


Python Hands-on Practice

Example

import numpy as np

A = np.array([[1, 2, 3], [4, 5, 6]])
print("A (2x3):\n", A)
print("A.T (3x2):\n", A.T)
print()

# Verify (AB)^T = B^T @ A^T
B = np.array([[1,2],[3,4],[5,6]])  # 3x2
AB = A @ B   # (2,3) @ (3,2) = (2,2)
print("(AB)^T:\n", AB.T)
print("B^T @ A^T:\n", B.T @ A.T)
print("Equal:", np.allclose(AB.T, B.T @ A.T))
print()

# Inverse matrix
C = np.array([[2, 1], [5, 3]])  # 2x2 square matrix
C_inv = np.linalg.inv(C)
print("C:\n", C)
print("C^{-1}:\n", C_inv)
print("C @ C^{-1} = I:\n", C @ C_inv)

# Singular matrix has no inverse
D = np.array([[1, 2], [2, 4]])  # Row 2 is twice row 1
try:
    np.linalg.inv(D)
except np.linalg.LinAlgError:
    print("\nSingular matrix D has no inverse (row 2 = 2 × row 1))
A (2x3):
 [[1 2 3]
 [4 5 6]]
A.T (3x2):
 [[1 4]
 [2 5]
 [3 6]]

C @ C^{-1} = I:
 [[1. 0.]
 [0. 1.]]

奇异矩阵 D 没有逆矩阵(第2行=2×第1行)

Application scenarios in AI

K^T operation in attention mechanism

In \( QK^T \), the transpose of the Key matrix allows Q of shape (seq_len, d_k) to multiply with K of shape (seq_len, d_k) via matrix multiplication—the transpose turns K into (d_k, seq_len), making the inner dimensions match. This K^T is the most critical dimension transformation in the attention computation of all Transformer models.

Closed-form solution of linear regression

\( \hat{\beta} = (X^TX)^{-1}X^Ty \)—transpose and inverse appear together. X^T X guarantees the result is a symmetric positive definite matrix (invertible), then inverting it gives the closed-form solution. sklearn's LinearRegression computes this under the hood when no regularization is added.

Parameter updates in gradient descent

In backpropagation, the shape of the weight gradient matrix must match the weight matrix to perform the subtraction update. For example, W is (d_out, d_in), and its gradient \nabla_W L is also (d_out, d_in)—transpose helps gradients propagate dimensions correctly through the computational graph in intermediate layers.


Other Extensions