Vector Spaces and Bases -- Seeing the World Through a Different Coordinate System

The same location has different longitude/latitude values in the GPS coordinate system and the BeiDou coordinate system — but the location itself hasn't changed.

The same applies to vectors:The same vector has different coordinates under different bases.Understanding this, you will understand the essence of PCA dimensionality reduction and word embedding spaces.

Vector Spaces

A vector space is a set whose elements (vectors) are closed under addition and scalar multiplication.

Closure = adding any two vectors in the space, or multiplying any vector by a scalar, still results in a vector within the space.

Basis: The Coordinate Ruler for Describing the Space

A basis is a set of vectors whose linear combinations can uniquely represent any vector in the space.

Standard basis: (1,0,0), (0,1,0), (0,0,1) — corresponding to the x, y, and z axis directions, respectively.

Linear Independence: No Redundancy

A set of vectors is linearly independent = none of the vectors can be represented by a linear combination of the other vectors.

Intuition: Each vector brings a genuinely 'new direction' with no waste.

Linearly independent

{(1,0), (0,1)}

Each vector provides a new direction

Rank = 2 (full rank)

Linearly dependent

{(1,2), (2,4)}

The second = the first × 2

Rank = 1 (redundant)

A basis must satisfy two conditions: linear independence (no redundancy) + spanning the whole space (can represent all vectors).

A basis of an n-dimensional space contains exactly n vectors.


Real-Life Examples

Changing coordinate systems

The location of Beijing is (x₁, y₁) in the Beijing 54 coordinate system and (x₂, y₂) in the WGS-84 coordinate system.

The location itself hasn't changed, but the 'ruler' (basis) used to describe it differs, so the coordinate numbers differ.

Word Embedding Space

The word embedding model maps each word to a vector in a 300-dimensional space. Each dimension of this space corresponds to a latent semantic direction.

"King - Man + Woman ≈ Queen" — these vector operations all take place in this 300-dimensional vector space.


Mathematical Definition

\( \{\mathbf{v}_1, \dots, \mathbf{v}_n\} \) is a basis = they are linearly independent + span{\(\mathbf{v}_1,\dots,\mathbf{v}_n\)} = the entire space.

Linear combination: \( c_1\mathbf{v}_1 + \cdots + c_n\mathbf{v}_n = \mathbf{0} \) has only the zero solution → linearly independent.


Hands-On Python Practice

Example

import numpy as np

# Standard basis
e1, e2, e3 = np.array([1,0,0]), np.array([0,1,0]), np.array([0,0,1])
v = np.array([4, 5, 6])
print("v = 4*e1 + 5*e2 + 6*e3:", 4*e1 + 5*e2 + 6*e3)

# Determine linear independence using rank
v1, v2 = np.array([1,2]), np.array([2,4])  # v2 = 2*v1
M = np.column_stack([v1, v2])
print(f"Rank of {{v1, v2}}: {np.linalg.matrix_rank(M)} (< 2 → linearly dependent)")

# Change of basis: the same vector has different coordinates under different bases
v_std = np.array([3.0, 2.0])
B_new = np.array([[2.0, 0], [0, 3.0]])  # New basis
v_new = np.linalg.inv(B_new) @ v_std
print(f"Standard coordinates: {v_std}, coordinates under the new basis: {v_new}")
print(f"Verify B @ new coordinates = {B_new @ v_new}")
v = 4*e1 + 5*e2 + 6*e3: [4 5 6]
{v1, v2} 的秩: 1 (< 2 → 线性相关)
标准坐标: [3. 2.], 新基下坐标: [1.5 0.67]
验证 B @ 新坐标 = [3. 2.]

Application Scenarios in AI

PCA Dimensionality Reduction = Changing the Basis

PCA finds a new set of orthonormal basis vectors (principal component directions) that maximize the variance of the data along the first k coordinates under the new basis. Dimensionality reduction = keeping only the first k new-basis coordinates and discarding the directions with small variance afterward. This is essentially a coordinate system transformation — changing from a 784-dimensional pixel space to a 50-dimensional principal component space.

Word Embedding Space

Word2Vec and GloVe map each word to a point in a 300-dimensional vector space. This space is not arbitrarily constructed — the geometric relationships between words reflect semantic relationships. Synonyms are close in distance, while antonyms are exactly opposite in certain directions. Linear operations (addition and subtraction) in the vector space correspond to semantic operations.

Feature Space Transformation

Each layer of a neural network can be seen as mapping the input to a new vector space. Data that is linearly inseparable in the original space (such as the XOR problem) becomes linearly separable after being transformed by hidden layers into the feature space. This is the core intuition of deep learning — layer-by-layer transformation, finally separable in the feature space with a simple classifier.


Other Extensions