Data Shapes -- Scalars, Vectors, Matrices, Tensors
In the world of AI, all data — whether an image, a piece of text, or an audio clip — is ultimately organized into a unified form:Tensor。
Tensor is a general term; depending on the number of dimensions, it has different names.
From Zero-Dimensional to Multi-Dimensional: Layered Nested Data Structures
Another way to understand it is 'nesting' — a scalar is at the innermost level; wrapping it one level creates a vector, wrapping it again creates a matrix, and continuing to wrap creates higher-dimensional tensors.
In deep learning frameworks such as PyTorch and TensorFlow, all data is uniformly represented as 'tensors'.
A scalar is a 0-dimensional tensor, a vector is a 1-dimensional tensor, and a matrix is a 2-dimensional tensor—they are special cases of tensors.
Understanding Data Shapes with Dimensions
It is just a number. For example, temperature 26°C. In NumPy, the shape is displayed as an empty tuple.()。
An ordered sequence of numbers. For example, temperature records for each hour of a day. The shape is(n,), where n is the number of elements.
A table with rows and columns. For example, a temperature table for 7 days a week and 24 hours a day. The shape is(number of rows, number of columns)。
Multiple layers of data stacked together. For example, temperatures for 300 cities nationwide, each city over 7 days a week and 24 hours a day. The shape is(number of cities, number of days, number of hours)i.e.(300, 7, 24)。
Dimensions can be simply understood as:How many "coordinates" you need to pinpoint a specific number.
Real-life Examples
Below, we use a progressive scenario to consolidate understanding.
Scenario 1: Recording the temperature right now
You only need to write down one number:26. This is ascalar. It has no dimensions; it is just an isolated value.
Scenario 2: Recording temperature changes over a day
You measure once every hour from morning to night, getting 24 numbers:[22, 23, 24, 26, 28, ...]。
This is avector. You need one coordinate (which hour) to locate a certain temperature value.
Scenario 3: Recording temperatures for a week
You arrange the seven days of data into a table: each row is a day, each column is an hour.
This is a 7-row, 24-columnmatrix. You need two coordinates (which day, which hour) to locate.
Scenario 4: Recording temperatures across cities nationwide
There are 300 cities nationwide, and each city has a 7×24 table.
Stacking these 300 tables together forms a3D tensor. You need three coordinates (which city, which day, which hour).
Mathematical Definitions
Scalar
A scalar is a single number, denoted as \( a = 26 \), or \( x \in \mathbb{R} \) (x belongs to the set of real numbers).
Vector
A vector is an ordered set of numbers, usually written vertically (column vector):
\[ \mathbf{v} = \begin{bmatrix} v_1 \\ v_2 \\ \vdots \\ v_n \end{bmatrix} \in \mathbb{R}^{n} \]\( \mathbb{R}^{n} \) means this is an n-dimensional real vector, \( v_1 \) is the 1st component, and \( v_n \) is the nth component.
Matrix
A matrix is a rectangular array with m rows and n columns:
\[ \mathbf{A} = \begin{bmatrix} a_{11} & a_{12} & \cdots & a_{1n} \\ a_{21} & a_{22} & \cdots & a_{2n} \\ \vdots & \vdots & \ddots & \vdots \\ a_{m1} & a_{m2} & \cdots & a_{mn} \end{bmatrix} \in \mathbb{R}^{m \times n} \]It is denoted as an m×n matrix (m rows, n columns), where \( a_{ij} \) represents the element in the i-th row and j-th column.
Tensor
A tensor is a generalization of vectors and matrices to arbitrary dimensions.
The shape of a k-dimensional tensor is denoted as \( (d_1, d_2, \dots, d_k) \), with a total of \( d_1 \times d_2 \times \cdots \times d_k \) elements.
Comparison Quick Reference Table
| Name | Dimensions | Shape example | Number of coordinates | Mathematical notation |
|---|---|---|---|---|
| Scalar | 0 | () | 0 | \( x \in \mathbb{R} \) |
| Vector | 1 | (n,) | 1 | \( \mathbf{v} \in \mathbb{R}^{n} \) |
| Matrix | 2 | (m, n) | 2 | \( \mathbf{A} \in \mathbb{R}^{m \times n} \) |
| 3D tensor | 3 | (d1, d2, d3) | 3 | \( \mathcal{T} \in \mathbb{R}^{d_1 \times d_2 \times d_3} \) |
Hands-on Python Practice
Next, we use NumPy to create and inspect these data structures.shape attributeis the key to understanding data dimensions.
Examples
# Scalar (0-dimensional tensor): shape = ()
scalar = np.array(26)
# Vector (1-dimensional tensor): shape = (n,)
vector = np.array([22, 23, 24, 26, 28, 30, 31, 29])
# Matrix (2-dimensional tensor): shape = (rows, columns)
matrix = np.array([[1, 2, 3], [4, 5, 6]])
# 3D tensor: shape = (layers, rows, columns)
tensor_3d = np.array([[[1,2,3],[4,5,6]], [[7,8,9],[10,11,12]]])
# 4D tensor: simulating 3 single-channel images with 2x2 pixels
batch_images = np.array([
[[[10],[20]],[[30],[40]]],
[[[50],[60]],[[70],[80]]],
[[[90],[100]],[[110],[120]]]
])
print("Scalar: shape=", scalar.shape, " ndim=", scalar.ndim)
print("Vector: shape=", vector.shape, " ndim=", vector.ndim)
print("Matrix: shape=", matrix.shape, " ndim=", matrix.ndim)
print("3D tensor: shape=", tensor_3d.shape, " ndim=", tensor_3d.ndim)
print("4D tensor: shape=", batch_images.shape, " ndim=", batch_images.ndim)
Output:
标量: shape= () ndim= 0 向量: shape= (8,) ndim= 1 矩阵: shape= (2, 3) ndim= 2 3D 张量: shape= (2, 2, 3) ndim= 3 4D 张量: shape= (3, 2, 2, 1) ndim= 4
The Pattern of the Shape Attribute
| Attribute | Meaning | Scalar | Vector (8 elements) | Matrix (2×3) | 3D(2×2×3) |
|---|---|---|---|---|---|
.shape | Size of each dimension | () | (8,) | (2, 3) | (2, 2, 3) |
.ndim | Number of dimensions = len(shape) | 0 | 1 | 2 | 3 |
.size | Total number of elements | 1 | 8 | 6 | 12 |
When debugging deep learning code, if you see a shape mismatch error, the first step is to print the shape to check the data dimensions.
Understanding shape is the first skill for learning NumPy and all deep learning frameworks.
3D Visualization of Tensors
Below, the 3×3×3 tensor is drawn as 27 points in space: the three coordinate axes correspond to "row, column, channel", and the color intensity of each point represents the value at that position. Drag the mouse to rotate the view.
Pay attention: keeping the "channel" layer fixed, you will see a 3×3 plane—that is exactly a second-order tensor (matrix). This is also why we say "a matrix is a special case of a tensor."
Application Scenarios in AI
Image Classification: From 3D to 4D Tensors
In the ImageNet classification task, input images are uniformly resized to 224×224 pixels. A single color image has shape(224, 224, 3)The 3D tensor.
During actual training, the GPU processes one batch at a time (e.g., 32 images), and the data becomes a 4D tensor.(32, 224, 224, 3)。
Note that different frameworks have different channel position conventions: TensorFlow/Keras use(N, H, W, C)(Channels last), PyTorch uses(N, C, H, W)(The passage is at the very front).
When you see the error "shape mismatch: (32, 224, 224, 3) vs (32, 3, 224, 224)", it is caused by inconsistent channel positions—use.permute()or.transpose()Just adjust it.
Natural Language Processing: Embedding Matrices
The vocabulary size of the BERT model is 30522, and the hidden dimension is 768. Its word embedding matrix shape is(30522, 768)—Each row is a 768-dimensional vector for a word.
The same is true for the GPT series: GPT-2's vocabulary is 50,257, and the embedding dimension ranges from 768 (small model) to 1600 (large model).
When you input a piece of text, the model first looks up the table to convert each word ID into the corresponding embedding vector; this is the "table lookup to extract rows" — a matrix indexing operation.
Fully Connected Layers: The Core Position of Matrix Multiplication
A classic MNIST handwritten digit recognition network: input 784 dimensions (28×28 pixels flattened), first hidden layer 256 dimensions.
The shape of the weight matrix W₁ is(256, 784)Each forward propagation computation.h = W₁ @ x, that is, a 256×784 matrix multiplied by a 784-dimensional vector.
The forward propagation of the entire neural network is essentially an alternation of layers of matrix multiplication and activation functions.
Video Data: 5D Tensors
A video can be understood as a sequence of multiple image frames. A 16-frame 224×224 color video clip has shape(16, 224, 224, 3)。
After adding the batch dimension, it becomes a 5D tensor.(8, 16, 224, 224, 3)——This is the standard input for video classification models (such as 3D-ResNet).
Other extensions