R scale() Function - Data Standardization
The R scale() function is used to standardize data (z-score standardization).
After standardization, the mean of the data is 0 and the standard deviation is 1, often used to eliminate the influence of different dimensions (units).
The syntax of the scale() function is as follows:
scale(x, center = TRUE, scale = TRUE)
Parameter description:
xInput numeric matrix or data frame.
centerWhether to perform centering (subtract the mean), default is TRUE.
scaleWhether to perform scaling (divide by the standard deviation), default is TRUE.
Example
# Original data: variables with different dimensions
height <- c(160, 170, 165, 175, 180) # Range 160-180
income <- c(50000, 60000, 55000, 80000, 90000) # Range 50000-90000
age <- c(25, 30, 28, 35, 40) # Range 25-40
# Combine into a matrix
data_matrix <- cbind(height, income, age)
print("Original data:")
print(data_matrix)
# Standardize
scaled_data <- scale(data_matrix)
print("After standardization (mean 0, standard deviation 1):")
print(round(scaled_data, 3))
# Verify the mean and standard deviation
print("Column means after standardization (should be close to 0):")
print(round(colMeans(scaled_data), 3))
print("Column standard deviations after standardization (should equal 1):")
print(round(apply(scaled_data, 2, sd), 3))
height <- c(160, 170, 165, 175, 180) # Range 160-180
income <- c(50000, 60000, 55000, 80000, 90000) # Range 50000-90000
age <- c(25, 30, 28, 35, 40) # Range 25-40
# Combine into a matrix
data_matrix <- cbind(height, income, age)
print("Original data:")
print(data_matrix)
# Standardize
scaled_data <- scale(data_matrix)
print("After standardization (mean 0, standard deviation 1):")
print(round(scaled_data, 3))
# Verify the mean and standard deviation
print("Column means after standardization (should be close to 0):")
print(round(colMeans(scaled_data), 3))
print("Column standard deviations after standardization (should equal 1):")
print(round(apply(scaled_data, 2, sd), 3))
Executing the above code produces the following output:
[1] "原始数据:"
height income age
[1,] 160 50000 25
[2,] 170 60000 30
[3,] 165 55000 28
[4,] 175 80000 35
[5,] 180 90000 40
[1] "标准化后(均值为0,标准差为1):"
height income age
[1,] -1.2649 -1.2647 -1.2649
[2,] 0.0000 -0.4739 0.0000
[3,] -0.6325 -0.8699 -0.6325
[4,] 0.6325 0.7109 0.6325
[5,] 1.2649 1.8976 1.2649
[1] "标准化后各列均值(应接近0):"
height income age
0 0 0
[1] "标准化后各列标准差(应等于1):"
height income age
1 1 1
Other extensions
R Language Examples