NumPy Statistical Functions

NumPy provides many statistical functions for finding the minimum element, maximum element, percentile, standard deviation, and variance from an array.

numpy.amin() and numpy.amax()

numpy.amin() is used to compute the minimum value of elements in an array along a specified axis.

numpy.amin(a, axis=None, out=None, keepdims=<no value>, initial=<no value>, where=<no value>)

Parameter description:

  • a: The input array, which can be a NumPy array or an array-like object.
  • axis: Optional parameter, used to specify on which axis to compute the minimum. If not provided, the minimum of the entire array is returned. It can be an integer representing the axis index, or a tuple representing multiple axes.
  • out: Optional parameter, used to specify the storage location for the result.
  • keepdims: Optional parameter. If True, the number of dimensions of the result array is kept the same as the input array. If False (default), axes with dimension 1 after computation are removed.
  • initial: Optional parameter, used to specify an initial value, and then the minimum is computed over the elements of the array.
  • where: Optional parameter, a boolean array used to specify that only elements satisfying the condition are considered.

numpy.amax() is used to compute the maximum value of elements in an array along a specified axis.

numpy.amax(a, axis=None, out=None, keepdims=<no value>, initial=<no value>, where=<no value>)

Parameter description:

  • a: The input array, which can be a NumPy array or an array-like object.
  • axis: Optional parameter, used to specify on which axis to compute the maximum. If not provided, the maximum of the entire array is returned. It can be an integer representing the axis index, or a tuple representing multiple axes.
  • out: Optional parameter, used to specify the storage location for the result.
  • keepdims: Optional parameter. If True, the number of dimensions of the result array is kept the same as the input array. If False (default), axes with dimension 1 after computation are removed.
  • initial: Optional parameter, used to specify an initial value, and then the maximum is computed over the elements of the array.
  • where: Optional parameter, a boolean array used to specify that only elements satisfying the condition are considered.

Example

import numpy as np a = np.array([[3,7,5],[8,4,3],[2,4,9]]) print ('Our array is:') print (a) print ('\n') print ('Calling the amin() function:') print (np.amin(a,1)) print ('\n') print ('Calling the amin() function again:') print (np.amin(a,0)) print ('\n') print ('Calling the amax() function:') print (np.amax(a)) print ('\n') print ('Calling the amax() function again:') print (np.amax(a, axis = 0))

The output result is:

我们的数组是:
[[3 7 5]
 [8 4 3]
 [2 4 9]]


调用 amin() 函数:
[3 3 2]


再次调用 amin() 函数:
[2 4 3]


调用 amax() 函数:
9


再次调用 amax() 函数:
[8 7 9]

numpy.ptp()

numpy.ptp()The function computes the difference between the maximum and minimum values of elements in the array (maximum - minimum).

numpy.ptp(a, axis=None, out=None, keepdims=<no value>, initial=<no value>, where=<no value>)

Parameter description:

  • a: The input array, which can be a NumPy array or an array-like object.
  • axis: Optional parameter, used to specify on which axis to compute the peak-to-peak value. If not provided, the peak-to-peak value of the entire array is returned. It can be an integer representing the axis index, or a tuple representing multiple axes.
  • out: Optional parameter, used to specify the storage location for the result.
  • keepdims: Optional parameter. If True, the number of dimensions of the result array is kept the same as the input array. If False (default), axes with dimension 1 after computation are removed.
  • initial: Optional parameter, used to specify an initial value, and then the peak-to-peak value is computed over the elements of the array.
  • where: Optional parameter, a boolean array used to specify that only elements satisfying the condition are considered.

Example

import numpy as np a = np.array([[3,7,5],[8,4,3],[2,4,9]]) print ('Our array is:') print (a) print ('\n') print ('Calling the ptp() function:') print (np.ptp(a)) print ('\n') print ('Calling the ptp() function along axis 1:') print (np.ptp(a, axis = 1)) print ('\n') print ('Calling the ptp() function along axis 0:') print (np.ptp(a, axis = 0))

The output result is:

我们的数组是:
[[3 7 5]
 [8 4 3]
 [2 4 9]]


调用 ptp() 函数:
7


沿轴 1 调用 ptp() 函数:
[4 5 7]


沿轴 0 调用 ptp() 函数:
[6 3 6]

numpy.percentile()

A percentile is a measure used in statistics that represents the percentage of observations less than this value. The numpy.percentile() function accepts the following parameters.

numpy.percentile(a, q, axis)

Parameter description:

  • a: input array
  • q: The percentile to compute, between 0 and 100
  • axis: The axis along which to compute the percentile

First, clarify percentiles:

The p-th percentile is a value such that at least p% of data items are less than or equal to this value, and at least (100-p)% of data items are greater than or equal to this value.

For example: college entrance exam scores are often reported in percentile form. For instance, suppose a candidate's raw score in the Chinese language section of the entrance exam is 54 points. It is not easy to know how he did relative to other students who took the same exam. But if the raw score of 54 points corresponds exactly to the 70th percentile, we can know that about 70% of students scored lower than him, and about 30% scored higher.

Here p = 70.

Example

import numpy as np a = np.array([[10, 7, 4], [3, 2, 1]]) print ('Our array is:') print (a) print ('Calling the percentile() function:') # The 50% percentile is the median of a after sorting print (np.percentile(a, 50)) # axis is 0, compute along columns print (np.percentile(a, 50, axis=0)) # axis is 1, compute along rows print (np.percentile(a, 50, axis=1)) # Keep the dimension unchanged print (np.percentile(a, 50, axis=1, keepdims=True))

The output result is:

我们的数组是:
[[10  7  4]
 [ 3  2  1]]
调用 percentile() 函数:
3.5
[6.5 4.5 2.5]
[7. 2.]
[[7.]
 [2.]]

numpy.median()

The numpy.median() function is used to compute the median (middle value) of elements in array a.

numpy.median(a, axis=None, out=None, overwrite_input=False, keepdims=<no value>)

Parameter description:

  • a: The input array, which can be a NumPy array or an array-like object.
  • axis: Optional parameter, used to specify on which axis to compute the median. If not provided, the median of the entire array is computed. It can be an integer representing the axis index, or a tuple representing multiple axes.
  • out: Optional parameter, used to specify the storage location for the result.
  • overwrite_input: Optional parameter. If True, the use of the input array's memory is allowed during computation. This may improve performance in some cases, but may modify the contents of the input array.
  • keepdims: Optional parameter. If True, the number of dimensions of the result array is kept the same as the input array. If False (default), axes with dimension 1 after computation are removed.

Example

import numpy as np a = np.array([[30,65,70],[80,95,10],[50,90,60]]) print ('Our array is:') print (a) print ('\n') print ('Calling the median() function:') print (np.median(a)) print ('\n') print ('Calling the median() function along axis 0:') print (np.median(a, axis = 0)) print ('\n') print ('Calling the median() function along axis 1:') print (np.median(a, axis = 1))

The output result is:

我们的数组是:
[[30 65 70]
 [80 95 10]
 [50 90 60]]


调用 median() 函数:
65.0


沿轴 0 调用 median() 函数:
[50. 90. 60.]


沿轴 1 调用 median() 函数:
[65. 80. 60.]

numpy.mean()

The numpy.mean() function returns the arithmetic mean of elements in an array. If an axis is provided, it is computed along that axis.

The arithmetic mean is the sum of elements along an axis divided by the number of elements.

numpy.mean(a, axis=None, dtype=None, out=None, keepdims=<no value>)

Parameter description:

  • a: The input array, which can be a NumPy array or an array-like object.
  • axis: Optional parameter, used to specify on which axis to compute the mean. If not provided, the mean of the entire array is computed. It can be an integer representing the axis index, or a tuple representing multiple axes.
  • dtype: Optional parameter, used to specify the data type of the output. If not provided, an appropriate data type is selected based on the type of the input data.
  • out: Optional parameter, used to specify the storage location for the result.
  • keepdims: Optional parameter. If True, the number of dimensions of the result array is kept the same as the input array. If False (default), axes with dimension 1 after computation are removed.

Example

import numpy as np a = np.array([[1,2,3],[3,4,5],[4,5,6]]) print ('Our array is:') print (a) print ('\n') print ('Calling the mean() function:') print (np.mean(a)) print ('\n') print ('Calling the mean() function along axis 0:') print (np.mean(a, axis = 0)) print ('\n') print ('Calling the mean() function along axis 1:') print (np.mean(a, axis = 1))

The output result is:

我们的数组是:
[[1 2 3]
 [3 4 5]
 [4 5 6]]


调用 mean() 函数:
3.6666666666666665


沿轴 0 调用 mean() 函数:
[2.66666667 3.66666667 4.66666667]


沿轴 1 调用 mean() 函数:
[2. 4. 5.]

numpy.average()

The numpy.average() function computes the weighted average of elements in an array based on their respective weights given in another array.

This function can accept an axis parameter. If no axis is specified, the array is flattened.

The weighted average is to multiply each value by its corresponding weight, sum them up to obtain the total value, and then divide by the number of units.

Consider the array [1,2,3,4] and the corresponding weights [4,3,2,1]. The weighted average is calculated by adding the products of the corresponding elements and dividing the sum by the sum of the weights.

加权平均值 = (1*4+2*3+3*2+4*1)/(4+3+2+1)

Function syntax:

numpy.average(a, axis=None, weights=None, returned=False)

Parameter description:

  • a: The input array, which can be a NumPy array or an array-like object.
  • axis: Optional parameter, used to specify on which axis to compute the weighted average. If not provided, the weighted average of the entire array is computed. It can be an integer representing the axis index, or a tuple representing multiple axes.
  • weights: Optional parameter, used to specify the weights for the corresponding data points. If no weights array is provided, equal weights are assumed by default.
  • returned: Optional parameter. If True, both the weighted average and the sum of weights are returned.

Example

import numpy as np a = np.array([1,2,3,4]) print ('Our array is:') print (a) print ('\n') print ('Calling the average() function:') print (np.average(a)) print ('\n') # When no weights are specified, it is equivalent to the mean function wts = np.array([4,3,2,1]) print ('Call the average() function again:') print (np.average(a,weights = wts)) print ('\n') # If the returned parameter is set to true, returns the sum of the weights print ('Sum of the weights:') print (np.average([1,2,3, 4],weights = [4,3,2,1], returned = True))

The output result is:

我们的数组是:
[1 2 3 4]


调用 average() 函数:
2.5


再次调用 average() 函数:
2.0


权重的和:
(2.0, 10.0)

In multi-dimensional arrays, you can specify the axis used for calculation.

Example

import numpy as np a = np.arange(6).reshape(3,2) print ('Our array is:') print (a) print ('\n') print ('The modified array is:') wt = np.array([3,5]) print (np.average(a, axis = 1, weights = wt)) print ('\n') print ('The modified array is:') print (np.average(a, axis = 1, weights = wt, returned = True))

The output result is:

我们的数组是:
[[0 1]
 [2 3]
 [4 5]]


修改后的数组:
[0.625 2.625 4.625]


修改后的数组:
(array([0.625, 2.625, 4.625]), array([8., 8., 8.]))

Standard Deviation

Standard deviation is a measure of the dispersion of a set of data values from their mean.

Standard deviation is the square root of the variance.

The standard deviation formula is as follows:

std = sqrt(mean((x - x.mean())**2))

If the array is [1, 2, 3, 4], then its mean is 2.5. Therefore, the squared differences are [2.25, 0.25, 0.25, 2.25], and then take the square root of the mean of the squared differences, i.e., sqrt(5/4), the result is 1.1180339887498949.

Example

import numpy as np print (np.std([1,2,3,4]))

The output result is:

1.1180339887498949

Variance

In statistics, variance (sample variance) is the mean of the squared differences between each sample value and the mean of all sample values, i.e., mean((x - x.mean())** 2).

In other words, the standard deviation is the square root of the variance.

Example

import numpy as np print (np.var([1,2,3,4]))

The output result is:

1.25
Other extensions