Pandas df.sort_values() Function
df.sort_values()It is a function in Pandas used to sort DataFrame data according to one or more columns.
Sorting is a basic operation in data analysis that helps you better understand data and discover patterns.sort_values()It supports sorting by a single column or multiple columns, ascending or descending order, and can handle missing values. It is an indispensable tool in data processing.
Basic Syntax and Parameters
sort_values()It is a member function of DataFrame, called through the dot operator.to invoke.
Syntax Format
DataFrame.sort_values(by, axis=0, ascending=True, inplace=False, kind='quicksort', na_position='last', ignore_index=False, key=None)
Parameter Description
| Parameter | Type | Required? | Description | Default Value |
|---|---|---|---|---|
| by | str or list of str | Required | Specifies the column name(s) used for sorting. If it is a list, sorting is performed according to the order in the list. | None |
| axis | int or str | Optional | Specifies the axis for sorting.0or'index'0 means sort by columns (default);1or'columns'1 means sort by rows. |
0 |
| ascending | bool or list of bool | Optional | Specifies the sort order.TrueTrue means ascending (from small to large);FalseFalse means descending (from large to small). If it is a list, it corresponds one-to-one withbythe columns in the 'by' parameter. |
True |
| inplace | bool | Optional | If TrueTrue, it modifies the original DataFrame directly and returns no new object; if FalseFalse, it returns a new DataFrame and the original data remains unchanged. |
False |
| kind | str | Optional | Specifies the sorting algorithm.'quicksort''quicksort' (quick sort),'mergesort''mergesort' (merge sort),'heapsort''heapsort' (heap sort),'stable''stable' (stable sort). |
'quicksort' |
| na_position | str | Optional | Specifies the position of missing values (NaN).'first''first' means placed at the beginning;'last''last' means placed at the end. |
'last' |
| ignore_index | bool | Optional | If TrueTrue, it resets the sorted index, numbering from 0; if FalseFalse, it retains the original index. |
False |
| key | callable | Optional | Applies a function to the data before sorting, commonly used for custom sorting rules. | None |
Return Value Description
- Returns a new DataFrame (if
inplace=False), orNone(ifinplace=True)。 - the returned DataFrame has been sorted according to the specified rules.
Examples
Let us thoroughly mastersort_values()the usage of df.sort_values().
Example 1: Sort by a Single Column in Ascending Order
The simplest usage is to sort ascending by the values of a column.
Example
# Create a student grade DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
Mathematics: [85, 92, 78, 90, 88],
English: [90, 85, 92, 88, 78]
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Sort ascending by math score
df_sorted = df.sort_values(by=Mathematics)
print(After sorting by math scores in ascending order:)
print(df_sorted)
Expected output:
原始数据: 姓名 数学 英语 0 张三 85 90 1 李四 92 85 2 王五 78 92 3 赵六 90 88 4 钱七 88 78 ================================================== 按数学成绩升序排序后: 姓名 数学 英语 2 王五 78 92 0 张三 85 90 4 钱七 88 78 3 赵六 90 88 1 李四 92 85
Code explanation:
- The default sort order is ascending (from small to large).
- The math score 78 is placed first, and 92 is placed last.
- After sorting, the index remains unchanged (2, 0, 4, 3, 1).
Example 2: Sort by a Single Column in Descending Order
Useascending=Falseto achieve descending sorting.
Example
# Create a student grade DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
Mathematics: [85, 92, 78, 90, 88],
English: [90, 85, 92, 88, 78]
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Sort descending by math score
df_sorted = df.sort_values(by=Mathematics, ascending=False)
print(After sorting by math scores in descending order:)
print(df_sorted)
Expected output:
原始数据: 姓名 数学 英语 0 张三 85 90 1 李四 92 85 2 王五 78 92 3 赵六 90 88 4 钱七 88 78 ================================================== 按数学成绩降序排序后: 姓名 数学 英语 1 李四 92 85 3 赵六 90 88 4 钱七 88 78 0 张三 85 90 2 王五 78 92
Code explanation:
- Use
ascending=Falsecan achieve descending sorting from large to small. - Li Si's math score of 92 is placed first.
Example 3: Sort by Multiple Columns
You can sort by multiple columns: first sort by the first column, and when the first column is the same, sort by the second column.
Example
# Create a student grade DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
Class: [Class 1, Class 2, Class 1, Class 2, Class 1],
Mathematics: [85, 92, 78, 90, 88]
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# First sort by class, then by math score (both ascending)
df_sorted = df.sort_values(by=[Class, Mathematics])
print("After sorting by class first, then by math score:")
print(df_sorted)
print("=" * 50)
# Sort by class ascending, math score descending
df_sorted2 = df.sort_values(by=[Class, Mathematics], ascending=[True, False])
print(Sort by class ascending, math score descending:)
print(df_sorted2)
Expected output:
原始数据: 姓名 班级 数学 0 张三 一班 85 1 李四 二班 92 2 王五 一班 78 3 赵六 二班 90 4 钱七 一班 88 ================================================== 先按班级,再按数学成绩排序后: 姓名 班级 数学 2 王五 一班 78 0 张三 一班 85 4 钱七 一班 88 3 赵六 二班 90 1 李四 二班 92 ================================================== 按班级升序,数学成绩降序: 姓名 班级 数学 4 钱七 一班 88 0 张三 一班 85 2 王五 一班 78 1 李四 二班 92 3 赵六 二班 90
Code explanation:
- Use a list
by=['班级', '数学']to sort by multiple columns. - First sort by "class", with Class 1 in front and Class 2 behind.
- Within the same class, further sort by "math" score.
- Use
ascending=[True, False]to specify different sorting directions for different columns.
Example 4: Sorting with Missing Values
Usena_positionparameter to control the position of missing values.
Example
import numpy as np
# Create a DataFrame with missing values
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'Score': [85, np.nan, 92, 78]
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Default: missing values placed last
df_sorted_last = df.sort_values(by='Score')
print("Missing values placed at the end:")
print(df_sorted_last)
print("=" * 50)
# Place missing values first
df_sorted_first = df.sort_values(by='Score', na_position='first')
print("Missing values placed at the front:")
print(df_sorted_first)
Expected output:
原始数据: 姓名 成绩 0 张三 85.0 1 李四 NaN 2 王五 92.0 3 赵六 78.0 ================================================== 缺失值放在最后: 姓名 成绩 3 赵六 78.0 0 张三 85.0 2 王五 92.0 1 李四 NaN ================================================== 缺失值放在最前: 姓名 成绩 1 李四 NaN 3 赵六 78.0 0 张三 85.0 2 王五 92.0
Code explanation:
- By default (
na_position='last'), missing values are placed last. - Set
na_position='first', and missing values are placed first.
Example 5: Resetting the Index After Sorting
After sorting, the index may become non-continuous; you can useignore_index=Trueto reset the index.
Example
# Create a DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
'Score': [85, 92, 78]
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Sort without resetting the index
df_sorted = df.sort_values(by='Score')
print(After sorting (with original indices preserved):)
print(df_sorted)
print("=" * 50)
# Sort and reset the index
df_sorted_reset = df.sort_values(by='Score', ignore_index=True)
print(After sorting (reset index):)
print(df_sorted_reset)
Expected output:
原始数据: 姓名 成绩 0 张三 85 1 李四 92 2 王五 78 ================================================== 排序后(保留原始索引): 姓名 成绩 2 王五 78 0 张三 85 1 李四 92 ================================================== 排序后(重置索引): 姓名 成绩 0 王五 78 1 张三 85 2 李四 92
Code explanation:
- Using
ignore_index=Trueafterwards, the index is renumbered starting from 0. - This is more convenient when you need to access data by index later.
Example 6: Custom Sorting Using the key Parameter
keyThe parameter allows you to apply a custom function to the data before sorting.
Example
# Create a DataFrame containing a mix of uppercase and lowercase strings
data = {
'Word': ['apple', 'Banana', 'cherry', 'APPLE', 'banana']
}
df = pd.DataFrame(data)
print(Original data:)
print(df)
print("=" * 50)
# Sort alphabetically in ascending order (case-sensitive)
df_sorted_case = df.sort_values(by='Word')
print(Alphabetical ascending (case-sensitive):)
print(df_sorted_case)
print("=" * 50)
# Use the key parameter to achieve case-insensitive sorting
df_sorted_nocase = df.sort_values(by='Word', key=lambda x: x.str.lower())
print(Alphabetical ascending (case-insensitive):)
print(df_sorted_nocase)
Expected output:
原始数据:
单词
0 apple
1 Banana
2 cherry
3 APPLE
4 banana
==================================================
按字母升序(区分大小写):
单词
3 APPLE
0 apple
1 Banana
4 banana
2 cherry
==================================================
按字母升序(不区分大小写):
单词
0 apple
3 APPLE
1 Banana
4 banana
2 cherry
Code explanation:
- By default, sorting is case-sensitive, with uppercase letters placed before lowercase letters.
- Using
key=lambda x: x.str.lower()all strings are converted to lowercase before sorting, achieving case-insensitive sorting.
Example 7: In-place Modification Using the inplace Parameter
Usinginplace=Trueyou can modify directly on the original DataFrame.
Example
# Create a DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
'Score': [85, 92, 78]
}
df = pd.DataFrame(data)
print(Original data:)
print(df)
print(fOriginal data id: {id(df)})
print("=" * 50)
# Sort in place using inplace=True
df.sort_values(by='Score', inplace=True)
print(After sorting with inplace=True:)
print(df)
print(fSorted data id: {id(df)})
Expected output:
原始数据: 姓名 成绩 0 张三 85 1 李四 92 2 王五 78 原始数据 id: 140234567890 ================================================== 使用 inplace=True 排序后: 姓名 成绩 2 王五 78 0 张三 85 1 李四 92 排序后数据 id: 140234567890 # 同一个对象
Code explanation:
- Using
inplace=Trueafterwards, the DataFrame itself is modified, and no new object is returned. - The id before and after sorting
idare the same, indicating it is the same object. - This method can save memory, but it will modify the original data.
Notes
sort_values()By default, a new DataFrame is returned and the original data is not modified. To modify in place, useinplace=Truethe parameter.- The sorting operation is not performed in place (unless specified
inplace=True), the original DataFrame remains unchanged. - Using
ignore_index=Truecan reset the index, which is more convenient in subsequent data processing. keyThe parameter is very powerful and can implement various custom sorts, such as case-insensitive, sorting by absolute value, etc.- When sorting by multiple columns,
ascendingthe parameter can be a list, corresponding one-to-one withbythe columns in ...
Other extensions
Pandas Common Functions