Pandas df.sort_values() Function

Pandas 常用函数Pandas Common Functions


df.sort_values()It is a function in Pandas used to sort DataFrame data according to one or more columns.

Sorting is a basic operation in data analysis that helps you better understand data and discover patterns.sort_values()It supports sorting by a single column or multiple columns, ascending or descending order, and can handle missing values. It is an indispensable tool in data processing.


Basic Syntax and Parameters

sort_values()It is a member function of DataFrame, called through the dot operator.to invoke.

Syntax Format

DataFrame.sort_values(by, axis=0, ascending=True, inplace=False, kind='quicksort', na_position='last', ignore_index=False, key=None)

Parameter Description

Parameter Type Required? Description Default Value
by str or list of str Required Specifies the column name(s) used for sorting. If it is a list, sorting is performed according to the order in the list. None
axis int or str Optional Specifies the axis for sorting.0or'index'0 means sort by columns (default);1or'columns'1 means sort by rows. 0
ascending bool or list of bool Optional Specifies the sort order.TrueTrue means ascending (from small to large);FalseFalse means descending (from large to small). If it is a list, it corresponds one-to-one withbythe columns in the 'by' parameter. True
inplace bool Optional If TrueTrue, it modifies the original DataFrame directly and returns no new object; if FalseFalse, it returns a new DataFrame and the original data remains unchanged. False
kind str Optional Specifies the sorting algorithm.'quicksort''quicksort' (quick sort),'mergesort''mergesort' (merge sort),'heapsort''heapsort' (heap sort),'stable''stable' (stable sort). 'quicksort'
na_position str Optional Specifies the position of missing values (NaN).'first''first' means placed at the beginning;'last''last' means placed at the end. 'last'
ignore_index bool Optional If TrueTrue, it resets the sorted index, numbering from 0; if FalseFalse, it retains the original index. False
key callable Optional Applies a function to the data before sorting, commonly used for custom sorting rules. None

Return Value Description

  • Returns a new DataFrame (ifinplace=False), orNone(ifinplace=True)。
  • the returned DataFrame has been sorted according to the specified rules.

Examples

Let us thoroughly mastersort_values()the usage of df.sort_values().

Example 1: Sort by a Single Column in Ascending Order

The simplest usage is to sort ascending by the values of a column.

Example

import pandas as pd

# Create a student grade DataFrame
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
    Mathematics: [85, 92, 78, 90, 88],
    English: [90, 85, 92, 88, 78]
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Sort ascending by math score
df_sorted = df.sort_values(by=Mathematics)

print(After sorting by math scores in ascending order:)
print(df_sorted)

Expected output:

原始数据:
   姓名  数学  英语
0  张三   85   90
1  李四   92   85
2  王五   78   92
3  赵六   90   88
4  钱七   88   78
==================================================
按数学成绩升序排序后:
   姓名  数学  英语
2  王五   78   92
0  张三   85   90
4  钱七   88   78
3  赵六   90   88
1  李四   92   85

Code explanation:

  1. The default sort order is ascending (from small to large).
  2. The math score 78 is placed first, and 92 is placed last.
  3. After sorting, the index remains unchanged (2, 0, 4, 3, 1).

Example 2: Sort by a Single Column in Descending Order

Useascending=Falseto achieve descending sorting.

Example

import pandas as pd

# Create a student grade DataFrame
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
    Mathematics: [85, 92, 78, 90, 88],
    English: [90, 85, 92, 88, 78]
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Sort descending by math score
df_sorted = df.sort_values(by=Mathematics, ascending=False)

print(After sorting by math scores in descending order:)
print(df_sorted)

Expected output:

原始数据:
   姓名  数学  英语
0  张三   85   90
1  李四   92   85
2  王五   78   92
3  赵六   90   88
4  钱七   88   78
==================================================
按数学成绩降序排序后:
   姓名  数学  英语
1  李四   92   85
3  赵六   90   88
4  钱七   88   78
0  张三   85   90
2  王五   78   92

Code explanation:

  • Useascending=Falsecan achieve descending sorting from large to small.
  • Li Si's math score of 92 is placed first.

Example 3: Sort by Multiple Columns

You can sort by multiple columns: first sort by the first column, and when the first column is the same, sort by the second column.

Example

import pandas as pd

# Create a student grade DataFrame
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
    Class: [Class 1, Class 2, Class 1, Class 2, Class 1],
    Mathematics: [85, 92, 78, 90, 88]
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# First sort by class, then by math score (both ascending)
df_sorted = df.sort_values(by=[Class, Mathematics])

print("After sorting by class first, then by math score:")
print(df_sorted)
print("=" * 50)

# Sort by class ascending, math score descending
df_sorted2 = df.sort_values(by=[Class, Mathematics], ascending=[True, False])

print(Sort by class ascending, math score descending:)
print(df_sorted2)

Expected output:

原始数据:
   姓名  班级  数学
0  张三  一班   85
1  李四  二班   92
2  王五  一班   78
3  赵六  二班   90
4  钱七  一班   88
==================================================
先按班级,再按数学成绩排序后:
   姓名  班级  数学
2  王五  一班   78
0  张三  一班   85
4  钱七  一班   88
3  赵六  二班   90
1  李四  二班   92
==================================================
按班级升序,数学成绩降序:
   姓名  班级  数学
4  钱七  一班   88
0  张三  一班   85
2  王五  一班   78
1  李四  二班   92
3  赵六  二班   90

Code explanation:

  • Use a listby=['班级', '数学']to sort by multiple columns.
  • First sort by "class", with Class 1 in front and Class 2 behind.
  • Within the same class, further sort by "math" score.
  • Useascending=[True, False]to specify different sorting directions for different columns.

Example 4: Sorting with Missing Values

Usena_positionparameter to control the position of missing values.

Example

import pandas as pd
import numpy as np

# Create a DataFrame with missing values
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
    'Score': [85, np.nan, 92, 78]
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Default: missing values placed last
df_sorted_last = df.sort_values(by='Score')

print("Missing values placed at the end:")
print(df_sorted_last)
print("=" * 50)

# Place missing values first
df_sorted_first = df.sort_values(by='Score', na_position='first')

print("Missing values placed at the front:")
print(df_sorted_first)

Expected output:

原始数据:
   姓名   成绩
0  张三  85.0
1  李四   NaN
2  王五  92.0
3  赵六  78.0
==================================================
缺失值放在最后:
   姓名   成绩
3  赵六  78.0
0  张三  85.0
2  王五  92.0
1  李四   NaN
==================================================
缺失值放在最前:
   姓名   成绩
1  李四   NaN
3  赵六  78.0
0  张三  85.0
2  王五  92.0

Code explanation:

  • By default (na_position='last'), missing values are placed last.
  • Setna_position='first', and missing values are placed first.

Example 5: Resetting the Index After Sorting

After sorting, the index may become non-continuous; you can useignore_index=Trueto reset the index.

Example

import pandas as pd

# Create a DataFrame
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
    'Score': [85, 92, 78]
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Sort without resetting the index
df_sorted = df.sort_values(by='Score')

print(After sorting (with original indices preserved):)
print(df_sorted)
print("=" * 50)

# Sort and reset the index
df_sorted_reset = df.sort_values(by='Score', ignore_index=True)

print(After sorting (reset index):)
print(df_sorted_reset)

Expected output:

原始数据:
   姓名  成绩
0  张三   85
1  李四   92
2  王五   78
==================================================
排序后(保留原始索引):
   姓名  成绩
2  王五   78
0  张三   85
1  李四   92
==================================================
排序后(重置索引):
   姓名  成绩
0  王五   78
1  张三   85
2  李四   92

Code explanation:

  • Usingignore_index=Trueafterwards, the index is renumbered starting from 0.
  • This is more convenient when you need to access data by index later.

Example 6: Custom Sorting Using the key Parameter

keyThe parameter allows you to apply a custom function to the data before sorting.

Example

import pandas as pd

# Create a DataFrame containing a mix of uppercase and lowercase strings
data = {
    'Word': ['apple', 'Banana', 'cherry', 'APPLE', 'banana']
}
df = pd.DataFrame(data)

print(Original data:)
print(df)
print("=" * 50)

# Sort alphabetically in ascending order (case-sensitive)
df_sorted_case = df.sort_values(by='Word')

print(Alphabetical ascending (case-sensitive):)
print(df_sorted_case)
print("=" * 50)

# Use the key parameter to achieve case-insensitive sorting
df_sorted_nocase = df.sort_values(by='Word', key=lambda x: x.str.lower())

print(Alphabetical ascending (case-insensitive):)
print(df_sorted_nocase)

Expected output:

原始数据:
    单词
0  apple
1  Banana
2  cherry
3  APPLE
4  banana
==================================================
按字母升序(区分大小写):
    单词
3  APPLE
0  apple
1  Banana
4  banana
2  cherry
==================================================
按字母升序(不区分大小写):
    单词
0  apple
3  APPLE
1  Banana
4  banana
2  cherry

Code explanation:

  • By default, sorting is case-sensitive, with uppercase letters placed before lowercase letters.
  • Usingkey=lambda x: x.str.lower()all strings are converted to lowercase before sorting, achieving case-insensitive sorting.

Example 7: In-place Modification Using the inplace Parameter

Usinginplace=Trueyou can modify directly on the original DataFrame.

Example

import pandas as pd

# Create a DataFrame
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
    'Score': [85, 92, 78]
}
df = pd.DataFrame(data)

print(Original data:)
print(df)
print(fOriginal data id: {id(df)})
print("=" * 50)

# Sort in place using inplace=True
df.sort_values(by='Score', inplace=True)

print(After sorting with inplace=True:)
print(df)
print(fSorted data id: {id(df)})

Expected output:

原始数据:
   姓名  成绩
0  张三   85
1  李四   92
2  王五   78
原始数据 id: 140234567890
==================================================
使用 inplace=True 排序后:
   姓名  成绩
2  王五   78
0  张三   85
1  李四   92
排序后数据 id: 140234567890  # 同一个对象

Code explanation:

  • Usinginplace=Trueafterwards, the DataFrame itself is modified, and no new object is returned.
  • The id before and after sortingidare the same, indicating it is the same object.
  • This method can save memory, but it will modify the original data.

Notes

  • sort_values()By default, a new DataFrame is returned and the original data is not modified. To modify in place, useinplace=Truethe parameter.
  • The sorting operation is not performed in place (unless specifiedinplace=True), the original DataFrame remains unchanged.
  • Usingignore_index=Truecan reset the index, which is more convenient in subsequent data processing.
  • keyThe parameter is very powerful and can implement various custom sorts, such as case-insensitive, sorting by absolute value, etc.
  • When sorting by multiple columns,ascendingthe parameter can be a list, corresponding one-to-one withbythe columns in ...

Pandas 常用函数Common Pandas functions

Other extensions