Pandas df.loc() Function
loc[]It is a label-based indexing method in Pandas, used to select data through row labels and column labels. It is one of the most intuitive and commonly used DataFrame indexing methods, allowing precise location of the data we need.
Unlike position-based indexing (such asiloc[]),loc[]it uses the index labels set in the data itself. This means that even if the position of the data changes, as long as the labels remain unchanged, we can still accurately select the target data.
Basic Syntax and Parameters
loc[]It is the indexer of DataFrame, accessed through square brackets[]to access. Various forms of parameters can be passed inside the square brackets to select rows, columns, or specific cells.
Syntax Format
# 选择单行(返回 Series) DataFrame.loc[行标签] # 选择多行(返回 DataFrame) DataFrame.loc[[行标签1, 行标签2, ...]] # 选择行和列 DataFrame.loc[行标签, 列标签] DataFrame.loc[行切片, 列切片] DataFrame.loc[行条件, 列条件]
Parameter Description
| Parameter Position | Parameter Type | Description |
|---|---|---|
| First parameter (rows) | Label, list of labels, label slice, boolean array | Used to select rows; can be a single label, multiple labels, a slice, or a condition expression. |
| Second parameter (columns) | Label, list of labels, label slice | Optional, used to select columns; the syntax is the same as row selection. |
Return Value Description
- Single element: Returns a scalar value (a single value).
- Single row: Returns a Series.
- Multiple rows: Returns a DataFrame.
- Combination of rows and columns: Returns a Series or DataFrame depending on the selection result.
Examples
Let's fully master through rich examplesloc[]the usage of ...
Example 1: Basic Usage - Using Default Index
When the DataFrame uses the default integer index,loc[]the behavior of ...iloc[]is similar, but uses labels rather than positions.
Example
# Create an example DataFrame
data = {
'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
'age': [18, 19, 17, 18, 20],
'score': [85, 92, 78, 90, 88],
'grade': ['A', 'A', 'B', 'A', 'B']
}
df = pd.DataFrame(data)
print("Original DataFrame:")
print(df)
print()
# Use loc to select a single row (via index label)
print("Select row 0:")
print(df.loc[0])
print()
# Select multiple rows
print("Select rows 1, 3, 4:")
print(df.loc[[1, 3, 4]])
print()
# Use slicing to select consecutive rows (note: loc slicing is inclusive on both ends)
print("Select rows 1 through 3 (inclusive):")
print(df.loc[1:3])
Output:
原始 DataFrame:
name age score grade
0 Alice 18 85 A
1 Bob 19 92 A
2 Charlie 17 78 B
3 David 18 90 A
4 Eve 20 88 B
选择第 0 行:
name Alice
age 18
score 85
grade A
Name: 0, dtype: object
选择第 1、3、4 行:
name age score grade
1 Bob 19 92 A
3 David 18 90 A
4 Eve 20 88 B
选择第 1 行到第 3 行(包含两端):
name age score grade
1 Bob 19 92 A
2 Charlie 17 78 B
3 David 18 90 A
Code explanation:
- By default, the DataFrame's index is the integers 0, 1, 2, 3, 4.
df.loc[0]Selecting the row with index label 0 returns a Series.df.loc[[1, 3, 4]]Selects multiple rows with specific indices.- Important:
loc[]Its slicing is inclusive on both ends (unlike Python slicing),1:3including indices 1, 2, 3.
Example 2: Using Custom Index
loc[]The real power of loc lies in the ability to use custom index labels, which is itsiloc[]fundamental difference.
Example
# Create a DataFrame with a custom index
data = {
'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
'age': [18, 19, 17, 18, 20],
'score': [85, 92, 78, 90, 88],
'grade': ['A', 'A', 'B', 'A', 'B']
}
df = pd.DataFrame(data, index=['a', 'b', 'c', 'd', 'e'])
print("DataFrame with custom index:")
print(df)
print()
# Use a custom label to select a single row
print("Select the row with index 'b':")
print(df.loc['b'])
print()
# Use custom labels for slicing
print("Select rows from index 'b' to 'd':")
print(df.loc['b':'d'])
print()
# Select multiple non-consecutive indices
print("Select rows with indices 'a', 'c', 'e':")
print(df.loc[['a', 'c', 'e']])
Output:
带自定义索引的 DataFrame:
name age score grade
a Alice 18 85 A
b Bob 19 92 A
c Charlie 17 78 B
d David 18 90 A
e Eve 20 88 B
选择索引为 'b' 的行:
name Bob
age 19
score 92
grade A
Name: b, dtype: object
选择索引 'b' 到 'd' 的行:
name age score grade
b Bob 19 92 A
c Charlie 17 78 B
d David 18 90 A
选择索引 'a', 'c', 'e' 的行:
name age score grade
a Alice 18 85 A
c Charlie 17 78 B
e Eve 20 88 B
Code explanation:
- Custom index uses strings 'a', 'b', 'c', 'd', 'e' instead of integer indices.
loc[]Using these labels to select data is very intuitive.- The slice 'b':'d' includes the labels 'b', 'c', 'd', also inclusive on both ends.
Example 3: Selecting Specific Columns
loc[]Not only can you select rows, but you can also select columns at the same time.
Example
data = {
'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
'age': [18, 19, 17, 18, 20],
'score': [85, 92, 78, 90, 88],
'grade': ['A', 'A', 'B', 'A', 'B']
}
df = pd.DataFrame(data, index=['a', 'b', 'c', 'd', 'e'])
# Select specific columns of specific rows
print("Select the name and score columns for index 'b':")
print(df.loc['b', ['name', 'score']])
print()
# Select multiple rows and columns
print("Select the name and age columns for indices 'a', 'c':")
print(df.loc[['a', 'c'], ['name', 'age']])
print()
# Select specific columns for all rows
print("The name and grade columns of all rows:")
print(df.loc[:, ['name', 'grade']])
print()
# Combine row slicing and column slicing
print("The name to score columns for indices 'b' through 'd':")
print(df.loc['b':'d', 'name':'score'])
Output:
选择索引 'b' 的 name 和 score 列:
name Bob
score 92
Name: b, dtype: object
选择索引 'a', 'c' 的 name 和 age 列:
name age
a Alice 18
c Charlie 17
所有行的 name 和 grade 列:
name grade
a Alice A
b Bob A
c Charlie B
d David A
e Eve B
索引 'b' 到 'd' 的 name 到 score 列:
name age score
b Bob 19 92
c Charlie 17 78
d David 18 90
Code explanation:
df.loc['b', ['name', 'score']]Selecting specific columns of a specific row returns a Series.df.loc[['a', 'c'], ['name', 'age']]Selecting multiple rows and columns returns a DataFrame.df.loc[:, ['name', 'grade']]A colon means selecting all rows.df.loc['b':'d', 'name':'score']Use slicing for both rows and columns simultaneously.
Example 4: Using Boolean Condition Selection
loc[]It supports using boolean arrays or condition expressions to filter data, which is a very powerful feature.
Example
data = {
'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
'age': [18, 19, 17, 18, 20],
'score': [85, 92, 78, 90, 88],
'grade': ['A', 'A', 'B', 'A', 'B']
}
df = pd.DataFrame(data)
# Use boolean condition to select rows with score greater than 85
print("Students with scores greater than 85:")
print(df.loc[df['score'] > 85])
print()
# Compound condition: score greater than 85 and age less than 19
print("Students with scores greater than 85 and age less than 19:")
print(df.loc[(df['score'] > 85) & (df['age'] 85, ['name', 'score']])
Output:
分数大于 85 的学生:
name age score grade
1 Bob 19 92 A
3 David 18 90 A
4 Eve 20 88 B
分数大于 85 且年龄小于 19 的学生:
name age score grade
3 David 18 90 A
使用布尔数组选择的行:
name age score grade
0 Alice 18 85 A
2 Charlie 17 78 B
4 Eve 20 88 B
分数大于 85 的学生的 name 和 score:
name score
1 Bob 92
3 David 90
4 Eve 88
Code explanation:
df.loc[df['score'] > 85]Using boolean condition filtering, returns rows that meet the condition.- Compound conditions need to be enclosed in parentheses and connected with
&(AND) and|(OR). - You can directly pass a boolean array (the length must be the same as the number of rows) to select rows.
- After the condition, you can also add column selection to achieve fine-grained data filtering.
Notes
loc[]It uses label indexing, and slicing is inclusive on both ends (unlike Python native slicing).- If the passed label does not exist, a KeyError will be raised.
- In boolean conditions, compound conditions must use parentheses, and use
&and|instead of Python'sandandor。 loc[]Reading and modifying data are both very convenient, but in big data scenarios, attention should be paid to performance.
Important notes:
loc[]andiloc[]They are two different indexers; although they look similar, the principles are completely different:loc[]Based on labels,iloc[]Based on positions.
Summary
loc[]It is a label-based data selector in Pandas, providing an intuitive and flexible way to access data. Through the combination of row labels and column labels, we can accurately select any data we need.
Its main advantages include: using custom indexes to make code more readable, supporting inclusive slicing syntax, and being able to combine boolean conditions for complex filtering. In actual data analysis,loc[]it is one of the most frequently used indexing methods in daily use.
Pandas Common Functions