Pandas df.loc() Function

Pandas 常用函数Pandas Common Functions


loc[]It is a label-based indexing method in Pandas, used to select data through row labels and column labels. It is one of the most intuitive and commonly used DataFrame indexing methods, allowing precise location of the data we need.

Unlike position-based indexing (such asiloc[]),loc[]it uses the index labels set in the data itself. This means that even if the position of the data changes, as long as the labels remain unchanged, we can still accurately select the target data.


Basic Syntax and Parameters

loc[]It is the indexer of DataFrame, accessed through square brackets[]to access. Various forms of parameters can be passed inside the square brackets to select rows, columns, or specific cells.

Syntax Format

# 选择单行(返回 Series)
DataFrame.loc[行标签]

# 选择多行(返回 DataFrame)
DataFrame.loc[[行标签1, 行标签2, ...]]

# 选择行和列
DataFrame.loc[行标签, 列标签]
DataFrame.loc[行切片, 列切片]
DataFrame.loc[行条件, 列条件]

Parameter Description

Parameter Position Parameter Type Description
First parameter (rows) Label, list of labels, label slice, boolean array Used to select rows; can be a single label, multiple labels, a slice, or a condition expression.
Second parameter (columns) Label, list of labels, label slice Optional, used to select columns; the syntax is the same as row selection.

Return Value Description

  • Single element: Returns a scalar value (a single value).
  • Single row: Returns a Series.
  • Multiple rows: Returns a DataFrame.
  • Combination of rows and columns: Returns a Series or DataFrame depending on the selection result.

Examples

Let's fully master through rich examplesloc[]the usage of ...

Example 1: Basic Usage - Using Default Index

When the DataFrame uses the default integer index,loc[]the behavior of ...iloc[]is similar, but uses labels rather than positions.

Example

import pandas as pd

# Create an example DataFrame
data = {
    'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
    'age': [18, 19, 17, 18, 20],
    'score': [85, 92, 78, 90, 88],
    'grade': ['A', 'A', 'B', 'A', 'B']
}
df = pd.DataFrame(data)

print("Original DataFrame:")
print(df)
print()

# Use loc to select a single row (via index label)
print("Select row 0:")
print(df.loc[0])
print()

# Select multiple rows
print("Select rows 1, 3, 4:")
print(df.loc[[1, 3, 4]])
print()

# Use slicing to select consecutive rows (note: loc slicing is inclusive on both ends)
print("Select rows 1 through 3 (inclusive):")
print(df.loc[1:3])

Output:

原始 DataFrame:
      name  age  score grade
0    Alice   18     85     A
1      Bob   19     92     A
2  Charlie   17     78     B
3    David   18     90     A
4      Eve   20     88     B

选择第 0 行:
name      Alice
age          18
score        85
grade         A
Name: 0, dtype: object

选择第 1、3、4 行:
    name  age  score grade
1    Bob   19     92     A
3  David   18     90     A
4      Eve   20     88     B

选择第 1 行到第 3 行(包含两端):
      name  age  score grade
1      Bob   19     92     A
2  Charlie   17     78     B
3    David   18     90     A

Code explanation:

  1. By default, the DataFrame's index is the integers 0, 1, 2, 3, 4.
  2. df.loc[0]Selecting the row with index label 0 returns a Series.
  3. df.loc[[1, 3, 4]]Selects multiple rows with specific indices.
  4. Important:loc[]Its slicing is inclusive on both ends (unlike Python slicing),1:3including indices 1, 2, 3.

Example 2: Using Custom Index

loc[]The real power of loc lies in the ability to use custom index labels, which is itsiloc[]fundamental difference.

Example

import pandas as pd

# Create a DataFrame with a custom index
data = {
    'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
    'age': [18, 19, 17, 18, 20],
    'score': [85, 92, 78, 90, 88],
    'grade': ['A', 'A', 'B', 'A', 'B']
}
df = pd.DataFrame(data, index=['a', 'b', 'c', 'd', 'e'])

print("DataFrame with custom index:")
print(df)
print()

# Use a custom label to select a single row
print("Select the row with index 'b':")
print(df.loc['b'])
print()

# Use custom labels for slicing
print("Select rows from index 'b' to 'd':")
print(df.loc['b':'d'])
print()

# Select multiple non-consecutive indices
print("Select rows with indices 'a', 'c', 'e':")
print(df.loc[['a', 'c', 'e']])

Output:

带自定义索引的 DataFrame:
   name  age  score grade
a  Alice   18     85     A
b    Bob   19     92     A
c  Charlie   17     78     B
d  David   18     90     A
e    Eve   20     88     B

选择索引为 'b' 的行:
name        Bob
age          19
score        92
grade         A
Name: b, dtype: object

选择索引 'b' 到 'd' 的行:
      name  age  score grade
b    Bob   19     92     A
c  Charlie   17     78     B
d  David   18     90     A

选择索引 'a', 'c', 'e' 的行:
      name  age  score grade
a  Alice   18     85     A
c  Charlie   17     78     B
e    Eve   20     88     B

Code explanation:

  • Custom index uses strings 'a', 'b', 'c', 'd', 'e' instead of integer indices.
  • loc[]Using these labels to select data is very intuitive.
  • The slice 'b':'d' includes the labels 'b', 'c', 'd', also inclusive on both ends.

Example 3: Selecting Specific Columns

loc[]Not only can you select rows, but you can also select columns at the same time.

Example

import pandas as pd

data = {
    'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
    'age': [18, 19, 17, 18, 20],
    'score': [85, 92, 78, 90, 88],
    'grade': ['A', 'A', 'B', 'A', 'B']
}
df = pd.DataFrame(data, index=['a', 'b', 'c', 'd', 'e'])

# Select specific columns of specific rows
print("Select the name and score columns for index 'b':")
print(df.loc['b', ['name', 'score']])
print()

# Select multiple rows and columns
print("Select the name and age columns for indices 'a', 'c':")
print(df.loc[['a', 'c'], ['name', 'age']])
print()

# Select specific columns for all rows
print("The name and grade columns of all rows:")
print(df.loc[:, ['name', 'grade']])
print()

# Combine row slicing and column slicing
print("The name to score columns for indices 'b' through 'd':")
print(df.loc['b':'d', 'name':'score'])

Output:

选择索引 'b' 的 name 和 score 列:
name     Bob
score    92
Name: b, dtype: object

选择索引 'a', 'c' 的 name 和 age 列:
    name  age
a  Alice   18
c  Charlie 17

所有行的 name 和 grade 列:
    name grade
a   Alice     A
b     Bob     A
c  Charlie   B
d    David    A
e     Eve      B

索引 'b' 到 'd' 的 name 到 score 列:
      name  age score
b    Bob   19    92
c  Charlie 17    78
d  David   18    90

Code explanation:

  1. df.loc['b', ['name', 'score']]Selecting specific columns of a specific row returns a Series.
  2. df.loc[['a', 'c'], ['name', 'age']]Selecting multiple rows and columns returns a DataFrame.
  3. df.loc[:, ['name', 'grade']]A colon means selecting all rows.
  4. df.loc['b':'d', 'name':'score']Use slicing for both rows and columns simultaneously.

Example 4: Using Boolean Condition Selection

loc[]It supports using boolean arrays or condition expressions to filter data, which is a very powerful feature.

Example

import pandas as pd

data = {
    'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],
    'age': [18, 19, 17, 18, 20],
    'score': [85, 92, 78, 90, 88],
    'grade': ['A', 'A', 'B', 'A', 'B']
}
df = pd.DataFrame(data)

# Use boolean condition to select rows with score greater than 85
print("Students with scores greater than 85:")
print(df.loc[df['score'] > 85])
print()

# Compound condition: score greater than 85 and age less than 19
print("Students with scores greater than 85 and age less than 19:")
print(df.loc[(df['score'] > 85) & (df['age']  85, ['name', 'score']])

Output:

分数大于 85 的学生:
      name  age  score grade
1      Bob   19     92     A
3    David   18     90     A
4      Eve   20     88     B

分数大于 85 且年龄小于 19 的学生:
      name  age  score grade
3    David   18     90     A

使用布尔数组选择的行:
      name  age  score grade
0    Alice   18     85     A
2  Charlie   17     78     B
4      Eve   20     88     B

分数大于 85 的学生的 name 和 score:
      name  score
1      Bob     92
3    David   90
4      Eve    88

Code explanation:

  1. df.loc[df['score'] > 85]Using boolean condition filtering, returns rows that meet the condition.
  2. Compound conditions need to be enclosed in parentheses and connected with&(AND) and|(OR).
  3. You can directly pass a boolean array (the length must be the same as the number of rows) to select rows.
  4. After the condition, you can also add column selection to achieve fine-grained data filtering.

Notes

  • loc[]It uses label indexing, and slicing is inclusive on both ends (unlike Python native slicing).
  • If the passed label does not exist, a KeyError will be raised.
  • In boolean conditions, compound conditions must use parentheses, and use&and|instead of Python'sandandor。
  • loc[]Reading and modifying data are both very convenient, but in big data scenarios, attention should be paid to performance.

Important notes:loc[]andiloc[]They are two different indexers; although they look similar, the principles are completely different:loc[]Based on labels,iloc[]Based on positions.


Summary

loc[]It is a label-based data selector in Pandas, providing an intuitive and flexible way to access data. Through the combination of row labels and column labels, we can accurately select any data we need.

Its main advantages include: using custom indexes to make code more readable, supporting inclusive slicing syntax, and being able to combine boolean conditions for complex filtering. In actual data analysis,loc[]it is one of the most frequently used indexing methods in daily use.

Pandas 常用函数Pandas Common Functions

Other Extensions