Pandas Series.str.contains() Function
Commonly Used Pandas Functions
Series.str.contains()is a function in Pandas used to check whether a string contains a specified substring.
In data processing, we often need to filter data based on text content, such as finding records containing specific keywords or filtering out text that matches a certain pattern.contains()The function can check whether each string element contains a specified substring or regular expression pattern.
Word Meaning:containsmeans 'to contain', indicating whether the string contains specified content.
Basic Syntax and Parameters
str.contains()is the string accessor method of Series, so you need to first have a Series containing strings, and then use the.straccessor to call it.
Syntax Format
Series.str.contains(pat, case=True, regex=True, na=None)
Parameter Description
| Parameter | Type | Required | Description | Default Value |
|---|---|---|---|---|
| pat | str | Required | The pattern to search for, which can be an ordinary string or a regular expression. | - |
| case | bool | Optional | Whether to distinguish case. Defaults to True (case-sensitive). | True |
| regex | bool | Optional | Whether to treat the pat parameter as a regular expression. Defaults to True. | True |
| na | object | Optional | The value to return when the element is NaN. Defaults to None (returns NaN). | None |
Function Description
- Return Value: returns a boolean Series indicating whether each element contains the specified pattern.
- Effect: checks each string element in the Series and returns True or False.
- Note: uses regular expression matching by default, which can match more complex patterns.
Examples
Let's thoroughly master, through a series of examples from simple to complex,str.contains()the usage of it.
Example 1: Basic Usage - Check if a Substring Is Contained
Example
# Create a Series containing text
s = pd.Series(['apple', 'banana', 'grape', 'pineapple', 'orange'])
# Check whether it contains 'ap'
result = s.str.contains('ap')
print("Original Series:")
print(s)
print("nContains 'ap':")
print(result)
Output result:
原始 Series: 0 apple 1 banana 2 grape 3 pineapple 4 orange dtype: object 是否包含 'ap': 0 True 1 False 2 True 3 True 4 False
Code explanation:
s.str.contains('ap')Check whether each string contains the substring 'ap'.- 'apple' contains 'ap', returns True.
- 'banana' does not contain 'ap', returns False.
- 'grape' contains 'ap', returns True.
- 'pineapple' contains 'ap', returns True.
Example 2: Using Regular Expression Matching
contains()Regular expressions are used by default, which can match more complex patterns.
Example
# Create a Series containing text
s = pd.Series(['hello123', 'world456', 'example789', 'python', 'test123'])
# Use a regular expression to match strings that start with a letter and are followed by digits
result = s.str.contains(r'^[a-zA-Z]+d+$')
print("Original Series:")
print(s)
print("nMatch the pattern starting with a letter and ending with a digit:")
print(result)
Output result:
原始 Series: 0 hello123 1 world456 2 example789 3 python 4 test123 dtype: object 匹配字母开头+数字结尾的模式: 0 True 1 True 2 True 3 False 4 True
Code explanation:
r'^[a-zA-Z]+d+$'It is a regular expression that matches strings starting with a letter and ending with a digit.^indicates the beginning of the string,$indicates the end of the string.- 'python' does not contain digits, so it does not match.
Example 3: Case-Insensitive Matching
By settingcase=Falsecase-insensitive matching can be achieved.
Example
# Create a Series containing text in different cases
s = pd.Series(['Apple', 'APPLE', 'apple', 'Banana', 'APPLE'])
# Case-sensitive matching
result_case = s.str.contains('APPLE')
# Case-insensitive matching
result_nocase = s.str.contains('APPLE', case=False)
print("Original Series:")
print(s)
print("nCase-sensitive match 'APPLE':")
print(result_case)
print("nCase-insensitive match 'APPLE':")
print(result_nocase)
Output result:
原始 Series: 0 Apple 1 APPLE 2 apple 3 Banana 4 APPLE dtype: object 区分大小写匹配 'APPLE': 0 False 1 True 2 False 3 False 4 True 不区分大小写匹配 'APPLE': 0 True 1 True 2 True 3 False 4 True
Code explanation:
case=True(default) Case-sensitive, only matches exactly 'APPLE'.case=FalseCase-insensitive, 'Apple', 'APPLE', and 'apple' all match.
Example 4: Handling Missing Values
Throughnathe parameter, you can specify how to handle NaN values.
Example
import numpy as np
# Create a Series containing NaN values
s = pd.Series(['apple', 'banana', np.nan, 'grape', None])
# Default handling of NaN (returns NaN)
result_default = s.str.contains('ap')
# Specify that NaN returns False
result_na_false = s.str.contains('ap', na=False)
print("Original Series:")
print(s)
print("nDefault handling (returns NaN):")
print(result_default)
print("nTreat NaN as False:")
print(result_na_false)
Output result:
原始 Series: 0 apple 1 banana 2 NaN 3 grape 4 None dtype: object 默认处理(返回 NaN): 0 True 1 False 2 NaN 3 True 4 NaN 将 NaN 视为 False: 0 True 1 False 2 False 3 True 4 False
Code explanation:
- By default, NaN and None return NaN.
- Setting
na=Falsecan treat NaN as not matching. - In actual data processing, this is useful and can avoid filtering issues caused by NaN.
Example 5: Filtering Data
contains()It is often combined with boolean indexing to filter data.
Example
# Create a simulated product data Series
products = pd.Series([
'iPhone 14 Pro',
'Samsung Galaxy S23',
'iPhone 13',
'Google Pixel 7',
'iPad Pro',
'MacBook Air',
'Dell XPS 15'
])
# Filter products containing 'iPhone'
iphone_products = products[products.str.contains('iPhone')]
# Filter products that do not contain 'i' (case-sensitive)
no_i_products = products[~products.str.contains('i')]
print("All products:")
print(products)
print("nProducts containing 'iPhone':")
print(iphone_products)
print("nProducts not containing uppercase 'I':")
print(no_i_products)
Output result:
所有产品: 0 iPhone 14 Pro 1 Samsung Galaxy S23 2 iPhone 13 3 Google Pixel 7 4 iPad Pro 5 MacBook Air 6 Dell XPS 15 dtype: object 包含 'iPhone' 的产品: 0 iPhone 14 Pro 2 iPhone 13 dtype: object 不包含大写 'I' 的产品: 6 Dell XPS 15
Code explanation:
products[products.str.contains('iPhone')]Filter out products containing 'iPhone'.~products.str.contains('i')Use~negation to filter out products that do not contain 'i'.- This is a common operation in data filtering.
Notes
str.contains()Regular expressions are used by default (regex=True)。- If you need to match ordinary strings (not as regular expressions), you can set
regex=False。 - By default, it is case-sensitive; use
case=Falseto make it case-insensitive. - By default, NaN values return NaN; use
naparameter to specify the return value. - This function returns a boolean Series, which can be directly used for boolean indexing to filter data.
Other Extensions