Pandas String Operations
Pandas provides powerful string processing capabilities. Through the.straccessor, you can process each element in a Series just like operating on Python strings.
str Accessor Overview
When the Series data type isobjectyou can use the.straccessor to perform string operations.
Example
# Create a Series containing strings
s = pd.Series(["hello", "world", "pandas", "Python"])
print("Original Series:")
print(s)
print()
# Convert to lowercase
print("Lowercase:")
print(s.str.lower())
print()
# Convert to uppercase
print("Uppercase:")
print(s.str.upper())
print()
# Capitalize first letter
print("Capitalized:")
print(s.str.capitalize())
Common String Functions
Case Conversion
| Method | Description | Example |
|---|---|---|
str.lower() |
Convert to lowercase | "Hello" → "hello" |
str.upper() |
Convert to uppercase | "Hello" → "HELLO" |
str.title() |
Capitalize first letter | "hello world" → "Hello World" |
str.capitalize() |
Capitalize first letter, lowercase the rest | "hELLO" → "Hello" |
str.swapcase() |
Swap case | "HeLLo" → "hEllO" |
String Search and Replace
Example
s = pd.Series(["hello world", "python pandas", "data science", "machine learning"])
print("Original data:")
print(s)
print()
# Check containment
print("Contains 'python':")
print(s.str.contains("python", case=False))
print()
# Check start/end
print("Starts with 'hello':")
print(s.str.startswith("hello"))
print()
# Replace
print("Replace spaces with underscores:")
print(s.str.replace(" ", "_"))
print()
# Split
print("Split by space:")
print(s.str.split(" "))
Whitespace Removal and Padding
Example
s = pd.Series([" hello ", "world ", " pandas", " python "])
print("Original data:")
print(s)
print()
# Strip leading and trailing whitespace
print("Strip leading/trailing whitespace:")
print(s.str.strip())
print()
# Strip left whitespace
print("Strip left whitespace:")
print(s.str.lstrip())
print()
# Strip right whitespace
print("Strip right whitespace:")
print(s.str.rstrip())
print()
# Pad
print("Pad with 0 on the left to length 10:")
print(s.str.pad(10, side="left", fillchar="0"))
Slicing and Concatenation
Example
s = pd.Series(["hello", "world", "python", "pandas"])
print("Original data:")
print(s)
print()
# Slice
print("First 3 characters:")
print(s.str[:3])
print()
# Take 3 characters starting from position 1
print("Take 3 characters from position 1:")
print(s.str[1:4])
print()
# String concatenation
s1 = pd.Series(["Hello", "World"])
s2 = pd.Series(["Python", "Pandas"])
print("Concatenate with +:")
print(s1 + " " + s2)
print()
# Join with join() (using a specified separator)
print("Join with join():")
print(s1.str.cat(s2, sep="-"))
Regular Expression Support
Pandas string functions support regular expressions, making them a powerful tool for handling complex string patterns.
Example
s = pd.Series([
"[email protected]",
"[email protected]",
"invalid-email",
"[email protected]"
])
print("Original data:")
print(s)
print()
# Extract email
print("Extract username:")
print(s.str.extract(r"(\w+)@"))
print()
print("Extract domain:")
print(s.str.extract(r"@(\w+\.\w+)"))
print()
# Replace
print("Replace email with [email]:")
print(s.str.replace(r"\w+@\w+\.\w+", "[email]", regex=True))
print()
# Match containment
print("Contains .com or .org:")
print(s.str.contains(r"\.com|\.org"))
Extraction and Splitting
Example
s = pd.Series(["2024-01-01", "2024-02-15", "2024-03-20"])
print("Original data:")
print(s)
print()
# Extract year, month, day
extract_df = s.str.extract(r"(\d{4})-(\d{2})-(\d{2})")
extract_df.columns = ["Year", "Month", "Day"]
print("Extract year/month/day:")
print(extract_df)
print()
# Split by delimiter
s2 = pd.Series(["a,b,c", "d,e,f", "g,h,i"])
print("Split by comma:")
print(s2.str.split(","))
print()
# Split and expand into a DataFrame
print("Split and expand:")
print(s2.str.split(",", expand=True))
Statistics and Conversion
Length and Counting
Example
s = pd.Series(["hello", "world", "python", "pandas"])
print("Original data:")
print(s)
print()
# String length
print("String length:")
print(s.str.len())
print()
# Character count
s2 = pd.Series(["hello", "world", "aaa", "bbb"])
print("Occurrences of character 'l':")
print(s2.str.count("l"))
Type Conversion
Example
# String to numeric
s = pd.Series(["1", "2", "3", "abc"])
print("String to numeric:")
print(pd.to_numeric(s, errors="coerce"))
print()
# Numeric to string
s2 = pd.Series([1, 2, 3, 4])
print("Numeric to string:")
print(s2.astype(str).str.zfill(4)) # Zero-pad to 4 digits
Practical: Data Cleaning
Example
import numpy as np
# Simulate dirty data
df = pd.DataFrame({
"Name": [" Zhang San ", "Li Si", "Wang Wu "],
"Phone": ["138-0000-0000", "13900000000", " 010-12345678 "],
"Email": ["[email protected]", "[email protected]", "[email protected] "]
})
print("Original data:")
print(df)
print()
# Cleaning steps
# 1. Remove whitespace
df["Name"] = df["Name"].str.strip()
df["Phone"] = df["Phone"].str.replace("-", "").str.strip()
df["Email"] = df["Email"].str.strip().str.lower()
print("After cleaning:")
print(df)
print()
# 2. Validate data
print("Email validation:")
print(df["Email"].str.contains(r"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$"))
Important Notes
1. Handling missing values
String operations skip missing values (NaN) by default. If you need to handle them, you can use.fillna()to fill them first.
2. Case sensitivity
Most functions are case-sensitive by default. Use thecase=Falseparameter for case-insensitive matching.
3. Regular expression performance
Regular expressions are powerful but relatively slow. Watch performance when processing large amounts of data.
Other ExtensionsString processing is an important part of data cleaning. Using the
.straccessor properly can efficiently accomplish most text processing tasks.