Pandas Series.str.split() Function

Pandas 常用函数Common Pandas Functions


Series.str.split()It is a function in Pandas used to split strings.

In data processing, we often need to split a string into multiple parts, for example, splitting a sentence into words, or splitting comma-separated values into lists.split()The function can split a string into multiple parts according to a specified delimiter.

Word Meaning:splitsplit means "separate, split", indicating dividing a string into multiple parts.


Basic Syntax and Parameters

str.split()It is a string accessor method of Series, so you need to first have a Series containing strings, and then use.strthe accessor to call it.

Syntax Format

Series.str.split(pat=None, n=-1, expand=False)

Parameter Description

Parameter Type Required? Description Default Value
pat str Optional Delimiter, can be a string or a regular expression. Defaults to whitespace. None (whitespace)
n int Optional Number of splits. -1 means no limit, split into as many parts as possible. -1
expand bool Optional Whether to expand the result into a DataFrame. Defaults to False (returns a Series of lists). False

Function Description

  • Return Value: By default, returns a Series containing lists, where each element in each list is a substring after splitting. Whenexpand=Trueexpand=True, it returns a DataFrame.
  • Effect: Splits each string into multiple parts according to the specified delimiter.
  • Note: By default, whitespace is used as the delimiter; you can specify a custom delimiter.

Examples

Let us, through a series of examples from simple to complex, thoroughly masterstr.split()the usage.

Example 1: Basic Usage - Splitting by Whitespace

Example

import pandas as pd

# Create a Series containing sentences
s = pd.Series(['hello world', 'example python', 'pandas data analysis'])

# Split by whitespace (default behavior)
result = s.str.split()

print("Original Series:")
print(s)
print("nResult after splitting:")
print(result)

Output:

原始 Series:
0         hello world
1      example python
2    pandas data analysis
dtype: object

拆分后的结果:
0         [hello, world]
1    [example, python]
2    [pandas, data, analysis]

Code explanation:

  1. s.str.split()By default, strings are split by whitespace (spaces, tabs, etc.).
  2. The return value is a Series containing lists, each list contains the substrings after splitting.
  3. 'pandas data analysis' is split into three parts.

Example 2: Splitting by a Specified Delimiter

You can usepatthe parameter to specify a custom delimiter.

Example

import pandas as pd

# Create a Series containing comma-separated values
s = pd.Series(['apple,banana,orange', 'dog,cat,bird', 'red,green,blue'])

# Split by comma
result = s.str.split(',')

print("Original Series:")
print(s)
print("nResult after splitting by comma:")
print(result)

Output:

原始 Series:
0    apple,banana,orange
1       dog,cat,bird
2       red,green,blue
dtype: object

拆分后的结果:
0    [apple, banana, orange]
1       [dog, cat, bird]
2       [red, green, blue]

Code explanation:

  • s.str.split(',')Split each string by comma.
  • This is a common method for processing CSV data or comma-separated values.

Example 3: Limiting the Number of Splits

Using thenparameter, you can limit the number of splits.

Example

import pandas as pd

s = pd.Series(['a,b,c,d,e', '1,2,3,4,5'])

# Limit to only 2 parts
result_2 = s.str.split(',', n=2)
# Limit to only 1 part (i.e., split only once)
result_1 = s.str.split(',', n=1)

print("Original Series:")
print(s)
print("nLimit to 2 parts:")
print(result_2)
print("nLimit to 1 part:")
print(result_1)

Output:

原始 Series:
0       a,b,c,d,e
1       1,2,3,4,5
dtype: object

限制拆分为 2 部分:
0       [a, b, c,d,e]
1       [1, 2, 3,4,5]

限制拆分为 1 部分:
0            [a, b,c,d,e]
1            [1, 2,3,4,5]

Code explanation:

  • n=2Means splitting into at most 2 parts, with the last part containing all the remaining content.
  • n=1Means splitting only at the first delimiter, resulting in 2 parts.

Example 4: Expanding to a DataFrame

Whenexpand=TrueWhen expand=True, the split result is expanded into a DataFrame.

Example

import pandas as pd

# Create a Series containing comma-separated values
s = pd.Series(['apple,banana,orange', 'dog,cat,bird', 'red,green,blue'])

# Expand to DataFrame
result = s.str.split(',', expand=True)

print("Original Series:")
print(s)
print("nExpanded to DataFrame:")
print(result)
print("nDataFrame type:", type(result))

Output:

原始 Series:
0    apple,banana,orange
1       dog,cat,bird
2       red,green,blue
dtype: object

展开为 DataFrame:
        0       1       2
0   apple   banana  orange
1     dog     cat    bird
2     red   green    blue

DataFrame 类型: <class 'pandas.core.frame.DataFrame'>

Code explanation:

  • expand=TrueExpand the split result into a DataFrame.
  • Each column represents a position after splitting.
  • This is very useful when handling structured data.

Example 5: Splitting with Regular Expressions

split()It also supports using regular expressions as delimiters.

Example

import pandas as pd

# Create a Series containing different delimiters
s = pd.Series(['hello-world', 'example_python', 'pandas#tutorial'])

# Split by regular expression (match any one of -, _, #)
result = s.str.split(r'[-_#]')

print("Original Series:")
print(s)
print("nSplit by regular expression:")
print(result)

Output:

原始 Series:
0     hello-world
1   example_python
2    pandas#tutorial
dtype: object

按正则表达式拆分:
0       [hello, world]
1       [example, python]
2    [pandas, tutorial]

Code explanation:

  • r'[-_#]'It is a regular expression that matches any one of '-', '_', or '#'.
  • It can handle multiple different delimiters at the same time.

Example 6: Handling Real Data

In practical applications,split()it is often used in combination with other functions.

Example

import pandas as pd

# Simulate data extracted from logs
logs = pd.Series([
    '2024-01-01 10:30:45 ERROR Connection failed',
    '2024-01-01 10:31:12 INFO User logged in',
    '2024-01-01 10:32:00 WARNING Memory usage high'
])

# Split the log content
log_parts = logs.str.split(r's+', expand=True)

print("Original log:")
print(logs)
print("nSplit DataFrame:")
print(log_parts)

# Rename columns
log_parts.columns = ['timestamp', 'level', 'message']
print("nResult after renaming:")
print(log_parts)

Output:

原始日志:
0    2024-01-01 10:30:45 ERROR Connection failed
1    2024-01-01 10:31:12 INFO User logged in
2    2024-01-01 10:32:00 WARNING Memory usage high
dtype: object

拆分后的 DataFrame:
                  0     1     2
0  2024-01-01 10:30:45  ERROR  Connection failed
1  2024-01-01 10:31:12   INFO  User logged in
2  2024-01-01 10:32:00  WARNING  Memory usage high

重命名后的结果:
               timestamp    level                message
0  2024-01-01 10:30:45    ERROR    Connection failed
1  2024-01-01 10:31:12     INFO      User logged in
2  2024-01-01 10:32:00  WARNING  Memory usage high

Code explanation:

  • r's+'Matches one or more whitespace characters, splitting the log into multiple parts.
  • expand=TrueExpand the split result into a DataFrame.
  • By renaming the columns, you can conveniently access and process each part.

Notes

  • str.split()By default, whitespace is used as the delimiter.
  • Whenexpand=FalseWhen expand=False, it returns a Series containing lists.
  • Whenexpand=TrueWhen expand=True, it returns a DataFrame, and the number of columns depends on the longest split result.
  • If a string does not contain the delimiter, the returned list has only one element (the original string).
  • Regular expressions can be used as delimiters, providing a more flexible way to split.
  • If the Series contains NaN values, NaN is returned.

Pandas 常用函数Common Pandas Functions

Other Extensions