Pandas Series.str_len() Function

Pandas 常用函数Pandas Common Functions


Series.str.len()It is a function in Pandas used to calculate string length.

In data processing, we often need to know string length information, such as validating input length, filtering text of specific length, analyzing text length distribution, etc.len()The function can return the number of characters in each string element.

Term Explanation:lenIt is an abbreviation of "length", indicating the number of characters in the returned string.


Basic Syntax and Parameters

str.len()It is a string accessor method of Series, so you need to first have a Series containing strings, and then use.strthe accessor to call it.

Syntax Format

Series.str.len()

Parameter Description

  • Parameter: No parameters. This function does not require any parameters; just call it directly.

Function Description

  • Return Value: Returns an integer Series, representing the number of characters in each string.
  • Effect: Calculates the number of characters in each string element of the Series (including spaces and special characters).
  • Note: For non-string elements, it returns the length of its string representation.

Examples

Let's, through a series of examples from simple to complex, thoroughly masterstr.len()its usage.

Example 1: Basic Usage - Calculating String Length

Example

import pandas as pd

# Create a Series containing strings of different lengths
s = pd.Series(['hello', 'example', 'python', 'ai', 'data'])

# Use str.len() to calculate the length of each string
result = s.str.len()

print("Original Series:")
print(s)
print("nString length:")
print(result)

Output result:

原始 Series:
0      hello
1     example
2     python
3        ai
4       data
dtype: object

字符串长度:
0    5
1    6
2    6
3    2
4    4

Code analysis:

  1. s.str.len()Calculate the number of characters in each string.
  2. 'hello' has 5 characters, returns 5.
  3. 'example' has 6 characters, returns 6.
  4. An empty string returns length 0.

Example 2: String Length Including Spaces

len()It calculates all characters, including spaces.

Example

import pandas as pd

# Create a Series containing spaces
s = pd.Series(['hello world', 'example python', 'a b c', '  space  '])

# Calculate string length
result = s.str.len()

print("Original Series:")
print(s)
print("nString length (including spaces):")
print(result)

Output result:

原始 Series:
0     hello world
1   example python
2           a b c
3          space
dtype: object

字符串长度(包括空格):
0    11
1    13
2          5
3          8

Code analysis:

  • 'hello world' has 11 characters (including the space in the middle).
  • ' space ' has 8 characters (including two spaces before and after).
  • Spaces are counted as characters.

Example 3: Filtering Strings of Specific Length

len()It is often combined with boolean indexing to filter data.

Example

import pandas as pd

# Create a product name Series
products = pd.Series(['iPhone', 'Samsung Galaxy', 'MacBook Pro', 'iPad', 'Dell XPS'])

# Filter products with length greater than 8
long_names = products[products.str.len() > 8]
# Filter products with length less than or equal to 5
short_names = products[products.str.len() <= 5]

print("All products:")
print(products)
print("nLength greater than8products:")
print(long_names)
print("nLength less than or equal to5products:")
print(short_names)

Output result:

所有产品:
0            iPhone
1    Samsung Galaxy
2       MacBook Pro
3              iPad
4            Dell XPS
dtype: object

长度大于 8 的产品:
1    Samsung Galaxy
2       MacBook Pro
4        Dell XPS
dtype: object

长度小于等于 5 的产品:
0       iPhone
3         iPad
dtype: object

Code analysis:

  • products.str.len() > 8Returns a boolean Series, indicating whether the length is greater than 8.
  • Use boolean indexing to filter out product names that meet the condition.
  • This is a common operation in data filtering.

Example 4: Counting Length Distribution

len()It can be used for statistical analysis.

Example

import pandas as pd

# Create a Series containing sentence lengths
sentences = pd.Series([
    'Hello',
    'Hello world',
    'Pandas is powerful',
    'Data science with machine learning',
    'AI'
])

# Calculate the length of each sentence
lengths = sentences.str.len()

print("Sentence list:")
print(sentences)
print("nCharacter length of each sentence:")
print(lengths)
print("nStatistics:")
print(f"Shortest length: {lengths.min()}")
print(f"Longest length: {lengths.max()}")
print(f"Average length: {lengths.mean():.2f}")
print(f"Total length: {lengths.sum()}")

Output result:

句子列表:
0                            Hello
1                       Hello world
2                  Pandas is powerful
3    Data science with machine learning
4                             AI
dtype: object

每个句子的字符长度:
0      5
1     11
2     18
3     29
4      2
dtype: int64

统计信息:
最短长度: 2
最长长度: 29
平均长度: 13.00
总长度: 65

Code analysis:

  • lengths.min()Returns the shortest string length.
  • lengths.max()Returns the longest string length.
  • lengths.mean()Returns the average length.
  • lengths.sum()Returns the total number of characters.

Example 5: Handling Mixed-Type Data

len()It also calculates the string representation length of non-string elements such as numbers.

Example

import pandas as pd
import numpy as np

# Create a Series containing mixed types
s = pd.Series(['hello', 12345, 'example', np.nan, 'py'])

# Calculate length
result = s.str.len()

print("Original Series:")
print(s)
print("nString length:")
print(result)

Output result:

原始 Series:
0       hello
1       12345
2      example
3        NaN
4         py
dtype: object

字符串长度:
0       5.0
1       5.0
2       6.0
3       NaN
4       2.0

Code analysis:

  • The number 12345 is converted to the string '12345', with a length of 5.
  • NaN values return NaN (no error is raised).
  • It returns a float type because NaN is float.

Example 6: Combining with split to Count Words

You canlen()andsplit()combine them to count the number of words.

Example

import pandas as pd

# Create a Series containing sentences
sentences = pd.Series([
    'hello world',
    'example python tutorial',
    'pandas data analysis',
    'machine learning'
])

# Split first, then calculate the length to get the word count
word_counts = sentences.str.split().str.len()

print("Sentence list:")
print(sentences)
print("nWord count:")
print(word_counts)

Output result:

句子列表:
0           hello world
1    example python tutorial
2    pandas data analysis
3    machine learning
dtype: object

单词数量:
0    2
1    3
2    3
3    2
dtype: int64

Code analysis:

  • str.split()Split the sentence into a list of words.
  • .str.len()Calculate the length of the list (i.e., the number of words).
  • This is very useful when counting the number of words in a text.

Notes

  • str.len()It counts the number of characters, not bytes.
  • For Chinese characters, each character counts as one character.
  • Special characters such as spaces and tabs are also counted.
  • If the Series contains NaN values, NaN is returned (instead of an error).
  • For non-string elements (such as numbers), they are first converted to strings and then the length is calculated.
  • This function returns a new Series and does not modify the original data.

Pandas 常用函数Pandas Common Functions

Other Extensions