Pandas df.astype() Function

Pandas Common Functions


df.astype()It is a function in Pandas for converting the data types of a DataFrame or Series.

During data processing, mismatched data types are a common problem. For example, numbers stored as strings, dates stored as text, etc.astype()It allows you to explicitly convert data to the required type, such as integer, float, string, date, etc., ensuring that the data can be correctly calculated and analyzed.


Basic Syntax and Parameters

astype()It is both a member function of DataFrame and a member function of Series, invoked via the dot operator.to call.

Syntax Format

DataFrame.astype(dtype, copy=True, errors='raise')
Series.astype(dtype, copy=True, errors='raise')

Parameter Description

Parameter Type Required? Description Default Value
dtype Python dtype or numpy dtype Required Target data type. It can be a Python type (e.g.,int、str), a NumPy type (e.g.,np.int64、np.float32), or a Pandas type (e.g.,'int64'、'float64'、'category')。 None
copy bool Optional If it isTrue, always return a new object; if it isFalse, modify the original object when possible. True
errors str Optional Controls error handling.'raise'Indicates that an exception is raised when conversion fails;'ignore'Indicates that the original data is returned when conversion fails, without raising an exception. 'raise'

Return Value Description

  • Returns a new DataFrame or Series in which the data types of all specified columns have been converted to the target type.

Examples

Let us thoroughly master through a series of examplesastype()the usage of.

Example 1: Convert the Data Type of a Single Column

Convert a column of a Series or DataFrame to the specified type.

Example

import pandas as pd

# Create a DataFrame where values are stored as strings
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
    'age': ['25', '30', '35', '28'],  # Store as string
    'Salary': ['5000', '6000', '5500', '7000']  # Store as string
}
df = pd.DataFrame(data)

print("Original Data Type:")
print(df.dtypes)
print("=" * 50)

# Convert the "Age" column to integer type
df['age'] = df['age'].astype(int)

print(Data type of the age column after conversion:)
print(df.dtypes)
print("=" * 50)

# Convert the 'Salary' column to float type
df['Salary'] = df['Salary'].astype(float)

print(Data type of the salary column after conversion:)
print(df.dtypes)
print("=" * 50)

print("Converted Data:")
print(df)

Expected output:

原始数据类型:
姓名     object
年龄    object  # 字符串类型
薪资    object  # 字符串类型
==================================================
转换年龄列后的数据类型:
姓名      object
年龄     int64  # 整数类型
薪资    object
==================================================
转换薪资列后的数据类型:
姓名      object
年龄     int64
薪资    float64  # 浮点数类型
==================================================
转换后的数据:
    姓名  年龄    薪资
0  张三   25  5000.0
1  李四   30  6000.0
2  王五   35  5500.0
3  赵六   28  7000.0

Code Explanation:

  1. In the original data, 'age' and 'salary' are both string types (object)。
  2. Usedf['列名'].astype(int)to convert strings to integers.
  3. After conversion, numerical calculations can be performed, such as sum, average, etc.

Example 2: Convert Data Types of an Entire DataFrame

You can specify data types for the entire DataFrame at once.

Example

import pandas as pd

# Create a DataFrame
data = {
    'A': [1, 2, 3, 4],
    'B': [1.5, 2.5, 3.5, 4.5],
    'C': ['a', 'b', 'c', 'd']
}
df = pd.DataFrame(data)

print("Original Data Type:")
print(df.dtypes)
print("=" * 50)

# Convert column A to float64 and column B to int64
df_converted = df.astype({'A': 'float64', 'B': 'int64'})

print(Converted data types:)
print(df_converted.dtypes)
print("=" * 50)

print("Converted Data:")
print(df_converted)

Expected output:

原始数据类型:
A      int64
B    float64
C     object
==================================================
转换后的数据类型:
A    float64
B      int64  # 小数部分被截断
C     object
==================================================
转换后的数据:
     A  B  C
0  1.0  2  a
1  2.0  2  b
2  3.0  3  c
3  4.0  4  d

Code Explanation:

  • Use a dictionary{'Column Name': 'Type'}You can specify the conversion types of multiple columns at once.
  • When converting the float values in column B to integers, the decimal part is truncated (2.5 becomes 2).

Example 3: Convert to Categorical Data Type

Categorical data type (category) can save memory, especially suitable for columns with limited values.

Example

import pandas as pd

# Create a DataFrame with many duplicate values
data = {
    'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen', 'Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen'] * 1000,
    Number: range(8000)
}
df = pd.DataFrame(data)

print("Original Data Type:")
print(df.dtypes)
print(fMemory usage: {df.memory_usage(deep=True).sum() / 1024:.2f} KB)
print("=" * 50)

# Convert the "City" column to categorical type
df['city'] = df['city'].astype('category')

print(Converted data types:)
print(df.dtypes)
print(fMemory usage: {df.memory_usage(deep=True).sum() / 1024:.2f} KB)

Expected output:

原始数据类型:
城市    object
编号    int64
内存使用: ~456 KB
==================================================
转换后的数据类型:
城市    category
编号     int64
内存使用: ~120 KB  # 内存使用大幅减少

Code Explanation:

  • The categorical data type assigns an integer identifier to each unique value, saving a lot of memory compared to storing full strings.
  • The original data has 8000 rows and uses about 456 KB of memory; after conversion it is about 120 KB, saving about 75% of memory.
  • This method is especially suitable for categorical variables with few values.

Example 4: Convert to Datetime Type

Convert strings to datetime type so that date-related operations can be performed.

Example

import pandas as pd

# Create a DataFrame with date strings
data = {
    Date: ['2024-01-01', '2024-01-02', '2024-01-03', '2024-01-04'],
    'Sales': [100, 150, 120, 180]
}
df = pd.DataFrame(data)

print("Original Data Type:")
print(df.dtypes)
print("=" * 50)

# Convert the "date" column to datetime type
df[Date] = pd.to_datetime(df[Date])

print(Converted data types:)
print(df.dtypes)
print("=" * 50)

print("Converted Data:")
print(df)

# Now you can perform date-related operations
print(nExtract month:)
print(df[Date].dt.month)

Expected output:

原始数据类型:
日期    object  # 字符串类型
销量     int64
==================================================
转换后的数据类型:
日期    datetime64[ns]  # 日期时间类型
销量     int64
==================================================
转换后的数据:
        日期   销量
0 2024-01-01  100
1 2024-01-02  150
2 2024-01-03  120
3 2024-01-04  180

提取月份:
0    1
1    1
2    1
3    1
Name: 日期, dtype: int64

Code Explanation:

  • Usepd.to_datetime()to convert strings to datetime type.
  • After conversion, you can use.dtaccessor to perform date operations, such as extracting month, day of the week, etc.
  • Note: Although you can also useastype('datetime64[ns]'), butpd.to_datetime()is more powerful and flexible.

Example 5: Handling Conversion Errors

Useerrors='ignore'parameter to handle conversion failures.

Example

import pandas as pd

# Create a Series containing values that cannot be converted to integers
s = pd.Series(['1', '2', 'hello', '4', 'world'])

print(“Raw data:”)
print(s)
print("=" * 50)

# Try conversion. Using errors='raise' (default) will raise an exception.
try:
    s_converted = s.astype(int)
except ValueError as e:
    print(fConversion failed: {e})
print("=" * 50)

# Use errors='ignore', return the original data when conversion fails
s_converted = s.astype(int, errors='ignore')

print("Data converted using errors='ignore':")
print(s_converted)
print("=" * 50)

# Use errors='coerce' to fill with NaN when conversion fails
s_converted2 = s.astype(float, errors='coerce')

print("Data converted using errors='coerce':")
print(s_converted2)

Expected output:

原始数据:
0       1
1       2
2    hello  # 无法转换为整数
3       4
4    world  # 无法转换为整数
==================================================
转换失败: invalid literal for int() with base 10: 'hello'
==================================================
使用 errors='ignore' 转换后的数据:
0       1
1       2
2    hello  # 保持原值
3       4
4    world  # 保持原值
==================================================
使用 errors='coerce' 转换后的数据:
0    1.0
1    2.0
2    NaN  # 无法转换,设为 NaN
3    4.0
4    NaN  # 无法转换,设为 NaN

Code Explanation:

  • When a string contains a value that cannot be converted to a number, by default it raisesValueErroran exception.
  • errors='ignore'will keep the original data unchanged.
  • errors='coerce'will set the non-convertible value toNaN, which is a common method for handling dirty data.

Example 6: Convert to Boolean Type

Convert data to boolean type, commonly used in conditional filtering and logical operations.

Example

import pandas as pd

# Create a DataFrame containing 0 and 1
data = {
    'Is Member': [0, 1, 1, 0, 1],
    'Is Active': ['Yes', 'No', 'Yes', 'Yes', 'No']
}
df = pd.DataFrame(data)

print("Raw data:")
print(df)
print("Raw data types:")
print(df.dtypes)
print("=" * 50)

# Convert 0/1 to boolean type
df['Is Member'] = df['Is Member'].astype(bool)

print("Data types after converting the Is Member column:")
print(df.dtypes)
print(df)
print("=" * 50)

# Convert "Yes"/"No" to boolean type
df['Is Active'] = df['Is Active'].map({'Yes': True, 'No': False})

print("Data after converting the Is Active column:")
print(df)
print(df.dtypes)

Expected running results:

原始数据:
  是否会员  是否激活
0       0       是
1       1       否
2       1       是
3       0       是
4       1       否
==================================================
转换是否会员列后的数据类型:
是否会员      bool
是否激活    object
==================================================
转换是否激活列后的数据:
  是否会员  是否激活
0    True    True
1   True   False
2    True    True
3   False   True
4    True   False

Code analysis:

  • 0 can be converted toFalse, 1 converted toTrue。
  • For strings like "Yes"/"No", you need to usemap()to convert.
  • After conversion, boolean indexing can be used for fast filtering.

Notes

  • astype()By default, a new object is returned, and the original data is not modified. If you want to modify in place, you can usecopy=Falseparameter, but this may affect the original data.
  • When converting floating-point numbers to integers, the decimal part is truncated, not rounded.
  • When converting strings to numbers, make sure the string is in a valid numeric format, otherwise an exception will be thrown.
  • Usingerrors='coerce'is the recommended method for handling dirty data, and can convert invalid values toNaN, then usefillna()to fill.
  • For columns with many duplicate values, converting to a categorical type can significantly reduce memory usage.

Pandas Common Functions

Other Extensions