Pandas Data Types (dtype and Type Conversion)

Pandas provides a rich data type system. Correctly understanding and using data types is the foundation of efficient data analysis. This section details Pandas' data type system, type inference, and type conversion methods.


Pandas Data Type Overview

dtype Description Python Type Example
int64 64-bit Integer int 1, 2, 100
float64 64-bit Float float 1.5, 3.14
object String or Mixed Type str "hello"
bool Boolean bool True, False
datetime64[ns] Datetime datetime 2024-01-01
timedelta64[ns] Timedelta timedelta 1 days
category Categorical Type - Finite Set

Example

import pandas as pd
import numpy as np

# Create a DataFrame with various data types
df = pd.DataFrame({
    "Integer": [1, 2, 3],
    "Float": [1.5, 2.5, 3.5],
    "String": ["a", "b", "c"],
    "Boolean": [True, False, True],
    "Date": pd.date_range("2024-01-01", periods=3)
})

print("Column data types:")
print(df.dtypes)

Type Inference and Specification

Automatic Type Inference

Example

import pandas as pd

# Type inference when reading CSV
# pandas attempts to infer the most suitable type for each column
df = pd.read_csv("data.csv")

# Or use the dtype parameter to explicitly specify types
df = pd.read_csv("data.csv", dtype={
    "Age": "int32",     # Specify as 32-bit integer to save memory
    "Salary": "float32",   # Specify as 32-bit float
    "Name": "string"     # Use PyArrow string
})

Specifying Types at Creation

Example

import pandas as pd
import numpy as np

# Specify data type to create Series
s = pd.Series([1, 2, 3], dtype="int8")  # Use a smaller integer type
print(f"int8 type: {s.dtype}")

s = pd.Series([1.5, 2.5, 3.5], dtype="float32")  # Use float32
print(f"float32 type: {s.dtype}")

# Use numpy type
s = pd.Series([1, 2, 3], dtype=np.int8)
print(f"np.int8 type: {s.dtype}")

Type Conversion

Converting with astype

Example

import pandas as pd
import numpy as np

# Create sample data
df = pd.DataFrame({
    "Integer": [1, 2, 3],
    "Float": [1.5, 2.5, 3.5],
    "String": ["1", "2", "3"],
    "Boolean": [1, 0, 1]
})

print("Original types:")
print(df.dtypes)
print()

# Convert to string
df["Integer_str"] = df["Integer"].astype(str)
print("After converting to string:")
print(df.dtypes)

# Convert string to numeric
df["String_int"] = df["String"].astype(int)
print("\n"String to integer:")
print(df.dtypes)

# Convert numeric to boolean (non-zero is True)
df["Boolean_int"] = df["Boolean"].astype(bool)
print("\n"Integer to boolean:")
print(df.dtypes)

# Convert float to integer (truncation)
df["Float_int"] = df["Float"].astype(int)
print("\n"Float to integer:")
print(df)

Converting with pd.to_numeric

Example

import pandas as pd
import numpy as np

# Handle numeric strings with special characters
s = pd.Series(["$1,000", "$2,500", "$3,200"])

# Clean and convert
s_cleaned = s.str.replace("$", "", regex=False).str.replace(",", "", regex=False)
s_numeric = pd.to_numeric(s_cleaned)
print("String to numeric:")
print(s_numeric)
print(f"Type: {s_numeric.dtype}")
print()

# Handle missing values
s_with_na = pd.Series(["1", "2", "NA", "4"])
s_numeric = pd.to_numeric(s_with_na, errors="coerce")  # Convert invalid values to NaN
print("Handling missing values:")
print(s_numeric)

Converting Dates with pd.to_datetime

Example

import pandas as pd

# Convert various date formats
dates = ["2024-01-01", "2024/01/02", "01/03/2024", "20240104"]

# Convert dates
dt = pd.to_datetime(dates, errors="coerce")
print("Date conversion result:")
print(dt)
print(f"Type: {dt.dtype}")
print()

# Specify date format
dates2 = ["20240101", "20240102", "20240103"]
dt2 = pd.to_datetime(dates2, format="%Y%m%d")
print("Conversion with specified format:")
print(dt2)

Memory Optimization

Choosing appropriate data types can significantly reduce memory usage.

Choosing Integer Types

Example

import pandas as pd
import numpy as np

# Create a Series with a large amount of data
s = pd.Series(np.random.randint(0, 100, 1000000))

# Default int64
print(f"int64 memory: {s.dtype} -> {s.memory_usage(deep=True) / 1024 / 1024:.2f} MB")

# Convert to a smaller type
s_int8 = s.astype("int8")
print(f"int8 memory: {s_int8.dtype} -> {s_int8.memory_usage(deep=True) / 1024 / 1024:.2f} MB")

# Choose an appropriate type based on the data range
# np.iinfo to view integer range
print(f"\n"int8 range: {np.iinfo('int8').min} to {np.iinfo('int8').max}")
print(f"int16 range: {np.iinfo('int16').min} to {np.iinfo('int16').max}")
print(f"int32 range: {np.iinfo('int32').min} to {np.iinfo('int32').max}")

Using the category Type

Example

import pandas as pd
import numpy as np

# Create a Series with many duplicate values
s = pd.Series(np.random.choice(["Beijing", "Shanghai", "Guangzhou", "Shenzhen"], 1000000))

# String type
print(f"object type memory: {s.dtype} -> {s.memory_usage(deep=True) / 1024 / 1024:.2f} MB")

# Convert to category type
s_cat = s.astype("category")
print(f"category type memory: {s_cat.dtype} -> {s_cat.memory_usage(deep=True) / 1024 / 1024:.2f} MB")

# Disadvantage: some operations may become slower
print("\n"category info:")
print(s_cat.cat.categories)

Type Checking and Validation

Example

import pandas as pd
import numpy as np

df = pd.DataFrame({
    "Integer": [1, 2, 3],
    "Float": [1.5, 2.5, 3.5],
    "String": ["a", "b", "c"]
})

# Check data types
print("Check if it is a numeric type:")
print(df.dtypes)

# Use is_integer_dtype
print(f"\n"Integer column: {pd.api.types.is_integer_dtype(df['Integer'])}")
print(f"Float column: {pd.api.types.is_float_dtype(df['Float'])}")
print(f"String column: {pd.api.types.is_object_dtype(df['String'])}")

# Check if it can be converted to numeric
print(f"\n"Can be converted to numeric: {pd.api.types.is_numeric_dtype(df['Integer'])}")
print(f"Is datetime: {pd.api.types.is_datetime64_any_dtype(pd.Series(pd.date_range('2024', periods=3)))}")

Common Issues

1. Float values appear in integer columns

When mixed into a DataFrame, it automatically converts to float64. UseNullable IntegerThe type can preserve integer null values.

2. Inconsistent date formats

Usepd.to_datetimeand specifyformatparameter to handle non-standard formats.

3. Type error when reading CSV

Usedtypethe parameter to explicitly specify the type, or useconvertersto convert.

For large datasets, prioritize using smaller data types (such as int8, int16, float32) and category types to reduce memory usage.

Other extensions