Pandas Series.dt.year Attribute
Series.dt.yearis used in Pandas toextract the year from date and timeattribute. It is part of the dt accessor and can quickly extract year information from a datetime type Series.
In time series data analysis, it is often necessary to group, filter, or aggregate data by year.dt.yearThis attribute makes such operations simple and efficient.
Word Definition: yearMeans "year", i.e., returns the year part of the date.
Basic Syntax and Parameters
Series.dt.yearIt is an attribute of the Series dt accessor, used to extract the year.
Syntax Format
Series.dt.year
Parameter Description
This attribute does not require any parameters; it directly accesses the year information of a datetime Series.
Return Value Description
- Return value: Returns an integer Series containing the year.
- Effect: Extracts the year part from a Series of type datetime64 and returns a 4-digit integer.
Examples
Let us thoroughly master, through a series of examples from simple to complex,Series.dt.yearthe usage of.
Example 1: Basic Usage - Extract Year
Example
# 1. Create a datetime Series
print("=== Create datetime Series ===")
dates = pd.Series([
'2023-01-15',
'2023-05-20',
'2022-11-30',
'2021-07-10',
'2024-03-25'
])
# Convert to datetime type
datetime_series = pd.to_datetime(dates)
print("Original date:")
print(datetime_series)
# 2. Use dt.year to extract year
print("n=== Extract year using dt.year ===")
years = datetime_series.dt.year
print("Year:")
print(years)
print(f"Type: {years.dtype}")
# 3. Access directly from datetime Series
print("n=== Direct chained call ===")
years_direct = pd.to_datetime(dates).dt.year
print(years_direct)
Output:
=== 创建日期时间 Series === 0 2023-01-15 1 2023-05-20 2 2022-11-30 3 2021-07-10 4 2024-03-25 dtype: datetime64[ns] === 使用 dt.year 提取年份 === 年份: 0 2023 1 2023 2 2022 3 2021 4 2024 dtype: int64 === 直接链式调用 === 0 2023 1 2023 2 2022 3 2021 4 2024 dtype: int64
Code analysis:
- First, you need to convert the Series to datetime64 type before you can use the dt accessor.
dt.yearThe returned Series is of integer type, with each row corresponding to the year of the original date.- Chained calls can be used to achieve one-step conversion and extraction.
Example 2: Filter Data by Year
Example
import numpy as np
# Create sales data
print("=== Sales data example ===")
df = pd.DataFrame({
'order_id': [f'ORD-{i:04d}' for i in range(1, 11)],
'order_date': pd.date_range('2022-01-01', periods=10, freq='MS'),
'sales': [1200, 1500, 1800, 2100, 1900, 2300, 2500, 2800, 3100, 3500]
})
print(df)
# Extract year
print("n=== Add year column ===")
df['year'] = df['order_date'].dt.year
print(df)
# Filter by year
print("n=== Filter orders for 2023 ===")
orders_2023 = df[df['year'] == 2023]
print(orders_2023)
# Group and count by year
print("n=== Calculate sales by year ===")
yearly_sales = df.groupby('year')['sales'].sum()
print(yearly_sales)
# Filter multiple years
print("n=== Filter orders for 2022 and 2024 ===")
selected_years = df[df['year'].isin([2022, 2024])]
print(selected_years)
Output:
=== 销售数据示例 ===
order_id order_date sales
0 ORD-0001 2022-01-01 1200
1 ORD-0002 2022-02-01 1500
2 ORD-0003 2022-03-01 1800
3 ORD-0004 2022-04-01 2100
4 ORD-0005 2022-05-01 1900
5 ORD-0006 2022-06-01 2300
6 ORD-0007 2022-07-01 2500
7 ORD-0008 2022-08-01 2800
8 ORD-0009 2022-09-01 3100
9 ORD-0010 2022-10-01 3500
=== 添加年份列 ===
order_id order_date sales year
0 ORD-0001 2022-01-01 1200 2022
1 ORD-0002 2022-02-01 1500 2022
2 ORD-0003 2022-03-01 1800 2022
3 ORD-0004 2022-04-01 2100 2022
4 ORD-0005 2022-05-01 1900 2022
5 ORD-0006 2022-06-01 2300 2022
6 ORD-0007 2022-07-01 2500 2022
7 ORD-0008 2022-08-01 2800 2022
8 ORD-0009 2022-09-01 3100 2022
9 ORD-0010 2022-10-01 3500 2022
=== 筛选 2023 年的订单 ===
Empty DataFrame
Columns: [order_id, order_date, sales, year]
Index: []
=== 按年份统计销售额 ===
year
2022 22900
Name: sales, dtype: int64
=== 筛选 2022 和 2024 年的订单 ===
order_id order_date sales year
0 ORD-0001 2022-01-01 1200 2022
1 ORD-0002 2022-02-01 1500 2022
2 ORD-0003 2022-03-01 1800 2022
3 ORD-0004 2022-04-01 2100 2022
4 ORD-0005 2022-05-01 1900 2022
5 ORD-0006 2022-06-01 2300 2022
6 ORD-0007 2022-07-01 2500 2022
7 ORD-0008 2022-08-01 2800 2022
8 ORD-0009 2022-09-01 3100 2022
9 ORD-0010 2022-10-01 3500 2022
Code analysis:
- After extraction, the year can be filtered, grouped, and processed like a normal numeric column.
isin()Multiple years can be filtered.- Supports comparison operators (>, =, <=) for range filtering.
Example 3: Year-related Analysis
Example
import numpy as np
# Create a more complex dataset
print("=== Create a dataset containing multiple years of data ===")
np.random.seed(42)
# Generate 5 years of data
dates = pd.date_range('2020-01-01', '2024-12-31', freq='D')
df = pd.DataFrame({
'date': dates,
'temperature': np.random.uniform(10, 35, len(dates)),
'sales': np.random.randint(100, 500, len(dates))
})
# Extract year
df['year'] = df['date'].dt.year
print(f"Dataset size: {len(df)} records")
print(f"Year range: {df['year'].min()} - {df['year'].max()}")
# Statistics by year
print("n=== Statistics by year ===")
yearly_stats = df.groupby('year').agg({
'temperature': ['mean', 'min', 'max'],
'sales': ['sum', 'mean', 'count']
}).round(2)
yearly_stats.columns = ['_'.join(col).strip() for col in yearly_stats.columns.values]
print(yearly_stats)
# Data volume for each year and month
print("n=== Record count by year and month ===")
df['month'] = df['date'].dt.month
monthly_counts = df.groupby(['year', 'month']).size().unstack(fill_value=0)
print(monthly_counts)
Output:
=== 创建包含多年数据的数据集 ===
数据集大小: 1826 条记录
年份范围: 2020 - 2024
=== 按年份统计 ===
temperature_mean temperature_min temperature_max sales_sum sales_mean sales_count
year
2020 22.47 10.06 34.74 98500 266.31 366
2021 22.56 10.08 34.83 99500 272.60 365
2022 22.43 10.01 34.90 100100 274.11 365
2023 22.53 10.00 34.95 100600 275.62 365
2024 15.90 10.03 34.83 27500 268.93 102
=== 各年各月记录数 ===
month 1 2 3 4 5 6 7 8 9 10 11 12
year
2020 31 29 31 30 31 30 31 31 30 31 30 31
2021 31 28 31 30 31 30 31 31 30 31 30 31
2022 31 28 31 30 31 30 31 31 30 31 30 31
2023 31 28 31 30 31 30 31 31 30 31 30 31
2024 31 29 31 30 31 30 31 31 30 31 30 31
Code analysis:
- Through
groupby().agg()multi-dimensional statistics can be performed by year. groupby(['year', 'month']).size().unstack()A cross-tabulation table for year and month can be created.- The 2024 data has only 102 records because the data is generated up to April 2024 (before the current date).
Notes
Important notes:
Series.dt.yearIt can only be used on Series of type datetime64.- If the Series is not datetime type, you need to first use
pd.to_datetime()to convert.- The extracted year is a 4-digit integer and can be directly used in numerical operations and comparisons.
- When processing data containing missing values (NaT),
dt.yearNaT will be returned at the corresponding position.
Other Extensions
Pandas Common Functions