Pandas Time Series Analysis

Time series analysis is an important part of data analysis. Pandas provides rich functionality to process and analyze time series data, including resampling, rolling calculations, moving averages, etc.


Basic Time Series Operations

Creating Time Series

Example

import pandas as pd
import numpy as np

# Create time series data
np.random.seed(42)
dates = pd.date_range("2024-01-01", periods=100, freq="D")

ts = pd.Series(
    np.random.randn(100).cumsum() + 100,
    index=dates
)

print("Time series data (first 10 items):")
print(ts.head(10))
print()

# View information
print(f"Index type: {type(ts.index)}")
print(f"Start time: {ts.index.min()}")
print(f"End time: {ts.index.max()}")
print(f"Time span: {ts.index.max() - ts.index.min()}")

Set Date as Index

Example

import pandas as pd
import numpy as np

# Create DataFrame and set date index
df = pd.DataFrame({
    "Date": pd.date_range("2024-01-01", periods=30, freq="D"),
    "Sales": np.random.randint(100, 500, 30),
    "Visitors": np.random.randint(50, 200, 30)
})

print("Before setting:")
print(df.head())
print()

# Set date as index
df = df.set_index("Date")
print("After setting:")
print(df.head())
print()

# Use loc to query by date
print("Query from 2024-01-05 to 2024-01-10:")
print(df.loc["2024-01-05":"2024-01-10"])

Resampling

Resampling is the process of converting time series data from one frequency to another, including upsampling (increasing data points) and downsampling (decreasing data points).

Downsampling

Example

import pandas as pd
import numpy as np

# Create daily-level data
np.random.seed(42)
dates = pd.date_range("2024-01-01", periods=90, freq="D")
ts = pd.Series(np.random.randint(100, 500, 90), index=dates)

print("Daily-level data (first 10 items):")
print(ts.head(10))
print()

# Resample by month (sum)
monthly = ts.resample("M").sum()
print("Monthly sum:")
print(monthly)
print()

# Resample by month (average)
monthly_mean = ts.resample("M").mean()
print("Monthly average:")
print(monthly_mean)

Upsampling

Example

import pandas as pd

# Create low-frequency data
dates = pd.date_range("2024-01-01", periods=3, freq="M")
ts = pd.Series([100, 200, 150], index=dates)

print("Monthly-level data:")
print(ts)
print()

# Upsample to daily level (needs fill method)
ts_daily = ts.resample("D").ffill()
print("Upsampled to daily level (first 10 items):")
print(ts_daily.head(10))

Common Resampling Methods

Method Description
sum() Sum
mean() Mean
max() / min() Max/Min value
first() / last() First/Last value
count() Count of non-null values
ohlc() Open, High, Low, Close

Rolling Calculations

Rolling calculation is a sliding window calculation on time series data, commonly used to compute moving averages, moving standard deviations, etc.

Moving Average

Example

import pandas as pd
import numpy as np

# Create data
np.random.seed(42)
dates = pd.date_range("2024-01-01", periods=30, freq="D")
ts = pd.Series(np.random.randint(100, 200, 30), index=dates)

# Calculate 7-day moving average
rolling_mean = ts.rolling(window=7).mean()
print("7-day moving average (first 10 items):")
print(rolling_mean.head(10))
print()

# Calculate 7-day moving standard deviation
rolling_std = ts.rolling(window=7).std()
print("7-day moving standard deviation:")
print(rolling_std.head(10))

Rolling Apply Custom Function

Example

import pandas as pd
import numpy as np

ts = pd.Series(range(1, 11))

print("Original data:")
print(ts)
print()

# Rolling sum
print("Rolling 3 sum:")
print(ts.rolling(3).sum())
print()

# Rolling max
print("Rolling 3 max:")
print(ts.rolling(3).max())
print()

# Use apply
print("Rolling 3 custom function (range):")
print(ts.rolling(3).apply(lambda x: x.max() - x.min()))

Time Series Data Visualization

Example

import pandas as pd
import numpy as np

# Create sample data
np.random.seed(42)
dates = pd.date_range("2024-01-01", periods=100, freq="D")
ts = pd.Series(
    np.random.randn(100).cumsum() + 100,
    index=dates
)

# Calculate moving average
ma_7 = ts.rolling(7).mean()
ma_30 = ts.rolling(30).mean()

# Display data
print("Time series + moving average:")
print(f"Original data (first 5 items): {ts.head().tolist()}")
print(f"7-day moving average (first 10 items): {ma_7.dropna().head().tolist()}")
print(f"30-day moving average (last 5 items): {ma_30.dropna().tail().tolist()}")

Time Series Feature Extraction

Example

import pandas as pd

# Create time series
ts = pd.Series(
    [1, 2, 3, 4, 5, 6, 7, 8, 9, 10],
    index=pd.date_range("2024-01-01", periods=10, freq="D")
)

# Difference
diff = ts.diff()
print("First-order difference:")
print(diff)
print()

# Percentage change
pct = ts.pct_change()
print("Percentage change:")
print(pct)
print()

# Shift
shifted = ts.shift(1)
print("Shift back by 1:")
print(shifted)

Hands-on: Stock Data Analysis

Example

import pandas as pd
import numpy as np

# Simulate stock data
np.random.seed(42)
n_days = 60

df = pd.DataFrame({
    "Date": pd.date_range("2024-01-01", periods=n_days, freq="D"),
    "Open price": 100 + np.random.randn(n_days).cumsum(),
    "Close price": 100 + np.random.randn(n_days).cumsum(),
    "Volume": np.random.randint(1000000, 10000000, n_days)
})
df = df.set_index("Date")

# Calculate daily return
df["Return"] = df["Close price"].pct_change()

# Calculate volatility (7-day rolling standard deviation)
df["Volatility"] = df["Return"].rolling(7).std() * np.sqrt(252)  # Annualize

# Calculate moving average line
df["MA5"] = df["Close price"].rolling(5).mean()
df["MA20"] = df["Close price"].rolling(20).mean()

# Generate trading signals (golden cross / death cross)
df["Signal"] = 0
df.loc[df["MA5"] > df["MA20"], "Signal"] = 1
df.loc[df["MA5"] < df["MA20"], "Signal"] = -1

print("Stock data analysis results:")
print(df.tail(10))

Common Issues

1. Dates are not continuous

Some time series do not have data for every day (holidays, etc.), need to useasfreqorreindexto handle.

2. Time zone issues

When handling cross-timezone data, usetz_localizeandtz_convertfor time zone setting and conversion.

3. Impact of missing values

Rolling calculations skip missing values by default, but this may affect the continuity of the results.

Time series analysis is a fundamental skill in finance, meteorology, IoT, and other fields. Pandas provides a complete toolchain for handling such data.

Other Extensions