Pandas df.tail() Function
tail()It is an important function in Pandas DataFrame and Series, used to view the end part of a dataset. It returns the last n rows of data, allowing us to quickly understand the ending situation and latest records of the data without traversing the entire dataset.
In time series data analysis,tail()it is particularly useful because we usually care about the latest data, such as recent stock prices, sales, or sensor readings. This function andhead()are complementary to each other, and together they constitute the basic tools for data exploration.
Basic Syntax and Parameters
tail()It is a member function of DataFrame and Series, invoked via the dot operator.to call. It andhead()have exactly the same usage, except the viewing direction is opposite.
Syntax Format
DataFrame.tail(n=5) Series.tail(n=5)
Parameter Description
| Parameter | Type | Required | Description | Default Value |
|---|---|---|---|---|
| n | int | Optional | Returns the last n rows of data. If n is greater than the total number of rows, returns all data. | 5 |
Return Value Description
- Return Value Type: When the caller is a DataFrame, returns a DataFrame; when the caller is a Series, returns a Series.
- Number of Rows Returned: Returns at most n rows; if the data has fewer than n rows, returns all data.
- Index Preservation: The returned data retains the original DataFrame's index values.
Examples
Let us comprehensively mastertail()its usage through examples.
Example 1: Basic Usage - View the Last Rows of a DataFrame
Create a DataFrame and usetail()to view the data at the end.
Example
# Create an example DataFrame
data = {
'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve', 'Frank', 'Grace', 'Henry', 'Iris', 'Jack'],
'age': [18, 19, 17, 18, 20, 19, 18, 17, 19, 18],
'score': [85, 92, 78, 90, 88, 95, 82, 76, 89, 91],
'grade': ['A', 'A', 'B', 'A', 'B', 'A', 'B', 'C', 'B', 'A']
}
df = pd.DataFrame(data)
# Returns the last 5 rows by default
print("Default last 5 rows:")
print(df.tail())
# Specify returning the last 3 rows
print("nLast 3 rows:")
print(df.tail(3))
# Returns the last 7 rows
print("nLast 7 rows:")
print(df.tail(7))
Output:
默认最后 5 行:
name age score grade
5 Frank 19 95 A
6 Grace 18 82 B
7 Henry 17 76 C
8 Iris 19 89 B
9 Jack 18 91 A
最后 3 行:
name age score grade
7 Henry 17 76 C
8 Iris 19 89 B
9 Jack 18 91 A
最后 7 行:
name age score grade
3 David 18 90 A
4 Eve 20 88 B
5 Frank 19 95 A
6 Grace 18 82 B
7 Henry 17 76 C
8 Iris 19 89 B
9 Jack 18 91 A
Code Explanation:
- The DataFrame has 10 rows of data, with indices from 0 to 9.
df.tail()Returns the last 5 rows with indices 5 to 9.df.tail(3)Only returns the last 3 rows, with indices 7, 8, and 9.- Note that the index values retain the original DataFrame's indices; they are not renumbered starting from 0.
Example 2: View the Last Rows of a Series
tail()It also applies to Series objects.
Example
# Create a Series
s = pd.Series([10, 20, 30, 40, 50, 60, 70, 80, 90, 100])
# View the last 3 elements
print("Last 3 elements:")
print(s.tail(3))
# Create a Series with indices
s2 = pd.Series([100, 200, 300, 400, 500], index=['a', 'b', 'c', 'd', 'e'])
print("nLast 2 elements of the indexed Series:")
print(s2.tail(2))
Output:
最后 3 个元素: 7 80 8 90 9 100 dtype: int64 带索引 Series 的最后 2 个元素: d 500 e 600 dtype: int64
Code Explanation:
- The Series'
tail()method returns the last n elements, preserving the original indices. - Series with custom indices also works, returning the corresponding elements at the end.
Example 3: Application in Time Series Data
When dealing with time series data,tail()it is especially useful for quickly viewing the latest data records.
Example
import numpy as np
# Create a time series DataFrame to simulate stock data
np.random.seed(42)
dates = pd.date_range('2024-01-01', periods=30, freq='D')
prices = 100 + np.cumsum(np.random.randn(30)) # Simulate stock price trends
stock_df = pd.DataFrame({
'date': dates,
'price': prices.round(2),
'volume': np.random.randint(1000, 10000, 30)
})
print("Comparison of the first few rows and the last few rows of the complete data:")
print("nFirst 5 rows:")
print(stock_df.head())
print("nLast 5 rows:")
print(stock_df.tail())
# Use tail to view the latest few days of data
latest_data = stock_df.tail(7)
print("nData for the last 7 days:")
print(latest_data)
# Further analysis can be performed on the tail result
print("nAverage price for the last 7 days:", latest_data['price'].mean())
print("Total trading volume for the last 7 days:", latest_data['volume'].sum())
Output:
完整数据的前几行和后几行对比:
前 5 行:
date price volume
0 2024-01-01 100.34 5234
1 2024-01-02 99.87 3421
2 2024-01-03 100.21 6789
3 2024-01-04 98.45 4321
4 2024-聚合-05 99.12 5654
最后 5 行:
date price volume
25 2024-01-26 102.45 4321
26 2024-01-27 103.12 5678
27 2024-01-28 101.89 6543
28 2024-01-29 104.23 3211
29 2024-01-30 103.87 4532
最近 7 天的数据:
date price volume
23 2024-01-24 101.56 4321
24 2024-01-25 102.01 5432
25 2024-01-26 <td> 45 4321
26 2024-01-27 103.12 5678
27 2024-01-28 101.89 6543
28 2024-01-29 104.23 3211
29 2024-01-30 103.87 4532
最近 7 天的平均价格: 102.79
最近 7 天的总成交量: 34938
Code Explanation:
- Created a DataFrame containing 30 days of stock price data.
tail()Quickly locates the latest data records.- Pair
tail()The returned data can be further processed by calling other functions for calculation and analysis. - This pattern is very useful in real-time data analysis scenarios.
Important Notes
tail()It does not modify the original DataFrame or Series; it returns a new object.- When n is less than or equal to 0, it returns an empty DataFrame or empty Series.
- For large datasets,
tail()it is more efficient than loading all the data because it does not need to read the entire file. tail()andhead()They are often used together to comprehensively understand both ends of the data.
Tip: In the exploration phase of data analysis, it is recommended to use
head()andtail()so that you can quickly determine whether the data is sorted as expected, as well as data completeness.
Summary
tail()It is a function in Pandas used to view the tail portion of data, andhead()they are complementary to each other. It is especially useful in scenarios such as time series analysis, real-time data monitoring, and log analysis.
Masteringhead()andtail()these two functions allows you to quickly understand the beginning and end of a dataset, which is a very important first step in the data exploration phase. Combined withinfo()anddescribe()and other functions, we can comprehensively grasp the characteristics and distribution of the data.
Pandas Common Functions