Pandas Index
Index is a core component of Pandas data structures, determining how data is organized and accessed. Pandas supports multiple types of indexes, from simple RangeIndex to complex MultiIndex. This section will introduce the usage of various index types in detail.
Basic Concepts of Index
Index is similar to a database primary key or Excel row number, used to uniquely identify each row of data. A DataFrame has row index (index) and column index (columns).
Example
# Create a simple DataFrame (uses RangeIndex by default)
df = pd.DataFrame({
"Name": ["Zhang San", "Li Si", "Wang Wu"],
"Age": [25, 30, 28],
"City": ["Beijing", "Shanghai", "Guangzhou"]
})
print("DataFrame info:")
print(f"Row index: {df.index.tolist()}")
print(f"Column index: {df.columns.tolist()}")
print("\n"Data:")
print(df)
RangeIndex
RangeIndex is the default integer index, similar to Python'srange(n), increasing from 0.
Creation and Usage
Example
# Create a RangeIndex
idx = pd.RangeIndex(start=0, stop=10, step=1)
print(f"RangeIndex: {idx}")
print(f"Type: {type(idx)}")
# DataFrame uses RangeIndex by default
df = pd.DataFrame({"A": [1, 2, 3]}, index=range(3))
print(f"\n"Default index type: {type(df.index)}")
# Resetting the index also returns a RangeIndex
df_reset = df.reset_index()
print(f"Index type after reset: {type(df_reset.index)}")
RangeIndex Features
- Minimal memory usage
- Supports default integer position access
- Easy conversion to and from other index types
Index Type Conversion
Index supports conversion between multiple types.
Convert to Other Index Types
Example
# Create an example DataFrame
df = pd.DataFrame({"Value": [1, 2, 3, 4]}, index=[10, 20, 30, 40])
# Convert to Index type
print("Original index:", df.index)
print("Index type:", type(df.index))
# Convert to list
idx_list = df.index.tolist()
print(f"Converted to list: {idx_list}")
# Convert to NumPy array
idx_array = df.index.values
print(f"Converted to array: {idx_array}")
# Reset the index to RangeIndex
df = df.reset_index(drop=True)
print(f"Reset to RangeIndex: {df.index}")
Setting Custom Index
You can use a column of the DataFrame as the row index.
Using set_index
Example
# Create DataFrame
df = pd.DataFrame({
"Student ID": ["S001", "S002", "S003", "S004"],
"Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu"],
"Score": [85, 92, 78, 90]
})
print("Original data:")
print(df)
print()
# Set the "Student ID" column as the index
df1 = df.set_index("Student ID")
print("Set Student ID as index:")
print(df1)
print()
# Set multiple indexes (this creates a MultiIndex)
df2 = df.set_index(["Student ID", "Name"])
print("Set multiple indexes:")
print(df2)
Creating with the index parameter
Example
# Specify the index directly when creating the DataFrame
df = pd.DataFrame(
{"Name": ["Zhang San", "Li Si", "Wang Wu"], "Age": [25, 30, 28]},
index=["A001", "A002", "A003"]
)
print(df)
print(f"\n"Index: {df.index.tolist()}")
# Use DatetimeIndex
dates = pd.date_range("2024-01-01", periods=3, freq="D")
df_date = pd.DataFrame({"Value": [100, 200, 300]}, index=dates)
print("\n"Using date index:")
print(df_date)
print(f"Index type: {type(df_date.index)}")
Index Operations
Resetting Index
Example
# Create a DataFrame with a custom index
df = pd.DataFrame(
{"Name": ["Zhang San", "Li Si"], "Score": [85, 92]},
index=["A001", "A002"]
)
# Reset the index to the default RangeIndex
df_reset = df.reset_index()
print("Reset index:")
print(df_reset)
# drop parameter: whether to discard the original index column
df_reset2 = df.reset_index(drop=True)
print("\n"Reset and discard the original index:")
print(df_reset2)
Reindexing
Example
# Create DataFrame
df = pd.DataFrame({"A": [1, 2, 3], "B": [4, 5, 6]}, index=[1, 2, 3])
# Reindex (change the index order)
df_reindex = df.reindex([1, 2, 3, 4, 5])
print("Reindexed (missing values filled with NaN):")
print(df_reindex)
# Use fill_value to fill missing values
df_reindex2 = df.reindex([1, 2, 3, 4, 5], fill_value=0)
print("\n"Reindexed (filled with 0):")
print(df_reindex2)
Index Level Operations
Example
# Create a MultiIndex DataFrame
df = pd.DataFrame({
"Chinese": [85, 92, 78],
"Math": [90, 88, 95]
}, index=pd.MultiIndex.from_tuples(
[("Grade 10", "Class A"), ("Grade 10", "Class B"), ("Grade 11", "Class A")],
names=["Grade", "Class"]
))
print("MultiIndex DataFrame:")
print(df)
print()
# Get the outer index
print(f"Outer index (grade): {df.index.get_level_values(0).tolist()}")
print(f"Inner index (class): {df.index.get_level_values(1).tolist()}")
Index Attributes and Methods
| Attribute/Method | Description | Example |
|---|---|---|
index.tolist() |
Convert to Python list | df.index.tolist() |
index.values |
Convert to NumPy array | df.index.values |
index.unique() |
Get unique values | df.index.unique() |
index.is_unique |
Check whether the index is unique | df.index.is_unique |
index.astype() |
Convert index type | df.index.astype(str) |
Example of Using Index Attributes
Example
# Create an example DataFrame
df = pd.DataFrame(
{"Value": [1, 2, 3, 4]},
index=["a", "b", "c", "d"]
)
# View index attributes
print(f"Index is unique: {df.index.is_unique}")
print(f"Index length: {len(df.index)}")
print(f"Index dtype: {df.index.dtype}")
# Convert to string type
str_index = df.index.astype(str)
print(f"Type after conversion: {str_index.dtype}")
Practical: Using Index to Improve Query Efficiency
In actual data analysis, setting a reasonable index can significantly improve query efficiency.
Example
# Simulate business data
df = pd.DataFrame({
"Order ID": range(1000),
"Customer ID": [f"C{i%100:03d}" for i in range(1000)],
"Product": [f"Product{i%20}" for i in range(1000)],
"Amount": [round(i * 1.5, 2) for i in range(1000)]
})
# Set frequently queried columns as the index
df_indexed = df.set_index(["Customer ID", "Product"])
# Using Index for Fast Query (Similar to Database Primary Key Query)
result = df_indexed.loc[("C001", "Product 1")]
print("Query a single customer's products using index:")
print(result)
# Group by Customer ID for Statistics
print("\nTotal spending by each customer:")
customer_total = df.groupby("Customer ID")["Amount"].sum()
print(customer_total.head(10))
Notes
1. Index values must be unique
If index values are not unique, some operations (such asloclookup) will return multiple matching rows.
2. Indexes can have names
Naming an index can improve code readability:df.index.name = "学号"
3. Indexes follow data operations
Operations such as slicing and filtering preserve the index; you need to pay attention to the correspondence between the index and the data.
Other ExtensionsThe index is key to Pandas performance. Setting frequently queried columns as indexes can greatly improve lookup speed, but too many indexes will increase write overhead, so a trade-off is needed.