Pandas Index

Index is a core component of Pandas data structures, determining how data is organized and accessed. Pandas supports multiple types of indexes, from simple RangeIndex to complex MultiIndex. This section will introduce the usage of various index types in detail.


Basic Concepts of Index

Index is similar to a database primary key or Excel row number, used to uniquely identify each row of data. A DataFrame has row index (index) and column index (columns).

Example

import pandas as pd

# Create a simple DataFrame (uses RangeIndex by default)
df = pd.DataFrame({
    "Name": ["Zhang San", "Li Si", "Wang Wu"],
    "Age": [25, 30, 28],
    "City": ["Beijing", "Shanghai", "Guangzhou"]
})

print("DataFrame info:")
print(f"Row index: {df.index.tolist()}")
print(f"Column index: {df.columns.tolist()}")
print("\n"Data:")
print(df)

RangeIndex

RangeIndex is the default integer index, similar to Python'srange(n), increasing from 0.

Creation and Usage

Example

import pandas as pd

# Create a RangeIndex
idx = pd.RangeIndex(start=0, stop=10, step=1)
print(f"RangeIndex: {idx}")
print(f"Type: {type(idx)}")

# DataFrame uses RangeIndex by default
df = pd.DataFrame({"A": [1, 2, 3]}, index=range(3))
print(f"\n"Default index type: {type(df.index)}")

# Resetting the index also returns a RangeIndex
df_reset = df.reset_index()
print(f"Index type after reset: {type(df_reset.index)}")

RangeIndex Features

  • Minimal memory usage
  • Supports default integer position access
  • Easy conversion to and from other index types

Index Type Conversion

Index supports conversion between multiple types.

Convert to Other Index Types

Example

import pandas as pd

# Create an example DataFrame
df = pd.DataFrame({"Value": [1, 2, 3, 4]}, index=[10, 20, 30, 40])

# Convert to Index type
print("Original index:", df.index)
print("Index type:", type(df.index))

# Convert to list
idx_list = df.index.tolist()
print(f"Converted to list: {idx_list}")

# Convert to NumPy array
idx_array = df.index.values
print(f"Converted to array: {idx_array}")

# Reset the index to RangeIndex
df = df.reset_index(drop=True)
print(f"Reset to RangeIndex: {df.index}")

Setting Custom Index

You can use a column of the DataFrame as the row index.

Using set_index

Example

import pandas as pd

# Create DataFrame
df = pd.DataFrame({
    "Student ID": ["S001", "S002", "S003", "S004"],
    "Name": ["Zhang San", "Li Si", "Wang Wu", "Zhao Liu"],
    "Score": [85, 92, 78, 90]
})

print("Original data:")
print(df)
print()

# Set the "Student ID" column as the index
df1 = df.set_index("Student ID")
print("Set Student ID as index:")
print(df1)
print()

# Set multiple indexes (this creates a MultiIndex)
df2 = df.set_index(["Student ID", "Name"])
print("Set multiple indexes:")
print(df2)

Creating with the index parameter

Example

import pandas as pd

# Specify the index directly when creating the DataFrame
df = pd.DataFrame(
    {"Name": ["Zhang San", "Li Si", "Wang Wu"], "Age": [25, 30, 28]},
    index=["A001", "A002", "A003"]
)

print(df)
print(f"\n"Index: {df.index.tolist()}")

# Use DatetimeIndex
dates = pd.date_range("2024-01-01", periods=3, freq="D")
df_date = pd.DataFrame({"Value": [100, 200, 300]}, index=dates)

print("\n"Using date index:")
print(df_date)
print(f"Index type: {type(df_date.index)}")

Index Operations

Resetting Index

Example

import pandas as pd

# Create a DataFrame with a custom index
df = pd.DataFrame(
    {"Name": ["Zhang San", "Li Si"], "Score": [85, 92]},
    index=["A001", "A002"]
)

# Reset the index to the default RangeIndex
df_reset = df.reset_index()
print("Reset index:")
print(df_reset)

# drop parameter: whether to discard the original index column
df_reset2 = df.reset_index(drop=True)
print("\n"Reset and discard the original index:")
print(df_reset2)

Reindexing

Example

import pandas as pd

# Create DataFrame
df = pd.DataFrame({"A": [1, 2, 3], "B": [4, 5, 6]}, index=[1, 2, 3])

# Reindex (change the index order)
df_reindex = df.reindex([1, 2, 3, 4, 5])
print("Reindexed (missing values filled with NaN):")
print(df_reindex)

# Use fill_value to fill missing values
df_reindex2 = df.reindex([1, 2, 3, 4, 5], fill_value=0)
print("\n"Reindexed (filled with 0):")
print(df_reindex2)

Index Level Operations

Example

import pandas as pd

# Create a MultiIndex DataFrame
df = pd.DataFrame({
    "Chinese": [85, 92, 78],
    "Math": [90, 88, 95]
}, index=pd.MultiIndex.from_tuples(
    [("Grade 10", "Class A"), ("Grade 10", "Class B"), ("Grade 11", "Class A")],
    names=["Grade", "Class"]
))

print("MultiIndex DataFrame:")
print(df)
print()

# Get the outer index
print(f"Outer index (grade): {df.index.get_level_values(0).tolist()}")
print(f"Inner index (class): {df.index.get_level_values(1).tolist()}")

Index Attributes and Methods

Attribute/Method Description Example
index.tolist() Convert to Python list df.index.tolist()
index.values Convert to NumPy array df.index.values
index.unique() Get unique values df.index.unique()
index.is_unique Check whether the index is unique df.index.is_unique
index.astype() Convert index type df.index.astype(str)

Example of Using Index Attributes

Example

import pandas as pd

# Create an example DataFrame
df = pd.DataFrame(
    {"Value": [1, 2, 3, 4]},
    index=["a", "b", "c", "d"]
)

# View index attributes
print(f"Index is unique: {df.index.is_unique}")
print(f"Index length: {len(df.index)}")
print(f"Index dtype: {df.index.dtype}")

# Convert to string type
str_index = df.index.astype(str)
print(f"Type after conversion: {str_index.dtype}")

Practical: Using Index to Improve Query Efficiency

In actual data analysis, setting a reasonable index can significantly improve query efficiency.

Example

import pandas as pd

# Simulate business data
df = pd.DataFrame({
    "Order ID": range(1000),
    "Customer ID": [f"C{i%100:03d}" for i in range(1000)],
    "Product": [f"Product{i%20}" for i in range(1000)],
    "Amount": [round(i * 1.5, 2) for i in range(1000)]
})

# Set frequently queried columns as the index
df_indexed = df.set_index(["Customer ID", "Product"])

# Using Index for Fast Query (Similar to Database Primary Key Query)
result = df_indexed.loc[("C001", "Product 1")]
print("Query a single customer's products using index:")
print(result)

# Group by Customer ID for Statistics
print("\nTotal spending by each customer:")
customer_total = df.groupby("Customer ID")["Amount"].sum()
print(customer_total.head(10))

Notes

1. Index values must be unique

If index values are not unique, some operations (such asloclookup) will return multiple matching rows.

2. Indexes can have names

Naming an index can improve code readability:df.index.name = "学号"

3. Indexes follow data operations

Operations such as slicing and filtering preserve the index; you need to pay attention to the correspondence between the index and the data.

The index is key to Pandas performance. Setting frequently queried columns as indexes can greatly improve lookup speed, but too many indexes will increase write overhead, so a trade-off is needed.

Other Extensions