Pandas pd.read_html() Function

Python math 模块Pandas Common Functions


read_html()It is a function in the pandas library used to parse HTML tables, which can read table data from web pages or HTML files and convert it into a DataFrame.

Table data in web pages is a very important data source. Many public data sets (such as stock information, statistical data, etc.) are presented in the form of HTML tables.read_html()UsinglxmlandBeautifulSoupthe library to parse HTML, it can automatically extract all tables or specified tables from the page.


Basic Syntax and Parameters

Syntax Format

pandas.read_html(io, match='.+', flavor=None, header=None, index_col=None,
                skiprows=None, attrs=None, parse_dates=False, thousands=',',
                decimal='.', converters=None, ...)

Parameter Description

ParameterTypeDescriptionDefault Value
iostr, path object, file-like objectHTML file path, URL, or stringRequired
matchstr, regexUse regular expressions to match the text content of the table'.+'
flavorstrParser: 'lxml', 'html5lib', 'bs4'None
headerint, list of intRow number used as column namesNone
index_colint, strColumn used as row indexNone
skiprowsint, list, sliceSkip specified rowsNone
attrsdictHTML tag attributes used to filter tablesNone
parse_datesbool, listWhether to parse date columnsFalse

Return Value

  • Return Type:list of DataFrames
  • Returns a list of DataFrames; the number of DataFrames returned corresponds to the number of tables on the page.
  • If no matching table is found, returns an empty list.

Examples

Through the following examples, comprehensively masterread_html()the various usages.

Example 1: Read a table from a local HTML file

First create an HTML file containing a table, then useread_html()to read it.

Example

import pandas as pd

# Create an HTML file containing a table
html_content = '''



<title>Employee Information Table</title>


<h1>Pandas pd-read-html() Function</h1>
    <table border="1" class="employee-table">
        <tr>
<th>Name</th>
<th>Age</th>
<th>City</th>
<th>Salary</th>
        </tr>
        <tr>
            <td>Tom</td>
            <td>28</td>
            <td>Beijing</td>
            <td>8000</td>
        </tr>
        <tr>
            <td>Jerry</td>
            <td>35</td>
            <td>Shanghai</td>
            <td>12000</td>
        </tr>
        <tr>
            <td>Mike</td>
            <td>42</td>
            <td>Guangzhou</td>
            <td>15000</td>
        </tr>
        <tr>
            <td>Lucy</td>
            <td>26</td>
            <td>Shenzhen</td>
            <td>7000</td>
        </tr>
    </table>

<h2>Department List</h2>
    <table border="1">
        <tr>
<th>Department</th>
<th>Headcount</th>
        </tr>
        <tr>
<td>Technology Department</td>
            <td>50</td>
        </tr>
        <tr>
<td>Sales Department</td>
            <td>30</td>
        </tr>
    </table>


'''


# Write the HTML to a file
with open('tables.html', 'w', encoding='utf-8') as f:
    f.write(html_content)

# Use read_html to read all tables
# io: HTML file path (required)
tables = pd.read_html('tables.html')

# View the read results
print(f"Found {len(tables)} tables")
print()

# Iterate over all tables
for i, df in enumerate(tables):
    print(f"--- Table {i+1} ---")
    print(df)
    print()

Expected output:

共找到 2 个表格

--- 表格 1 ---
    姓名  年龄       城市    薪资
0   Tom   28    Beijing   8000
1  Jerry   35   Shanghai  12000
2   Mike   42  Guangzhou  15000
3   Lucy   26  Shenzhen   7000

--- 表格 2 ---
    部门  人数
0  技术部   50
1  销售部   30

Code explanation:

  • read_html()Returns a list of DataFrames, each table corresponding to one DataFrame.
  • By default, all tables are read.
  • The first row is automatically recognized as the column names (because there arethtags).

Example 2: Filter tables using attrs and match

When a page has multiple tables, you can use attribute or text matching to filter the desired tables.

Example

import pandas as pd

# Create an HTML file with attributes
html_with_attrs = '''



    <table id="employees" class="data-table">
        <tr><th>name</th><th>age</th></tr>
        <tr><td>Tom</td><td>28</td></tr>
        <tr><td>Jerry</td><td>35</td></tr>
    </table>

    <table id="products" class="data-table">
        <tr><th>product</th><th>price</th></tr>
        <tr><td>A</td><td>100</td></tr>
        <tr><td>B</td><td>200</td></tr>
    </table>

    <table class="summary">
        <tr><td>Total</td><td>2</td></tr>
    </table>


'''


with open('tables_attrs.html', 'w', encoding='utf-8') as f:
    f.write(html_with_attrs)

# Example 2a: Use attrs to filter by id attribute
# Read the table with id="employees"
tables_by_id = pd.read_html('tables_attrs.html', attrs={'id': 'employees'})
print("Using id to filter:")
print(tables_by_id[0])
print()

# Example 2b: Use attrs to filter by class attribute
# Read all tables with class="data-table"
tables_by_class = pd.read_html('tables_attrs.html', attrs={'class': 'data-table'})
print("Using class to filter (found {} tables):".format(len(tables_by_class)))
for i, df in enumerate(tables_by_class):
    print(f"Table {i+1}:")
    print(df)
    print()

# Example 2c: Use match to filter tables containing specific text
# match uses regular expressions to match text in tables
tables_by_text = pd.read_html('tables_attrs.html', match='Tom')
print("Tables containing the text 'Tom':")
print(tables_by_text[0])

Expected output:

使用 id 筛选:
   name  age
0   Tom   28
1  Jerry   35

使用 class 筛选 (找到 2 个表格):
表格 1:
   name  age
0   Tom   28
1  Jerry   35

表格 2:
  product  price
0       A    100
1       B    200

包含 'Tom' 文本的表格:
   name  age
0   Tom   28
1  Jerry   35

Code explanation:

  • attrsThe parameter can specify HTML tag attributes (such as id, class, style, etc.) to filter tables.
  • matchThe parameter uses regular expressions to match the text content of the table, returning tables that contain the matching text.
  • When only one table is returned, you can usetables[0]to get the first DataFrame.

Example 3: Handle headers and indexes

HTML tables may have complex formats, so headers and indexes need to be handled flexibly.

h2 class="example">Example
import pandas as pd

# Create HTML with a complex structure
html_complex = '''



<!-- Table without a header -->
    <table id="no_header">
        <tr><td>Tom</td><td>28</td></tr>
        <tr><td>Jerry</td><td>35</td></tr>
    </table>

<!-- Multi-line header -->
    <table id="multi_header">
<tr><th>Name</th><th colspan="2">Contact Information</th></tr>
<tr><th></th><th>Phone</th><th>Email</th></tr>
        <tr><td>Tom</td><td>123456</td><td>[email protected]</td></tr>
        <tr><td>Jerry</td><td>789012</td><td>[email protected]</td></tr>
    </table>

<!-- Table with attributes -->
    <table id="with_index" data-type="employee">
<tr><th>Name</th><th>Age</th><th>City</th></tr>
        <tr><th></th><th></th><th></th></tr>
        <tr><td>Tom</td><td>28</td><td>Beijing</td></tr>
    </table>


'''


with open('tables_complex.html', 'w', encoding='utf-8') as f:
    f.write(html_complex)

# Example 3a: Case without a header
# header=None does not use the first row as column names
df_no_header = pd.read_html('tables_complex.html', attrs={'id': 'no_header'}, header=None)[0]
print("No header:")
print(df_no_header)
print()

# Example 3b: Multi-line header
# Use row 0 and row 1 together as the header
df_multi_header = pd.read_html('tables_complex.html', attrs={'id': 'multi_header'})[0]
print("Multi-line header:")
print(df_multi_header)
print()

# Example 3c: Set the index column
# index_col specifies column 0 as the index
df_with_index = pd.read_html('tables_complex.html', attrs={'id': 'with_index'}, index_col=0)[0]
print("Setting index:")
print(df_with_index)

Expected output:

没有表头:
      0     1
0  Tom    28
1  Jerry  35

多行表头:
       姓名    联系方式
        NaN    电话          邮箱
0     Tom  123456  [email protected]
1   Jerry  789012  [email protected]

设置索引:
                年龄       城市
姓名
Tom            28    Beijing
</空行>
空字符串    NaN       NaN
Tom            28    Beijing

Code explanation:

  • header=NoneYou can disable automatic header recognition and use the default integer index as column names.
  • Multi-line headers are processed as a MultiIndex.
  • index_colThe parameter can specify a column to be used as the row index.

Notes

  • Usingread_html()requires installinglxmlthe library:pip install lxml。
  • It returns a list of DataFrames; you need to select a specific table based on the index.
  • When there are multiple tables on the page, useattrsormatchthe parameter to filter.
  • read_html()It will try to parse all tables, which may be slow. For large pages, you can usematchthe parameter to narrow down the scope.
  • Reading web pages requires network support, and there may be access restrictions or anti-crawling mechanisms.

Summary

read_html()It is a powerful tool in pandas for reading HTML table data. It can automatically parse tables in web pages and convert them into structured DataFrame format.

In practical work, if you need to obtain data from web pages,read_html()it is a very practical choice. Mastering the use of the attrs and match parameters allows you to efficiently extract the required data from complex pages. It is recommended that readers give priority to this function when they need to scrape web tables.

Python math 模块Pandas Common Functions

Other Extensions