Python Web Scraping - BeautifulSoup
Python web scraping refers to the process of automatically extracting information from the internet by writing Python programs.
The basic web scraping process usually includes sending HTTP requests to fetch web page content, parsing the web pages and extracting data, and then storing the data.
Python's rich ecosystem makes it a popular language for developing web scrapers, especially because of its powerful library support.
Generally speaking, the web scraping process can be divided into the following steps:
- Send HTTP requests: The scraper obtains HTML pages from the target website via HTTP requests. Commonly used libraries include
requests。 - Parse HTML content: After obtaining the HTML page, the scraper needs to parse the content and extract data. Common libraries include
BeautifulSoup、lxml、Scrapy, etc. - Extract data: Extract the required data by locating HTML elements (such as tags, attributes, class names, etc.).
- Store data: Store the extracted data in formats such as databases, CSV files, and JSON files for later use or analysis.
This chapter mainly introduces BeautifulSoup, which is a Python library for parsing HTML and XML documents. It can extract data from web pages and is commonly used for web scraping and data mining.

BeautifulSoup
BeautifulSoup is a Python library for extracting data from web pages, especially suitable for parsing HTML and XML files.
BeautifulSoup can extract and manipulate web page content by providing a simple API, making it very suitable for web scraping and data extraction tasks.
Installing BeautifulSoup
To use BeautifulSoup, you need to install beautifulsoup4 and lxml or html.parser (an HTML parser).
We can use pip to install these dependencies:
pip install beautifulsoup4 pip install lxml # 推荐使用 lxml 作为解析器(速度更快)
If you don't have lxml, you can use Python's built-in html.parser as the parser.
Basic Usage
BeautifulSoup is used to parse HTML or XML data and provides methods to navigate, search, and modify the parse tree.
Common BeautifulSoup operations include finding tags, getting tag attributes, extracting text, etc.
To use BeautifulSoup, you need to first import BeautifulSoup and load the HTML page into a BeautifulSoup object.
Usually, you will first use a web scraping library (such asrequests) to obtain the web page content:
Example
import requests
# Use requests to get web page content
url = 'https://cn.bing.com/' # Scrape the web page content of the Bing search engine
response = requests.get(url)
# Use BeautifulSoup to parse the web page
soup = BeautifulSoup(response.text, 'lxml') # Use the lxml parser
# Parse web page content with the html.parser parser
# soup = BeautifulSoup(response.text, 'html.parser')
Get the web page title:
Example
import requests
# Specify the website whose title you want to get
url = 'https://cn.bing.com/' # Scrape the web page content of the Bing search engine
# Send an HTTP request to get the web page content
response = requests.get(url)
# Chinese garbled text issue
response.encoding = 'utf-8'
# Ensure the request was successful
if response.status_code == 200:
# Use BeautifulSoup to parse the web page content
soup = BeautifulSoup(response.text, 'lxml')
# Find the <title> tag
title_tag = soup.find('title')
# Print the title text
if title_tag:
print(title_tag.get_text())
else:
print("The <title> tag was not found")
else:
print("Request failed, status code:", response.status_code)
After running the above code, the output title is:
搜索 - Microsoft 必应
Chinese Garbled Text Issue
When using the requests library to scrape Chinese web pages, encoding issues may occur, causing Chinese content to not display correctly. To ensure that Chinese web pages are correctly scraped and displayed, you usually need to handle the character encoding of the web page.
Automatic encoding detection: requests usually automatically infers the web page encoding based on the Content-Type in the response headers, but this may sometimes be inaccurate. In such cases, you can use chardet to automatically detect the encoding.
Example
url = 'https://cn.bing.com/'
response = requests.get(url)
# Use chardet to automatically detect the encoding
import chardet
encoding = chardet.detect(response.content)['encoding']
print(encoding)
response.encoding = encoding
After running the above code, the output is:
utf-8
If you know the encoding of the web page (e.g., utf-8 or gbk), you can directly setresponse.encoding:
response.encoding = 'utf-8' # 或者 'gbk',根据实际情况选择
Finding Tags
BeautifulSoup provides multiple methods for finding tags in a web page. The most commonly used ones includefind()andfind_all()。
find()Returns the first matching tagfind_all()Returns all matching tags
Example
import requests
# Specify the website whose title you want to get
url = 'https://www.baidu.com/' # Scrape the web page content of the Bing search engine
# Send an HTTP request to get the web page content
response = requests.get(url)
# Chinese garbled text issue
response.encoding = 'utf-8'
soup = BeautifulSoup(response.text, 'lxml')
# Find the first <a> tag
first_link = soup.find('a')
print(first_link)
print("----------------------------")
# Get the href attribute of the first <a> tag
first_link_url = first_link.get('href')
print(first_link_url)
print("----------------------------")
# Find all <a> tags
all_links = soup.find_all('a')
print(all_links)
The output result is similar to the following:
<a class="mnav" href="http://news.baidu.com" name="tj_trnews">新闻</a> ---------------------------- http://news.baidu.com ---------------------------- [<a class="mnav" href="http://news.baidu.com" name="tj_trnews">新闻</a>, <a class="mnav" href="https://www.hao123.com" name="tj_trhao123">hao123</a>, <a class="mnav" href="http://map.baidu.com" name="tj_trmap">地图</a>, <a class="mnav" href="http://v.baidu.com" name="tj_trvideo">视频</a>,
Getting Tag Text
Throughget_text()method, you can extract the text content within a tag:
Example
import requests
# Specify the website whose title you want to get
url = 'https://www.baidu.com/' # Scrape the web page content of the Bing search engine
# Send an HTTP request to get the web page content
response = requests.get(url)
# Chinese garbled text issue
response.encoding = 'utf-8'
soup = BeautifulSoup(response.text, 'lxml')
# Get the text content of the first <p> tag
paragraph_text = soup.find('p').get_text()
# Get all text content on the page
all_text = soup.get_text()
print(all_text)
The output result is similar to the following:
Baidu it, and you'll know.
Finding Child and Parent Tags
You can useparentandchildrenproperties to access a tag's parent and child tags:# 获取当前标签的父标签 parent_tag = first_link.parent # 获取当前标签的所有子标签 children = first_link.children
Example
import requests
# Specify the website whose title you want to get
url = 'https://www.baidu.com/' # Scrape the web page content of the Bing search engine
# Send an HTTP request to get the web page content
response = requests.get(url)
# Chinese garbled text issue
response.encoding = 'utf-8'
soup = BeautifulSoup(response.text, 'lxml')
# Find the first <a> tag
first_link = soup.find('a')
print(first_link)
print("----------------------------")
# Get the parent tag of the current tag
parent_tag = first_link.parent
print(parent_tag.get_text())
The output result is similar to the following:
<a class="mnav" href="http://news.baidu.com" name="tj_trnews">新闻</a> ---------------------------- 新闻 hao123 地图 视频 贴吧 登录 更多产品
Finding Tags with Specific Attributes
You can find tags with specific attributes by passing attribute parameters.
For example, find all div tags with the class name example-class:
# 查找所有 class="example-class" 的 <div> 标签
divs_with_class = soup.find_all('div', class_='example-class')
# 查找具有 id="unique-id" 的 <p> 标签
unique_paragraph = soup.find('p', id='unique-id')
Get the search button with the id su:

Example
import requests
# Specify the website whose title you want to get
url = 'https://www.baidu.com/' # Scrape the web page content of the Bing search engine
# Send an HTTP request to get the web page content
response = requests.get(url)
# Chinese garbled text issue
response.encoding = 'utf-8'
soup = BeautifulSoup(response.text, 'lxml')
# Find the <input> tag with id="unique-id"
unique_input = soup.find('input', id='su')
input_value = unique_input['value'] # Get the value of the input field
print(input_value)
The output result is:
Baidu it.
Advanced Usage
CSS Selectors
BeautifulSoup also supports finding tags using CSS selectors.
The select() method allows you to use jQuery-like selector syntax to find tags:
# 使用 CSS 选择器查找所有 class 为 'example' 的 <div> 标签
example_divs = soup.select('div.example')
# 查找所有 <a> 标签中的 href 属性
links = soup.select('a[href]')
Handling Nested Tags
BeautifulSoup supports deeply nested HTML structures. You can handle these structures by recursively finding child tags:
# 查找嵌套的 <div> 标签
nested_divs = soup.find_all('div', class_='nested')
for div in nested_divs:
print(div.get_text())
Modifying Web Page Content
BeautifulSoup allows you to modify HTML content.
We can modify tag attributes, text, or delete tags:
Example
first_link['href'] = 'http://new-url.com'
# Modify the text content of the first <p> tag
first_paragraph = soup.find('p')
first_paragraph.string = 'Updated content'
# Delete a tag
first_paragraph.decompose()
Converting to String
You can convert the parsed BeautifulSoup object back to an HTML string:
# 转换为字符串 html_str = str(soup)
BeautifulSoup Attributes and Methods
The following are commonly used attributes and methods in BeautifulSoup:
| Method/Attribute | Description | Example |
|---|---|---|
BeautifulSoup() | Used to parse HTML or XML documents and return a BeautifulSoup object. | soup = BeautifulSoup(html_doc, 'html.parser') |
.prettify() | Formats and beautifies the document content, generating a structured string. | print(soup.prettify()) |
.find() | Find the first matching tag. | tag = soup.find('a') |
.find_all() | Find all matching tags and return a list. | tags = soup.find_all('a') |
.find_all_next() | Find all tags that match the conditions after the current tag. | tags = soup.find('div').find_all_next('p') |
.find_all_previous() | Find all tags that match the conditions before the current tag. | tags = soup.find('div').find_all_previous('p') |
.find_parent() | Return the parent tag of the current tag. | parent = tag.find_parent() |
.find_all_parents() | Find all parent tags of the current tag. | parents = tag.find_all_parents() |
.find_next_sibling() | Find the next sibling tag of the current tag. | next_sibling = tag.find_next_sibling() |
.find_previous_sibling() | Find the previous sibling tag of the current tag. | prev_sibling = tag.find_previous_sibling() |
.parent | Get the parent tag of the current tag. | parent = tag.parent |
.next_sibling | Get the next sibling tag of the current tag. | next_sibling = tag.next_sibling |
.previous_sibling | Get the previous sibling tag of the current tag. | prev_sibling = tag.previous_sibling |
.get_text() | Extract the text content inside the tag, ignoring all HTML tags. | text = tag.get_text() |
.attrs | Return all attributes of the tag, represented as a dictionary. | href = tag.attrs['href'] |
.string | Get the string content inside the tag. | string_content = tag.string |
.name | Return the name of the tag. | tag_name = tag.name |
.contents | Return all child elements of the tag, returned as a list. | children = tag.contents |
.descendants | Return all descendant elements of the tag, in generator form. | for child in tag.descendants: print(child) |
.parent | Get the parent tag of the current tag. | parent = tag.parent |
.previous_element | Get the previous element of the current tag (excluding text). | prev_elem = tag.previous_element |
.next_element | Get the next element of the current tag (excluding text). | next_elem = tag.next_element |
.decompose() | Remove the current tag and its contents from the tree. | tag.decompose() |
.unwrap() | Remove the tag itself, keeping only its child content. | tag.unwrap() |
.insert() | Insert a new tag or text into the tag. | tag.insert(0, new_tag) |
.insert_before() | Insert a new tag before the current tag. | tag.insert_before(new_tag) |
.insert_after() | Insert a new tag after the current tag. | tag.insert_after(new_tag) |
.extract() | Delete the tag and return the tag. | extracted_tag = tag.extract() |
.replace_with() | Replace the current tag and its contents. | tag.replace_with(new_tag) |
.has_attr() | Check whether the tag has the specified attribute. | if tag.has_attr('href'): |
.get() | Get the value of the specified attribute. | href = tag.get('href') |
.clear() | Clear all contents of the tag. | tag.clear() |
.encode() | Encode the tag content as a byte stream. | encoded = tag.encode() |
.is_empty_element | Check whether the tag is a void element (e.g.<br>、<img>etc.). | if tag.is_empty_element: |
.is_ancestor_of() | Check whether the current tag is an ancestor of the specified tag. | if tag.is_ancestor_of(another_tag): |
.is_descendant_of() | Check whether the current tag is a descendant of the specified tag. | if tag.is_descendant_of(another_tag): |
Other Attributes
| Method/Attribute | Description | Example |
|---|---|---|
.style | Get the inline style of the tag. | style = tag['style'] |
.id | Get the tag'sidattribute. | id = tag['id'] |
.class_ | Get the tag'sclassattribute. | class_name = tag['class'] |
.string | Get the string content inside the tag, ignoring other tags. | content = tag.string |
.parent | Get the parent element of the tag. | parent = tag.parent |
Other
| Method/Attribute | Description | Example |
|---|---|---|
find_all(string) | Use a string to find matching tags. | tag = soup.find_all('div', class_='container') |
find_all(id) | Find the specifiedidtag. | tag = soup.find_all(id='main') |
find_all(attrs) | Find tags with the specified attributes. | tag = soup.find_all(attrs={"href": "http://example.com"}) |