Python urllib

The Python urllib library is used to operate on web URLs and scrape web content.

This article mainly introduces Python3's urllib.

The urllib package contains the following modules:

  • urllib.request- Open and read URLs.
  • urllib.error- Contains exceptions raised by urllib.request.
  • urllib.parse- Parse URLs.
  • urllib.robotparser- Parse robots.txt files.


urllib.request

urllib.request defines some functions and classes for opening URLs, including authorization verification, redirects, browser cookies, etc.

urllib.request can simulate a browser's request initiation process.

We can use the urlopen method of urllib.request to open a URL, the syntax is as follows:

urllib.request.urlopen(url, data=None, [timeout, ]*, cafile=None, capath=None, cadefault=False, context=None)
  • url: url address.
  • data: Other data objects sent to the server, default is None.
  • timeout: Set access timeout.
  • cafile and capath: cafile is the CA certificate, capath is the path to the CA certificate, required for HTTPS.
  • cadefault: Deprecated.
  • context: ssl.SSLContext type, used to specify SSL settings.

Example is as follows:

Example

from urllib.request import urlopen

myURL = urlopen("https://www.example.com/")
print(myURL.read())

The above code uses urlopen to open a URL, then uses the read() function to get the HTML source code of the webpage.

read() reads the entire webpage content; we can specify the length to read:

Example

from urllib.request import urlopen

myURL = urlopen("https://www.example.com/")
print(myURL.read(300))

In addition to the read() function, it also includes the following two functions for reading webpage content:

  • readline()- Read one line of the file

    from urllib.request import urlopen
    
    myURL = urlopen("https://www.example.com/")
    print(myURL.readline()) #读取一行内容
  • readlines()- Read all content of the file, it assigns the read content to a list variable.

    from urllib.request import urlopen
    
    myURL = urlopen("https://www.example.com/")
    lines = myURL.readlines()
    for line in lines:
        print(line) 

When scraping webpages, we often need to determine whether the webpage can be accessed normally. Here we can use the getcode() function to get the webpage status code. Returning 200 means the webpage is normal, returning 404 means the webpage does not exist:

Example

import urllib.request

myURL1 = urllib.request.urlopen("https://www.example.com/")
print(myURL1.getcode())   # 200

try:
    myURL2 = urllib.request.urlopen("https://www.example.com/no.html")
except urllib.error.HTTPError as e:
    if e.code == 404:
        print(404)   # 404

For more webpage status codes, refer to:https://www.example.com/http/http-status-codes.html。

If you want to save the scraped webpage locally, you can usePython3 File write() methodfunction:

Example

from urllib.request import urlopen

myURL = urlopen("https://www.example.com/")
f = open("example_urllib_test.html", "wb")
content = myURL.read()  # Read webpage content
f.write(content)
f.close()

After executing the above code, a example_urllib_test.html file will be generated locally, containing the content of the https://www.example.com/ webpage.

For more Python File handling, refer to:https://www.example.com/python3/python3-file-methods.html

。

URL encoding and decoding can useurllib.request.quote()andurllib.request.unquote()methods:

Example

import urllib.request

encode_url = urllib.request.quote("https://www.example.com/")  # Encoding
print(encode_url)

unencode_url = urllib.request.unquote(encode_url)    # Decoding
print(unencode_url)

The output result is:

https%3A//www.example.com/
https://www.example.com/

Simulate Headers

When scraping webpages, we generally need to simulate headers (webpage header information). At this time, we need to use the urllib.request.Request class:

class urllib.request.Request(url, data=None, headers={}, origin_req_host=None, unverifiable=False, method=None)
  • url: url address.
  • data: Other data objects sent to the server, default is None.
  • headers: HTTP request header information, in dictionary format.
  • origin_req_host: The host address of the request, IP or domain name.
  • unverifiable: This parameter is rarely used. It is used to set whether the webpage requires verification. Default is False.
  • method: Request method, such as GET, POST, DELETE, PUT, etc.

Example - py3_urllib_test.py file code

import urllib.request
import urllib.parse

url = 'https://www.example.com/?s='  # Example Tutorial search page
keyword = 'Python Tutorial'
key_code = urllib.request.quote(keyword)  # Encode the request
url_all = url+key_code
header = {
    'User-Agent':'Mozilla/5.0 (X11; Fedora; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36'
}   # Header information
request = urllib.request.Request(url_all,headers=header)
reponse = urllib.request.urlopen(request).read()

fh = open("./urllib_test_example_search.html","wb")    # Write the file to the current directory
fh.write(reponse)
fh.close()

After executing the above Python code, an urllib_test_example_search.html file will be generated in the current directory. Open the urllib_test_example_search.html file (you can open it with a browser), the content is as follows:

For form POST data transmission, let's first create a form. The code is as follows. Here I use PHP code to get the form data:

Example - py3_urllib_test.php file code:

<!DOCTYPE html>
<html>
<head>
<meta charset="utf-8">
<title>Example Tutorial (example.com) urllib POST Test</title>
</head>
<body>
<form action="" method="post" name="myForm">
    Name: <input type="text" name="name"><br>
    Tag: <input type="text" name="tag"><br>
    <input type="submit" value="Submit">
</form>
<hr>
<?php
//Use PHP to get the data submitted by the form, you can replace it with other languages
if(isset($_POST['name']) && $_POST['tag'] ) {
   echo $_POST["name"] . ', ' . $_POST['tag'];
}
?>
</body>
</html>

Example

import urllib.request
import urllib.parse

url = 'https://www.example.com/try/py3/py3_urllib_test.php'  # Submit to the form page
data = {'name':'EXAMPLE', 'tag' : 'Rookie Tutorial'}   # Submit data
header = {
    'User-Agent':'Mozilla/5.0 (X11; Fedora; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.36'
}   # Header information
data = urllib.parse.urlencode(data).encode('utf8')  # Encode the parameters, use urllib.parse.urldecode for decoding
request=urllib.request.Request(url, data, header)   # Request processing
reponse=urllib.request.urlopen(request).read()      # Read result

fh = open("./urllib_test_post_example.html","wb")    # Write the file to the current directory
fh.write(reponse)
fh.close()

Executing the above code will submit form data to the py3_urllib_test.php file, and write the output result to the urllib_test_post_example.html file.

Open the urllib_test_post_example.html file (you can open it with a browser), the displayed result is as follows:


urllib.error

The urllib.error module defines exception classes for exceptions raised by urllib.request, and the base exception class is URLError.

urllib.error contains two exception classes, URLError and HTTPError.

URLError is a subclass of OSError, used to handle exceptions (or derived exceptions) that the program raises when encountering problems. The attribute reason it contains is the cause of the exception.

HTTPError is a subclass of URLError, used to handle special HTTP errors such as authentication requests. The attributes it contains are:codecode: the HTTP status code,reasonreason: the cause of the exception,headersheaders: the HTTP response headers of the specific HTTP request that caused the HTTPError.

Scrape a non-existent webpage and handle exceptions:

Example

import urllib.request
import urllib.error

myURL1 = urllib.request.urlopen("https://www.example.com/")
print(myURL1.getcode())   # 200

try:
    myURL2 = urllib.request.urlopen("https://www.example.com/no.html")
except urllib.error.HTTPError as e:
    if e.code == 404:
        print(404)   # 404


urllib.parse

urllib.parse is used to parse URLs, the format is as follows:

urllib.parse.urlparse(urlstring, scheme='', allow_fragments=True)

urlstring is the string url address, scheme is the protocol type,

If the allow_fragments parameter is false, fragment identifiers cannot be recognized. Instead, they are parsed as part of the path, parameters, or query components, and fragment is set to an empty string in the return value.

Example

from urllib.parse import urlparse

o = urlparse("https://www.example.com/?s=python+%E6%95%99%E7%A8%8B")
print(o)

The output result of the above example is:

ParseResult(scheme='https', netloc='www.example.com', path='/', params='', query='s=python+%E6%95%99%E7%A8%8B', fragment='')

From the result, it can be seen that the content is a tuple containing 6 strings: protocol, location, path, parameters, query, and fragment.

We can directly read the protocol content:

Example

from urllib.parse import urlparse

o = urlparse("https://www.example.com/?s=python+%E6%95%99%E7%A8%8B")
print(o.scheme)

The output result of the above example is:

https

The complete content is as follows:

Attribute

Index

Value

Value (if not present)

scheme

0

URL protocol

schemeParameters

netloc

1

Network location part

Empty string

path

2

Hierarchical path

Empty string

params

3

Parameters of the last path element

Empty string

query

4

Query component

Empty string

fragment

5

Fragment identifier

Empty string

username

Username

None

password

Password

None

hostname

Hostname (lowercase)

None

port

Port number as an integer (if present)

None


urllib.robotparser

urllib.robotparser is used to parse robots.txt files.

robots.txt (all lowercase) is a robots protocol stored in the root directory of a website. It is usually used to tell search engines the crawling rules for the website.

urllib.robotparser provides the RobotFileParser class, the syntax is as follows:

class urllib.robotparser.RobotFileParser(url='')

This class provides methods that can read and parse robots.txt files:

  • set_url(url) - Sets the URL of the robots.txt file.

  • read() - Reads the robots.txt URL and feeds it to the parser.

  • parse(lines) - Parses the lines parameter.

  • can_fetch(useragent, url) - Returns True if the useragent is allowed to fetch the url according to the rules in the parsed robots.txt file.

  • mtime() - Returns the time the robots.txt file was last fetched. This is useful for long-running web crawlers that need to periodically check for robots.txt updates.

  • modified() - Sets the time the robots.txt file was last fetched to the current time.

  • crawl_delay(useragent) - Returns the Crawl-delay parameter from robots.txt for the specified useragent. Returns None if this parameter does not exist, does not apply to the specified useragent, or if the robots.txt entry for this parameter has a syntax error.

  • request_rate(useragent) - Returns the content of the Request-rate parameter from robots.txt as a named tuple RequestRate(requests, seconds). Returns None if this parameter does not exist, does not apply to the specified useragent, or if the robots.txt entry for this parameter has a syntax error.

  • site_maps() - Returns the content of the Sitemap parameter from robots.txt as a list(). Returns None if this parameter does not exist or if the robots.txt entry for this parameter has a syntax error.

Example

>>> import urllib.robotparser
>>> rp = urllib.robotparser.RobotFileParser()
>>> rp.set_url("http://www.musi-cal.com/robots.txt")
>>> rp.read()
>>> rrate = rp.request_rate("*")
>>> rrate.requests
3
>>> rrate.seconds
20
>>> rp.crawl_delay("*")
6
>>> rp.can_fetch("*", "http://www.musi-cal.com/cgi-bin/search?city=San+Francisco")
False
>>> rp.can_fetch("*", "http://www.musi-cal.com/")
True
Other extensions