Python3 XML Parsing
What Is XML?
XML refers to the Extensible Markup Language (eXtensible Markup LXtensible Markup Language), a subset of the Standard Generalized Markup Language, and is a markup language used for marking electronic files so that they are structured. You can learn from this siteXML Tutorial
XML was designed to transmit and store data.
XML is a set of rules for defining semantic tags. These tags divide a document into many parts and identify those parts.
It is also a meta-markup language, that is, a syntactic language used to define other semantic, structured markup languages related to specific domains.
Python's XML Parsing
Common XML programming interfaces include DOM and SAX. These two interfaces handle XML files differently, so they are used in different situations.
Python has three methods for parsing XML: ElementTree, SAX, and DOM.
1. ElementTree
xml.etree.ElementTree is a module in the Python standard library for handling XML. It provides a simple and efficient API for parsing and generating XML documents.
2.SAX (simple API for XML )
The Python standard library includes a SAX parser. SAX uses an event-driven model, processing XML files by triggering events one by one and calling user-defined callback functions during XML parsing.
3.DOM(Document Object Model)
It parses XML data into a tree in memory and operates on XML by manipulating the tree.
The XML example files used in this chaptermovies.xmlThe content is as follows:
Examples
<movie title="Enemy Behind">
<type>War, Thriller</type>
<format>DVD</format>
<year>2003</year>
<rating>PG</rating>
<stars>10</stars>
<description>Talk about a US-Japan war</description>
</movie>
<movie title="Transformers">
<type>Anime, Science Fiction</type>
<format>DVD</format>
<year>1989</year>
<rating>R</rating>
<stars>8</stars>
<description>A schientific fiction</description>
</movie>
<movie title="Trigun">
<type>Anime, Action</type>
<format>DVD</format>
<episodes>4</episodes>
<rating>PG</rating>
<stars>10</stars>
<description>Vash the Stampede!</description>
</movie>
<movie title="Ishtar">
<type>Comedy</type>
<format>VHS</format>
<rating>PG</rating>
<stars>2</stars>
<description>Viewable boredom</description>
</movie>
</collection>
Python uses ElementTree to parse XML
xml.etree.ElementTreeis a module in the Python standard library for handling XML.
The following are some key concepts and usage of the xml.etree.ElementTree module:
ElementTree and Element objects:
- ElementTree: The ElementTree class is a tree representation of an XML document. It contains one or more Element objects and represents the entire XML document.
- Element: The Element object represents an element in an XML document. Each element has a tag, a set of attributes, and zero or more child elements.
Parsing XML
fromstring() method: You can use the fromstring() method to convert a string containing XML data into an Element object:
Example
xml_string = '<root><element>Some data</element></root>'
root = ET.fromstring(xml_string)
parse() method: If the XML data is stored in a file, you can use the parse() method to parse the entire XML document:
tree = ET.parse('example.xml')
root = tree.getroot()
Traversing the XML tree
find() method: Use the find() method to find the first child element with the specified tag:
title_element = root.find('title')findall() method: Use the findall() method to find all child elements with the specified tag:
book_elements = root.findall('book')
Accessing element attributes and text content
attribAttribute: You can access an element's attributes through the attrib attribute:
price = book_element.attrib['price']
textText content: You can access the element's text content through the text attribute:
title_text = title_element.text
Creating XML
Element() constructor: Use the Element() constructor to create a new element:
new_element = ET.Element('new_element')SubElement() function: Use the SubElement() function to create a child element with the specified tag:
new_sub_element = ET.SubElement(root, 'new_sub_element')
Modifying XML
Modify the element's attributes and text content: directly modify the element's attrib and text attributes.
Delete an element: You can delete an element using the remove() method:
root.remove(title_element)
Simple reading of XML content:
Example
# Define an XML string
xml_string = '''
<bookstore>
<book>
<title>Introduction to Python</title>
<author>John Doe</author>
<price>29.99</price>
</book>
<book>
<title>Data Science with Python</title>
<author>Jane Smith</author>
<price>39.95</price>
</book>
</bookstore>
'''
# Use ElementTree to parse the XML string
root = ET.fromstring(xml_string)
# Traverse the XML tree
for book in root.findall('book'):
title = book.find('title').text
author = book.find('author').text
price = book.find('price').text
print(f'Title: {title}, Author: {author}, Price: {price}')
The output after running the above code is:
Title: Introduction to Python, Author: John Doe, Price: 29.99 Title: Data Science with Python, Author: Jane Smith, Price: 39.95
Next, we will create an XML document containing book information, and then use this module to parse and manipulate it:
Example
# Create an XML document
root = ET.Element('bookstore')
# Add the first book
book1 = ET.SubElement(root, 'book')
title1 = ET.SubElement(book1, 'title')
title1.text = 'Introduction to Python'
author1 = ET.SubElement(book1, 'author')
author1.text = 'John Doe'
price1 = ET.SubElement(book1, 'price')
price1.text = '29.99'
# Add the second book
book2 = ET.SubElement(root, 'book')
title2 = ET.SubElement(book2, 'title')
title2.text = 'Data Science with Python'
author2 = ET.SubElement(book2, 'author')
author2.text = 'Jane Smith'
price2 = ET.SubElement(book2, 'price')
price2.text = '39.95'
# Save the XML document to a file
tree = ET.ElementTree(root)
tree.write('books.xml')
# Parse the XML document from the file
parsed_tree = ET.parse('books.xml')
parsed_root = parsed_tree.getroot()
# Traverse the XML tree and print book information
for book in parsed_root.findall('book'):
title = book.find('title').text
author = book.find('author').text
price = book.find('price').text
print(f'Title: {title}, Author: {author}, Price: {price}')
In the above example, we first create an XML document containing information about two books. Then we save this document to the file books.xml. Next, we use the ET.parse() method to parse the XML document in the file, traverse the document tree, and extract and print the title, author, and price of each book.
Python uses SAX to parse XML
SAX is an event-driven API.
Using SAX to parse XML documents involves two parts:ParserandEvent handler。
The parser is responsible for reading the XML document and sending events to the event handler, such as element start and element end events.
The event handler, in turn, is responsible for responding to events and processing the passed XML data.
- 1. Processing large files;
- 2. Only needing part of the file's content, or only needing to get specific information from the file.
- 3. When you want to build your own object model.
To process XML using the SAX approach in Python, you must first import the parse function from xml.sax, as well as ContentHandler from xml.sax.handler.
Introduction to ContentHandler class methods
characters(content) method
When it is called:
From the start of a line, before a tag is encountered, if characters exist, the value of content is those strings.
From one tag to before the next tag is encountered, if characters exist, the value of content is those strings.
From a tag to before the line terminator is encountered, if characters exist, the value of content is those strings.
Tags can be start tags or end tags.
startDocument() method
Called when the document starts.
endDocument() method
Called when the parser reaches the end of the document.
startElement(name, attrs) method
Called when an XML start tag is encountered. name is the name of the tag, and attrs is a dictionary of the tag's attribute values.
endElement(name) method
Called when an XML end tag is encountered.
make_parser method
The following method creates a new parser object and returns it.
xml.sax.make_parser( [parser_list] )
Parameter description:
- parser_list- Optional parameter, a list of parsers
parser method
The following method creates a SAX parser and parses an XML document:
xml.sax.parse( xmlfile, contenthandler[, errorhandler])
Parameter description:
- xmlfile- XML filename
- contenthandler- Must be a ContentHandler object
- errorhandler- If this parameter is specified, the errorhandler must be a SAX ErrorHandler object
parseString method
The parseString method creates an XML parser and parses the XML string:
xml.sax.parseString(xmlstring, contenthandler[, errorhandler])
Parameters:
- xmlstring- XML string
- contenthandler- Must be a ContentHandler object
- errorhandler- If this parameter is specified, errorhandler must be a SAX ErrorHandler object.
Python XML Parsing Examples
Example
import xml.sax
class MovieHandler( xml.sax.ContentHandler ):
def __init__(self):
self.CurrentData = ""
self.type = ""
self.format = ""
self.year = ""
self.rating = ""
self.stars = ""
self.description = ""
# Called when an element starts
def startElement(self, tag, attributes):
self.CurrentData = tag
if tag == "movie":
print ("*****Movie*****")
title = attributes["title"]
print ("Title:", title)
# Called when an element ends
def endElement(self, tag):
if self.CurrentData == "type":
print ("Type:", self.type)
elif self.CurrentData == "format":
print ("Format:", self.format)
elif self.CurrentData == "year":
print ("Year:", self.year)
elif self.CurrentData == "rating":
print ("Rating:", self.rating)
elif self.CurrentData == "stars":
print ("Stars:", self.stars)
elif self.CurrentData == "description":
print ("Description:", self.description)
self.CurrentData = ""
# Called when characters are read
def characters(self, content):
if self.CurrentData == "type":
self.type = content
elif self.CurrentData == "format":
self.format = content
elif self.CurrentData == "year":
self.year = content
elif self.CurrentData == "rating":
self.rating = content
elif self.CurrentData == "stars":
self.stars = content
elif self.CurrentData == "description":
self.description = content
if ( __name__ == "__main__"):
# Create an XMLReader
parser = xml.sax.make_parser()
# Disable namespaces
parser.setFeature(xml.sax.handler.feature_namespaces, 0)
# Override ContextHandler
Handler = MovieHandler()
parser.setContentHandler( Handler )
parser.parse("movies.xml")
The result of executing the above code is as follows:
*****Movie***** Title: Enemy Behind Type: War, Thriller Format: DVD Year: 2003 Rating: PG Stars: 10 Description: Talk about a US-Japan war *****Movie***** Title: Transformers Type: Anime, Science Fiction Format: DVD Year: 1989 Rating: R Stars: 8 Description: A schientific fiction *****Movie***** Title: Trigun Type: Anime, Action Format: DVD Rating: PG Stars: 10 Description: Vash the Stampede! *****Movie***** Title: Ishtar Type: Comedy Format: VHS Rating: PG Stars: 2 Description: Viewable boredom
Please refer to the complete SAX API documentation.Python SAX APIs
Parsing XML with xml.dom
The Document Object Model (DOM) is a standard programming interface recommended by the W3C organization for processing extensible markup languages.
A DOM parser, when parsing an XML document, reads the entire document at once, saves all elements in the document into a tree structure in memory. Then you can use the various functions provided by DOM to read or modify the document's content and structure, and also write the modified content to an XML file.
In Python, use xml.dom.minidom to parse XML files. An example is as follows:
Example
from xml.dom.minidom import parse
import xml.dom.minidom
# Open the XML document using the minidom parser
DOMTree = xml.dom.minidom.parse("movies.xml")
collection = DOMTree.documentElement
if collection.hasAttribute("shelf"):
print ("Root element : %s" % collection.getAttribute("shelf"))
# Get all movies from the collection
movies = collection.getElementsByTagName("movie")
# Print detailed information for each movie
for movie in movies:
print ("*****Movie*****")
if movie.hasAttribute("title"):
print ("Title: %s" % movie.getAttribute("title"))
type = movie.getElementsByTagName('type')[0]
print ("Type: %s" % type.childNodes[0].data)
format = movie.getElementsByTagName('format')[0]
print ("Format: %s" % format.childNodes[0].data)
rating = movie.getElementsByTagName('rating')[0]
print ("Rating: %s" % rating.childNodes[0].data)
description = movie.getElementsByTagName('description')[0]
print ("Description: %s" % description.childNodes[0].data)
The result of executing the above program is as follows:
Root element : New Arrivals *****Movie***** Title: Enemy Behind Type: War, Thriller Format: DVD Rating: PG Description: Talk about a US-Japan war *****Movie***** Title: Transformers Type: Anime, Science Fiction Format: DVD Rating: R Description: A schientific fiction *****Movie***** Title: Trigun Type: Anime, Action Format: DVD Rating: PG Description: Vash the Stampede! *****Movie***** Title: Ishtar Type: Comedy Format: VHS Rating: PG Description: Viewable boredom
Please refer to the complete DOM API documentation.Python DOM APIs。
Other Extensions