Python Regular Expressions

A regular expression is a special sequence of characters that helps you easily check whether a string matches a certain pattern.

Python added the re module since version 1.5, which provides Perl-style regular expression patterns.

The re module gives the Python language all the functionality of regular expressions.

The compile function generates a regular expression object based on a pattern string and optional flag parameters. This object has a series of methods for regular expression matching and replacement.

The re module also provides functions that are fully consistent with the functionality of these methods. These functions use a pattern string as their first parameter.

This chapter mainly introduces the commonly used regular expression processing functions in Python.


re.match function

re.match attempts to match a pattern from the start of the string. If the match is not successful at the starting position, match() returns None.

Function syntax:

re.match(pattern, string, flags=0)

Function parameter description:

ParameterDescription
patternThe regular expression to match.
stringThe string to be matched.
flagsFlags, used to control the matching method of the regular expression, such as: case sensitivity, multiline matching, etc. See:Regular Expression Modifiers - Optional Flags

If the match is successful, the re.match method returns a match object, otherwise it returns None.

We can use group(num) or groups() match object functions to obtain the matched expression.

Match object methodsDescription
group(num=0)The string of the entire matched expression. group() can take multiple group numbers at once, in which case it returns a tuple containing the values corresponding to those groups.
groups()Returns a tuple containing all subgroup strings, from 1 to the number of subgroups contained.

Example

#!/usr/bin/python # -*- coding: UTF-8 -*- import re print(re.match('www', 'www.example.com').span()) # Match at the beginning print(re.match('com', 'www.example.com')) # Not match at the beginning

The output of the above example is:

(0, 3)
None

Example

#!/usr/bin/python import re line = "Cats are smarter than dogs" matchObj = re.match( r'(.*) are (.*?) .*', line, re.M|re.I) if matchObj: print "matchObj.group() : ", matchObj.group() print "matchObj.group(1) : ", matchObj.group(1) print "matchObj.group(2) : ", matchObj.group(2) else: print "No match!!"

The result of the above example is as follows:

matchObj.group() :  Cats are smarter than dogs
matchObj.group(1) :  Cats
matchObj.group(2) :  smarter

re.search method

re.search scans the entire string and returns the first successful match.

Function syntax:

re.search(pattern, string, flags=0)

Function parameter description:

ParameterDescription
patternThe regular expression to match.
stringThe string to be matched.
flagsFlags, used to control the matching method of the regular expression, such as: case sensitivity, multiline matching, etc.

If the match is successful, the re.search method returns a match object, otherwise it returns None.

We can use group(num) or groups() match object functions to obtain the matched expression.

Match object methodsDescription
group(num=0)The string of the entire matched expression. group() can take multiple group numbers at once, in which case it returns a tuple containing the values corresponding to those groups.
groups()Returns a tuple containing all subgroup strings, from 1 to the number of subgroups contained.

Example

#!/usr/bin/python # -*- coding: UTF-8 -*- import re print(re.search('www', 'www.example.com').span()) # Match at the beginning print(re.search('com', 'www.example.com').span()) # Not match at the beginning

The output of the above example is:

(0, 3)
(11, 14)

Example

#!/usr/bin/python import re line = "Cats are smarter than dogs"; searchObj = re.search( r'(.*) are (.*?) .*', line, re.M|re.I) if searchObj: print "searchObj.group() : ", searchObj.group() print "searchObj.group(1) : ", searchObj.group(1) print "searchObj.group(2) : ", searchObj.group(2) else: print "Nothing found!!"
The result of the above example is as follows:
searchObj.group() :  Cats are smarter than dogs
searchObj.group(1) :  Cats
searchObj.group(2) :  smarter

Difference between re.match and re.search

re.match only matches the beginning of the string. If the beginning of the string does not conform to the regular expression, the match fails and the function returns None; while re.search matches the entire string until a match is found.

Example

#!/usr/bin/python import re line = "Cats are smarter than dogs"; matchObj = re.match( r'dogs', line, re.M|re.I) if matchObj: print "match --> matchObj.group() : ", matchObj.group() else: print "No match!!" matchObj = re.search( r'dogs', line, re.M|re.I) if matchObj: print "search --> searchObj.group() : ", matchObj.group() else: print "No match!!"
The output of the above example is as follows:
No match!!
search --> searchObj.group() :  dogs

Search and Replace

Python's re module provides re.sub for replacing matches in a string.

Syntax:

re.sub(pattern, repl, string, count=0, flags=0)

Parameters:

  • pattern: The pattern string in the regular expression.
  • repl: The replacement string, can also be a function.
  • string: The original string to be searched and replaced.
  • count: The maximum number of replacements after pattern matching, default 0 means replace all matches.

Example

#!/usr/bin/python # -*- coding: UTF-8 -*- import re phone = "2004-959-559 # This is a foreign phone number" # Delete Python comments in the string num = re.sub(r'#.*$', "", phone) print "The phone number is:", num # Delete non-digit (-) strings num = re.sub(r'\D', "", phone) print "The phone number is:", num
The result of the above example is as follows:
电话号码是:  2004-959-559 
电话号码是 :  2004959559

repl parameter is a function

In the following example, the matched numbers in the string are multiplied by 2:

Example

#!/usr/bin/python # -*- coding: UTF-8 -*- import re # Multiply the matched numbers by 2 def double(matched): value = int(matched.group('value')) return str(value * 2) s = 'A23G4HFD567' print(re.sub('(?P<value>\d+)', double, s))

The output after execution is:

A46G8HFD1134

re.compile function

The compile function is used to compile a regular expression and generate a regular expression (Pattern) object, which can be used by functions such as match(), search(), and findall.

The syntax format is:

re.compile(pattern[, flags])

Parameters:

  • pattern: A regular expression in string form

  • flags: Optional, representing the matching mode, such as ignoring case, multiline mode, etc. The specific parameters are:

    1. re.IIgnore case
    2. re.LIndicates that the special character sets \w, \W, \b, \B, \s, \S depend on the current environment
    3. re.MMultiline mode
    4. re.SThat is.And any character including newline (.Excluding newline)
    5. re.UIndicates that the special character sets \w, \W, \b, \B, \d, \D, \s, \S depend on the Unicode character property database
    6. re.XTo increase readability, ignore spaces and#Comments after

Example

Example

>>>import re >>> pattern = re.compile(r'\d+') # Used to match at least one digit >>> m = pattern.match('one12twothree34four') # Search the beginning, no match >>> print m None >>> m = pattern.match('one12twothree34four', 2, 10) # Start matching from the 'e' position, no match >>> print m None >>> m = pattern.match('one12twothree34four', 3, 10) # Start matching from the '1' position, matches exactly >>> print m # Returns a Match object <_sre.SRE_Match object at 0x10a42aac0> >>> m.group(0) # 0 can be omitted '12' >>> m.start(0) # 0 can be omitted 3 >>> m.end(0) # 0 can be omitted 5 >>> m.span(0) # 0 can be omitted (3, 5)

Above, upon a successful match, a Match object is returned, where:

  • group([group1, …])The method is used to obtain the string matched by one or more groups. When you want to obtain the entire matched substring, you can directly usegroup()orgroup(0);
  • start([group])The method is used to get the starting position of the substring matched by a group in the entire string (the index of the first character of the substring), with a default parameter value of 0;
  • end([group])The method is used to get the ending position of the substring matched by a group in the entire string (the index of the last character of the substring + 1), with a default parameter value of 0;
  • span([group])The method returns(start(group), end(group))。

Let's look at another example:

Example

>>>import re >>> pattern = re.compile(r'([a-z]+) ([a-z]+)', re.I) # re.I means ignore case >>> m = pattern.match('Hello World Wide Web') >>> print m # Match successful, returns a Match object <_sre.SRE_Match object at 0x10bea83e8> >>> m.group(0) # Returns the entire matched substring 'Hello World' >>> m.span(0) # Returns the index of the entire matched substring (0, 11) >>> m.group(1) # Returns the substring matched by the first group 'Hello' >>> m.span(1) # Returns the index of the substring matched by the first group (0, 5) >>> m.group(2) # Returns the substring matched by the second group 'World' >>> m.span(2) # Returns the substring matched by the second group (6, 11) >>> m.groups() # Equivalent to (m.group(1), m.group(2), ...) ('Hello', 'World') >>> m.group(3) # No third group exists Traceback (most recent call last): File "<stdin>", line 1, in <module> IndexError: no such group

findall

Finds all substrings matched by the regular expression in the string and returns a list. If there are multiple matching patterns, it returns a list of tuples. If no matches are found, it returns an empty list.

Note:match and search match once, while findall matches all.

The syntax format is:

findall(string[, pos[, endpos]])

Parameters:

  • string: The string to be matched.
  • pos: Optional parameter, specifies the starting position of the string, defaults to 0.
  • endpos: Optional parameter, specifies the ending position of the string, defaults to the length of the string.

Find all numbers in the string:

Example

# -*- coding:UTF8 -*- import re pattern = re.compile(r'\d+') # Find numbers result1 = pattern.findall('example 123 google 456') result2 = pattern.findall('run88oob123google456', 0, 10) print(result1) print(result2)

Output result:

['123', '456']
['88', '12']

Multiple matching patterns, returns a list of tuples:

Example

import re

result = re.findall(r'(\w+)=(\d+)', 'set width=20 and height=10')
print(result)
[('width', '20'), ('height', '10')]

re.finditer

Similar to findall, finds all substrings matched by the regular expression in the string and returns them as an iterator.

re.finditer(pattern, string, flags=0)

Parameters:

ParameterDescription
patternThe regular expression to match.
stringThe string to be matched.
flagsFlags, used to control the matching behavior of the regular expression, such as case sensitivity, multiline matching, etc. See:Regular expression modifiers - optional flags

Example

# -*- coding: UTF-8 -*- import re it = re.finditer(r"\d+","12a32bc43jf3") for match in it: print (match.group() )

Output result:

12 
32 
43 
3

re.split

The split method splits the string according to the substrings that can be matched and returns a list. Its usage form is as follows:

re.split(pattern, string[, maxsplit=0, flags=0])

Parameters:

ParameterDescription
patternThe regular expression to match.
stringThe string to be matched.
maxsplitMaximum number of splits. maxsplit=1 means split once. Default is 0, which means no limit.
flagsFlags, used to control the matching behavior of the regular expression, such as case sensitivity, multiline matching, etc. See:Regular expression modifiers - optional flags

Example

>>>import re >>> re.split('\W+', 'example, example, example.') ['example', 'example', 'example', ''] >>> re.split('(\W+)', ' example, example, example.') ['', ' ', 'example', ', ', 'example', ', ', 'example', '.', ''] >>> re.split('\W+', ' example, example, example.', 1) ['', 'example, example, example.'] >>> re.split('a*', 'hello world') # For a string that cannot find a match, split will not split it ['hello world']

Regular Expression Object

re.RegexObject

re.compile() returns a RegexObject object.

re.MatchObject

group() returns the string matched by the RE.

  • start()Returns the starting position of the match
  • end()Returns the ending position of the match
  • span()Returns a tuple containing the (start, end) positions of the match

Regular Expression Modifiers - Optional Flags

Regular expressions can contain optional flag modifiers to control the matching pattern. Modifiers are specified as an optional flag. Multiple flags can be specified by bitwise ORing (|) them. For example, re.I | re.M sets both the I and M flags:

ModifierDescription
re.IMakes the matching case-insensitive
re.LPerforms locale-aware matching
re.MMulti-line matching, affects ^ and $
re.SMakes . match all characters including newline
re.UParses characters according to the Unicode character set. This flag affects \w, \W, \b, \B.
re.XThis flag gives you a more flexible format so you can write regular expressions that are easier to understand.

Regular Expression Patterns

Pattern strings use special syntax to represent a regular expression:

Letters and numbers represent themselves. Letters and numbers in a regular expression pattern match the same strings.

Most letters and numbers have different meanings when preceded by a backslash.

Punctuation characters match themselves only when escaped; otherwise they represent special meanings.

The backslash itself needs to be escaped with a backslash.

Since regular expressions often contain backslashes, you'd better use raw strings to represent them. Pattern elements (such as r'\t', equivalent to '\\t') match the corresponding special characters.

The following table lists the special elements in regular expression pattern syntax. If you provide optional flag parameters while using a pattern, the meaning of certain pattern elements may change.

PatternDescription
^Matches the beginning of the string.
$Matches the end of the string.
.Matches any character except newline. When the re.DOTALL flag is specified, it can match any character including newline.
[...]Used to represent a set of characters, listed individually: [amk] matches 'a', 'm', or 'k'
[^...]Characters not in []: [^abc] matches any character except a, b, c.
re*Matches 0 or more occurrences of the expression.
re+Matches 1 or more occurrences of the expression.
re?Matches 0 or 1 occurrences of the fragment defined by the preceding regular expression, non-greedy way
re{ n}Exactly matches n occurrences of the preceding expression. For example,o{2}It cannot match the 'o' in "Bob", but it can match the two o's in "food".
re{ n,}Matches n or more occurrences of the preceding expression. For example, o{2,} cannot match the 'o' in "Bob", but it can match all the o's in "foooood". "o{1,}" is equivalent to "o+". "o{0,}" is equivalent to "o*".
re{ n, m}Matches n to m occurrences of the fragment defined by the preceding regular expression, greedy way
a| bMatches a or b
(re)Groups the regular expression and remembers the matched text
(?imx)Regular expression contains three optional flags: i, m, or x. Only affects the area within the parentheses.
(?-imx)Regular expression turns off i, m, or x optional flags. Only affects the area within the parentheses.
(?: re)Similar to (...), but does not represent a group
(?imx: re)Uses i, m, or x optional flags inside the parentheses
(?-imx: re)Does not use i, m, or x optional flags inside the parentheses
(?#...)Comment.
(?= re)Positive lookahead assertion. If the contained regular expression, represented by ..., matches successfully at the current position, it succeeds; otherwise it fails. But once the contained expression has been attempted, the matching engine does not advance; the rest of the pattern still has to try to the right of the assertion.
(?! re)Negative lookahead assertion. Opposite of the positive assertion; succeeds when the contained expression cannot match at the current position in the string.
(?> re)Matches an independent pattern, eliminating backtracking.
\wMatches letters, digits, and underscores
\WMatches anything except letters, digits, and underscores
\sMatches any whitespace character, equivalent to[ \t\n\r\f]。
\SMatches any non-whitespace character
\dMatches any digit, equivalent to [0-9].
\DMatches any non-digit
\AMatches the start of the string
\ZMatches the end of the string. If there is a newline, it only matches up to the end of the string before the newline.
\zMatches the end of the string
\GMatches the position where the last match completed.
\bMatches a word boundary, that is, the position between a word and a space. For example, 'er\b' can match the 'er' in "never", but cannot match the 'er' in "verb".
\BMatches a non-word boundary. 'er\B' can match the 'er' in "verb", but cannot match the 'er' in "never".
\n, \t, etc.Matches a newline character. Matches a tab character. etc.
\1...\9Matches the content of the nth group.
\10Matches the content of the nth group if it has been matched. Otherwise it refers to an octal character code expression.

Regular Expression Examples

Character matching

ExampleDescription
pythonMatches "python".

Character classes

ExampleDescription
[Pp]ython Matches "Python" or "python"
rub[ye]Matches "ruby" or "rube"
[aeiou]Matches any one letter in the brackets
[0-9]Matches any digit. Similar to
[a-z]Matches any lowercase letter
[A-Z]Matches any uppercase letter
[a-zA-Z0-9]Matches any letter and digit
[^aeiou]All characters except the letters aeiou
[^0-9]Matches characters other than digits

Special character classes

ExampleDescription
.Matches any single character except "\n". To match any character including '\n', use a pattern like '[.\n]'.
\dMatches a digit character. Equivalent to [0-9].
\D Matches a non-digit character. Equivalent to [^0-9].
\sMatches any whitespace character, including space, tab, form feed, etc. Equivalent to [ \f\n\r\t\v].
\S Matches any non-whitespace character. Equivalent to [^ \f\n\r\t\v].
\wMatches any word character including underscore. Equivalent to '[A-Za-z0-9_]'.
\WMatches any non-word character. Equivalent to '[^A-Za-z0-9_]'.
Other extensions