Perl Regular Expressions
A regular expression describes a pattern for string matching. It can be used to check whether a string contains a certain substring, replace the matched substring, or extract substrings that meet a certain condition from a string, etc.
Perl's regular expression functionality is very powerful, basically the most powerful among commonly used languages. Many languages refer to Perl's regular expressions when designing their regular expression support.
Perl has three forms of regular expressions: matching, substitution, and translation:
Match: m/
/ (can also be abbreviated as / /, omitting m) Substitution: s/
/ / Translation: tr/
/ /
These three forms are generally used with=~or!~in combination. =~ means a match, !~ means no match.
Match Operator
The match operator m// is used to match a string statement or a regular expression. For example, to match "run" in the scalar $bar, the code is as follows:
Example
Executing the above program produces the following output:
First match Second match
Pattern Matching Modifiers
There are some commonly used modifiers for pattern matching, as shown in the following table:
| Modifier | Description |
|---|---|
| i | Ignore case in the pattern |
| m | Multi-line mode |
| o | Assign only once |
| s | Single-line mode, "." matches "\n" (does not match by default) |
| x | Ignore whitespace in the pattern |
| g | Global match |
| cg | After a global match fails, allow searching for the matching string again |
Regular Expression Variables
After processing, Perl stores the matched values in three special variable names:
- $`: The string before the matched part
- $&: The matched string
- $': The remaining string that has not been matched
If you put these three variables together, you will get the original string.
The example is as follows:
Example
Executing the above program outputs the following result:
匹配前的字符串: welcome to 匹配的字符串: run 匹配后的字符串: oob site.
Substitution Operator
The substitution operator s/// is an extension of the match operator. It uses a new string to replace the specified string. The basic format is as follows:
s/PATTERN/REPLACEMENT/;
PATTERN is the match pattern, REPLACEMENT is the replacement string.
For example, we replace "google" with "example" in the following string:
Example
Executing the above program outputs the following result:
welcome to example site.
Substitution Operation Modifiers
The substitution operation modifiers are shown in the following table:
| Modifier | Description |
|---|---|
| i | If "i" is added to the modifier, the regex will become case-insensitive, that is, "a" and "A" are the same. |
| m | By default, the regex start "^" and end "$" only apply to the entire string. If "m" is added to the modifier, the start and end will refer to each line of the string: the beginning of each line is "^", and the end is "$". |
| o | The expression is executed only once. |
| s | If "s" is added to the modifier, then the default "." which represents any character except newline will become any character, including newline! |
| x | If this modifier is added, whitespace characters in the expression will be ignored unless they have been escaped. |
| g | Replace all matched strings. |
| e | Treat the replacement string as an expression. |
Translation Operator
The following are the modifiers related to the translation operator:
| Modifier | Description |
|---|---|
| c | Translate all unspecified characters |
| d | Delete all specified characters |
| s | Squash multiple identical output characters into one |
The following example converts all lowercase letters in the variable $string to uppercase letters:
Example
Executing the above program outputs the following result:
WELCOME TO EXAMPLE SITE.
The following example uses /s to delete repeated characters in the variable $string:
Example
Executing the above program outputs the following result:
runob
More examples:
$string =~ tr/\d/ /c; # 把所有非数字字符替换为空格 $string =~ tr/\t //d; # 删除tab和空格 $string =~ tr/0-9/ /cs # 把数字间的其它字符替换为一个空格。
More Regular Expression Rules
| Expression | Description |
|---|---|
| . | Matches any character except newline |
| x? | Matches 0 or 1 occurrence of the x string |
| x* | Matches 0 or more occurrences of the x string, but matches as few as possible |
| x+ | Matches 1 or more occurrences of the x string, but matches as few as possible |
| .* | Matches 0 or more occurrences of any character |
| .+ | Matches 1 or more occurrences of any character |
| {m} | Matches exactly m occurrences of the specified string |
| {m,n} | Matches between m and n occurrences of the specified string |
| {m,} | Matches m or more occurrences of the specified string |
| [] | Matches characters within [] |
| [^] | Matches characters not within [] |
| [0-9] | Matches all digit characters |
| [a-z] | Matches all lowercase letter characters |
| [^0-9] | Matches all non-digit characters |
| [^a-z] | Matches all non-lowercase letter characters |
| ^ | Matches at the beginning of a string |
| $ | Matches at the end of a string |
| \d | Matches a digit character, same syntax as [0-9] |
| \d+ | Matches one or more digit characters, same syntax as [0-9]+ |
| \D | Non-digit, otherwise same as \d |
| \D+ | Non-digit, otherwise same as \d+ |
| \w | A string of English letters or digits, same syntax as [a-zA-Z0-9_] |
| \w+ | Same syntax as [a-zA-Z0-9_]+ |
| \W | A string of non-English letters or non-digits, same syntax as [^a-zA-Z0-9_] |
| \W+ | Same syntax as [^a-zA-Z0-9_]+ |
| \s | Whitespace, same syntax as [\n\t\r\f] |
| \s+ | Same as [\n\t\r\f]+ |
| \S | Non-whitespace, same syntax as [^\n\t\r\f] |
| \S+ | Same syntax as [^\n\t\r\f]+ |
| \b | Matches strings with English letters or digits as boundaries |
| \B | Matches strings not bounded by English letters or digits |
| a|b|c | Matches a string that matches the character a, or b, or c |
| abc | Matches a string containing abc. (pattern) () This symbol remembers the found string; it is a very practical syntax. The string found in the first () becomes the variable $1 or \1, the string found in the second () becomes the variable $2 or \2, and so on. |
| /pattern/i | The i parameter means ignoring English case, that is, when matching strings, English case is not considered. \ If you need to find a special character in the pattern, such as "*", add a \ symbol before this character, so that the special character loses its effect. |
More References
Regular Expressions:https://www.example.com/regexp/regexp-tutorial.html
Perl Regular Expressions:https://perldoc.perl.org/perlre#Regular-Expressions
Other Extensions