Java Regular Expressions
Regular expressions define patterns for strings.
Regular expressions can be used to search, edit, or process text.
Regular expressions are not limited to a particular language, but there are subtle differences in each language.
Java provides the java.util.regex package, which contains the Pattern and Matcher classes for handling regular expression matching operations.
Regular Expression Examples
A string is actually a simple regular expression. For example,Hello WorldThe regular expression matches the string "Hello World".
.(dot) is also a regular expression. It matches any single character such as "a" or "1".
The following table lists some regular expression examples and descriptions:
| Regular Expression | Description |
|---|---|
this is text |
Matches the string "this is text" |
this\s+is\s+text |
Note the ... in the string\s+。 Matches the ... after the word "this"\s+It can match multiple spaces, then match the is string, and then\s+match multiple spaces and then follow with the text string. It can match this example: this is text |
^\d+(\.\d+)? |
^ defines the beginning \d+ matches one or more digits ? sets the option inside the parentheses as optional \. matches "." Examples that can be matched: "5", "1.5", and "2.21". |
For more regular expression content, you can refer to:Regular Expressions - Tutorial
java.util.regex Package
java.util.regexThe package is a package in the Java standard library used to support regular expression operations.
The java.util.regex package mainly includes the following three classes:
-
Pattern Class:
A pattern object is a compiled representation of a regular expression. The Pattern class has no public constructor. To create a Pattern object, you must first call its public static compile method, which returns a Pattern object. This method accepts a regular expression as its first argument.
-
Matcher Class:
A Matcher object is the engine that interprets and performs matching operations on an input string. Like the Pattern class, Matcher also has no public constructor. You need to call the matcher method of a Pattern object to obtain a Matcher object.
-
PatternSyntaxException:
PatternSyntaxException is an unchecked exception class that indicates a syntax error in a regular expression pattern.
The following example uses a regular expression.*example.*to find whether the string containsexamplethe substring:
Example
The example output is:
Capturing Groups
Capturing groups are a way to treat multiple characters as a single unit. They are created by grouping the characters inside parentheses.
For example, the regular expression (dog) creates a single group containing "d", "o", and "g".
Capturing groups are numbered by counting their opening parentheses from left to right. For example, in the expression ((A)(B(C))), there are four such groups:
- ((A)(B(C)))
- (A)
- (B(C))
- (C)
You can use the groupCount method of a matcher object to see how many groups the expression has. The groupCount method returns an int value indicating how many capturing groups the matcher object currently has.
There is also a special group (group(0)) that always represents the entire expression. This group is not included in the return value of groupCount.
Example
The following example illustrates how to find digit strings from a given string:
RegexMatches.java file code:
The above example compiles and runs with the following result:
Found value: This order was placed for QT3000! OK? Found value: This order was placed for QT Found value: 3000 Found value: ! OK?
Regular Expression Syntax
In other languages,\\it means:I want to insert an ordinary (literal) backslash into the regular expression; please do not give it any special meaning.
In Java,\\it means:I want to insert a regular expression backslash, so the character after it has special meaning.
Therefore, in other languages (such as Perl), a single backslash\is enough to have an escaping effect, whereas in Java regular expressions, two backslashes are needed to be parsed as the escaping effect in other languages. It can also be simply understood that in Java regular expressions, two\\represent one in other languages\, which is why the regular expression for a single digit is\\d, and a normal backslash is represented by\\。
System.out.print("\\"); // 输出为 \
System.out.print("\\\\"); // 输出为 \\
Character | Description |
|---|---|
\ | Marks the next character as a special character, a literal, a backreference, or an octal escape. For example,nmatches the charactern。\nmatches a newline. The sequence\\\\matches\\ ,\\(matches(。 |
^ | Matches the position at the beginning of the input string. If theRegExpobject'sMultilineproperty is set, ^ will also match the position after "\n" or "\r". |
$ | Matches the position at the end of the input string. If theRegExpobject'sMultilineproperty is set, $ will also match the position before "\n" or "\r". |
* | Matches the preceding character or subexpression zero or more times. For example, zo* matches "z" and "zoo". * is equivalent to {0,}. |
+ | Matches the preceding character or subexpression one or more times. For example, "zo+" matches "zo" and "zoo", but does not match "z". + is equivalent to {1,}. |
? | Matches the preceding character or subexpression zero or one time. For example, "do(es)?" matches "do" or the "do" in "does". ? is equivalent to {0,1}. |
{n} | n is a nonnegative integer. Match exactlyntimes. For example, "o{2}" does not match the "o" in "Bob", but matches the two "o"s in "food". |
{n,} | n is a nonnegative integer. Match at leastn times. For example, "o{2,}" does not match the "o" in "Bob", but matches all the o's in "foooood". "o{1,}" is equivalent to "o+". "o{0,}" is equivalent to "o*". |
{n,m} | mandnis a nonnegative integer, wheren <= m. Match at leastntimes, at mostmtimes. For example, "o{1,3}" matches the first three o's in "fooooood". 'o{0,1}' is equivalent to 'o?'. Note: You cannot insert spaces between the comma and the digits. |
? | When this character immediately follows any other quantifier (*, +, ?, {n}、{n,}、{n,m}) after, the matching mode is "non-greedy". A "non-greedy" mode matches the shortest possible string that is searched, while the default "greedy" mode matches the longest possible string that is searched. For example, in the string "oooo", "o+?" matches only a single "o", while "o+" matches all "o"s. |
. | Matches any single character except "\r\n". To match any character including "\r\n", use a pattern such as "[\s\S]". |
(pattern) | Matchespatternand captures the matched subexpression. You can use$0…$9the property to retrieve the captured match from the resulting "match" collection. To match the parenthesis character ( ), use "\(" or "\)". |
(?:pattern) | MatchespatternBut a subexpression that does not capture the match, i.e., it is a non-capturing match and does not store the match for later use. This is useful when combining pattern parts with the "or" character (|). For example, 'industr(?:y|ies)' is a more economical expression than 'industry|industries'. |
(?=pattern) | A subexpression that performs a positive lookahead search, which matchespatternthe string at the starting point of the matched string. It is a non-capturing match, i.e., it cannot capture the match for later use. For example, 'Windows (?=95|98|NT|2000)' matches "Windows" in "Windows 2000", but does not match "Windows" in "Windows 3.1". Lookahead does not consume characters; that is, after a match occurs, the search for the next match begins immediately after the previous match, not after the characters that make up the lookahead. |
(?!pattern) | A subexpression that performs a negative lookahead search, which matches a string that is not at the starting point of apatternmatched search string. It is a non-capturing match, i.e., it cannot capture the match for later use. For example, 'Windows (?!95|98|NT|2000)' matches "Windows" in "Windows 3.1", but does not match "Windows" in "Windows 2000". Lookahead does not consume characters; that is, after a match occurs, the search for the next match begins immediately after the previous match, not after the characters that make up the lookahead. |
x|y | Matchesxory. For example, 'z|food' matches "z" or "food". '(z|f)ood' matches "zood" or "food". |
[xyz] | Character set. Matches any character included. For example, "[abc]" matches "a" in "plain". |
[^xyz] | Negative character set. Matches any character not included. For example, "[^abc]" matches "p", "l", "i", "n" in "plain". |
[a-z] | Character range. Matches any character in the specified range. For example, "[a-z]" matches any lowercase letter in the range "a" to "z". |
[^a-z] | Negative range character. Matches any character not in the specified range. For example, "[^a-z]" matches any character not in the range "a" to "z". |
\b | Matches a word boundary, i.e., the position between a word and a space. For example, "er\b" matches "er" in "never", but does not match "er" in "verb". |
\B | Non-word boundary match. "er\B" matches "er" in "verb", but does not match "er" in "never". |
\cx | Matchesxthe indicated control character. For example, \cM matches Control-M or carriage return.xThe value must be between A-Z or a-z. If not, then c is assumed to be the character "c" itself. |
\d | Digit character match. Equivalent to [0-9]. |
\D | Non-digit character match. Equivalent to [^0-9]. |
\f | Form-feed character match. Equivalent to \x0c and \cL. |
\n | Newline character match. Equivalent to \x0a and \cJ. |
\r | Matches a carriage return. Equivalent to \x0d and \cM. |
\s | Matches any whitespace character, including space, tab, form feed, etc. Equivalent to [ \f\n\r\t\v]. |
\S | Matches any non-whitespace character. Equivalent to [^ \f\n\r\t\v]. |
\t | Tab character match. Equivalent to \x09 and \cI. |
\v | Vertical tab character match. Equivalent to \x0b and \cK. |
\w | Matches any word character, including underscore. Equivalent to "[A-Za-z0-9_]". |
\W | Matches any non-word character. Equivalent to "[^A-Za-z0-9_]". |
\xn | Matchesn, where thenis a hexadecimal escape code. The hexadecimal escape code must be exactly two digits long. For example, "\x41" matches "A". "\x041" is equivalent to "\x04"&"1". ASCII codes are allowed in regular expressions. |
\num | Matchesnum, where thenumis a positive integer. A backreference to a captured match. For example, "(.)\1" matches two consecutive identical characters. |
\n | Identifies an octal escape code or a backreference. If \nis preceded by at leastncapturing subexpressions, thennis a backreference. Otherwise, ifnis an octal digit (0-7), thennis an octal escape code. |
\nm | Identifies an octal escape code or a backreference. If \nmis preceded by at leastnmcapturing subexpressions, thennmis a backreference. If \nmis preceded by at leastncaptures, thennis a backreference, followed by the characterm. If neither of the preceding cases exists, then \nmmatches the octal valuenm, wheren andmis an octal digit (0-7). |
\nml | Whennis an octal digit (0-3),mandlwhen it is an octal digit (0-7), matches the octal escape codenml。 |
\un | Matchesn, wherenis a Unicode character represented by a four-digit hexadecimal number. For example, \u00A9 matches the copyright symbol (©). |
According to the requirements of the Java Language Specification, backslashes in strings in Java source code are interpreted as Unicode escapes or other character escapes. Therefore, two backslashes must be used in string literals to protect the regular expression from being interpreted by the Java bytecode compiler. For example, when interpreted as a regular expression, the string literal "\b" matches a single backspace character, while "\\b" matches a word boundary. The string literal "\(hello\)" is illegal and will cause a compile-time error; to match the string (hello), the string literal "\\(hello\\)" must be used.
Methods of the Matcher Class
Index Methods
Index methods provide useful index values that precisely indicate where a match can be found in the input string:
| No. | Method and description |
|---|---|
| 1 |
public int start() Returns the initial index of the previous match. |
| 2 |
public int start(int group) Returns the initial index of the subsequence captured by the given group during the previous match operation. |
| 3 |
public int end() Returns the offset after the last matched character. |
| 4 |
public int end(int group) Returns the offset after the last character of the subsequence captured by the given group during the previous match operation. |
Study Methods
Search methods are used to check the input string and return a boolean value indicating whether the pattern is found:
| No. | Method and description |
|---|---|
| 1 |
public boolean lookingAt() Attempts to match the input sequence starting at the beginning of the region to the pattern. |
| 2 |
public boolean find() Attempts to find the next subsequence of the input sequence that matches the pattern. |
| 3 |
public boolean find(int start) Resets this matcher and then attempts to find the next subsequence of the input sequence that matches the pattern, starting at the specified index. |
| 4 |
public boolean matches() Attempts to match the entire region to the pattern. |
Replacement Methods
Replacement methods are methods for replacing text in the input string:
| No. | Method and description |
|---|---|
| 1 |
public Matcher appendReplacement(StringBuffer sb, String replacement) Implements a non-terminal append-and-replace step. |
| 2 |
public StringBuffer appendTail(StringBuffer sb) Implements a terminal append-and-replace step. |
| 3 |
public String replaceAll(String replacement) Replaces every subsequence of the input sequence that matches the pattern with the given replacement string. |
| 4 |
public String replaceFirst(String replacement) Replaces the first subsequence of the input sequence that matches the pattern with the given replacement string. |
| 5 |
public static String quoteReplacement(String s) Returns the literal replacement string for the specified string. This method returns a string that works as if it were a literal string passed to the appendReplacement method of the Matcher class. |
start and end Methods
The following is an example that counts the number of times the word "cat" appears in the input string:
RegexMatches.java file code:
The above example compiles and runs with the following result:
Match number 1 start(): 0 end(): 3 Match number 2 start(): 4 end(): 7 Match number 3 start(): 8 end(): 11 Match number 4 start(): 19 end(): 22
It can be seen that this example uses word boundaries to ensure that the letters "c", "a", "t" are not merely a substring of a longer word. It also provides some useful information about where matches occur in the input string.
The start method returns the initial index of the subsequence captured by the given group during the previous match operation, and the end method returns the index of the last matched character plus 1.
matches and lookingAt Methods
Both the matches and lookingAt methods are used to attempt to match an input sequence against a pattern. The difference is that matches requires the entire sequence to match, while lookingAt does not.
The lookingAt method does not require the entire string to match, but it does require matching from the first character.
These two methods are often used at the beginning of the input string.
We use the following example to explain this functionality:
RegexMatches.java file code:
The above example compiles and runs with the following result:
Current REGEX is: foo Current INPUT is: fooooooooooooooooo Current INPUT2 is: ooooofoooooooooooo lookingAt(): true matches(): false lookingAt(): false
replaceFirst and replaceAll Methods
The replaceFirst and replaceAll methods are used to replace text that matches the regular expression. The difference is that replaceFirst replaces the first match, while replaceAll replaces all matches.
The following example explains this functionality:
RegexMatches.java file code:
The result of compiling and running the above example is:
The cat says meow. All cats say meow.
appendReplacement and appendTail Methods
The Matcher class also provides the appendReplacement and appendTail methods for text replacement:
See the following example to explain this functionality:
RegexMatches.java file code:
The result of compiling and running the above example is:
-foo-foo-foo-kkk
Methods of the PatternSyntaxException Class
PatternSyntaxException is a non-checked exception class that indicates a syntax error in a regular expression pattern.
The PatternSyntaxException class provides the following methods to help us see what error occurred.
| No. | Method and Description |
|---|---|
| 1 |
public String getDescription() Gets the description of the error. |
| 2 |
public int getIndex() Gets the index of the error. |
| 3 |
public String getPattern() Gets the erroneous regular expression pattern. |
| 4 |
public String getMessage() Returns a multi-line string containing the description of the syntax error and its index, the erroneous regular expression pattern, and a visual indication of the error index in the pattern. |