Regular Expressions -Matching Rules
Basic Pattern Matching
It all starts with the basics. Patterns are the most fundamental elements of regular expressions; they are sets of characters describing string characteristics. Patterns can be very simple, consisting of ordinary strings, or very complex, often using special characters to represent a range of characters, repetition, or context. For example:
^once
This pattern contains a special character^, which means the pattern only matches strings that begin withonce. For example, the pattern matches the string"once upon a time", and does not match"There once was a man from NewYork". Just as the^symbol denotes the beginning, the$symbol is used to match strings that end with a given pattern.
bucket$
This pattern matches"Who kept all of this cash in a bucket", and does not match"buckets". When the characters^and$are used together, they indicate an exact match (the string is exactly the same as the pattern). For example:
^bucket$
only matches the string"bucket". If a pattern does not include^and$, it matches any string that contains the pattern. For example, the pattern:
once
matches the string
There once was a man from NewYork Who kept all of his cash in a bucket.
.
In this pattern, the letters(o-n-c-e)are literal characters; that is, they represent the letters themselves, and the same is true for numbers. Some other slightly more complex characters, such as punctuation and whitespace (spaces, tabs, etc.), require escape sequences. All escape sequences begin with a backslash\. The escape sequence for a tab is\t. So if we want to test whether a string begins with a tab, we can use this pattern:
^\t
Similarly, use\nto represent"newline",\rto represent carriage return. For other special symbols, you can put a backslash in front of them; for example, the backslash itself is represented by\\, the period by.use\., and so on.
Character Clusters
In Internet programs, regular expressions are often used to validate user input. After a user submits a FORM, to determine whether the entered phone number, address, EMAIL address, credit card number, etc. are valid, ordinary literal-based characters are not sufficient.
So a more flexible way to describe the pattern we want is needed, and that is the character cluster. To create a character cluster representing all vowel characters, put all vowel characters in square brackets:
[AaEeIiOoUu]
This pattern matches any vowel character, but can only represent one character. A hyphen can be used to represent a range of characters, such as:
[a-z] // 匹配所有的小写字母 [A-Z] // 匹配所有的大写字母 [a-zA-Z] // 匹配所有的字母 [0-9] // 匹配所有的数字 [0-9\.\-] // 匹配所有的数字,句号和减号 [ \f\r\t\n] // 匹配所有的白字符
Similarly, these also represent only one character, which is a very important point. To match a string consisting of one lowercase letter and one digit, such as "z2", "t6" or "g7", but not "ab2", "r2d3" or "b52", use this pattern:
^[a-z][0-9]$
Although[a-z]represents the range of 26 letters, here it can only match strings whose first character is a lowercase letter.
Earlier it was mentioned that ^ represents the beginning of a string, but it has another meaning. When used within a set of square brackets,^it means "notor "exclude", and is often used to eliminate a character. Using the previous example, we require that the first character cannot be a digit:
^[^0-9][0-9]$
This pattern matches "&5", "g7" and "-2", but does not match "12" or "66". The following are a few examples of excluding specific characters:
[^a-z] //除了小写字母以外的所有字符
[^\\\/\^] //除了(\)(/)(^)之外的所有字符
[^\"\'] //除了双引号(")和单引号(')之外的所有字符
Special Characters.(dot, period) is used in regular expressions to represent all characters except "newline". So the pattern^.5$matches any two-character string ending with the digit 5 and starting with any other non-newline character. The pattern.can match any string,except newline characters (\n, \r)。
PHP regular expressions have some built-in general character clusters, listed as follows:
| Character Cluster | Description |
|---|---|
| [[:alpha:]] | Any letter |
| [[:digit:]] | Any digit |
| [[:alnum:]] | Any letter or digit |
| [[:space:]] | Any whitespace character |
| [[:upper:]] | Any uppercase letter |
| [[:lower:]] | Any lowercase letter |
| [[:punct:]] | Any punctuation mark |
| [[:xdigit:]] | Any hexadecimal digit, equivalent to [0-9a-fA-F] |
Determining Repeated Occurrences
Up to now, you know how to match a letter or digit, but more often you may need to match a word or a group of digits. A word is composed of several letters, and a group of digits is composed of several single digits. Braces ({}) following a character or character cluster determine the number of repeated occurrences of the preceding content.
| Character Cluster | Description |
|---|---|
| ^[a-zA-Z_]$ | All letters and underscores |
| ^[[:alpha:]]{3}$ | All 3-letter words |
| ^a$ | The letter a |
| ^a{4}$ | aaaa |
| ^a{2,4}$ | aa, aaa, or aaaa |
| ^a{1,3}$ | a, aa, or aaa |
| ^a{2,}$ | Strings containing more than two a's |
| ^a{2,} | e.g., aardvark and aaab, but not apple |
| a{2,} | e.g., baad and aaa, but not Nantucket |
| \t{2} | Two tab characters |
| .{2} | All two-character strings |
These examples describe three different uses of braces. A number{x} meansthe preceding character or character cluster appears exactly x times; a number followed by a comma{x,}meansthe preceding content appears x or more times; two numbers separated by a comma{x,y}meansthe preceding content appears at least x times, but no more than y times. We can extend the pattern to more words or digits:
^[a-zA-Z0-9_]{1,}$ // 所有包含一个以上的字母、数字或下划线的字符串
^[1-9][0-9]{0,}$ // 所有的正整数
^\-{0,1}[0-9]{1,}$ // 所有的整数
^[-]?[0-9]+\.?[0-9]+$ // 所有的浮点数
The last example is a bit hard to understand, isn't it? Look at it this way: start with an optional minus sign ([-]?) at the beginning (^), followed by one or more digits ([0-9]+), and a decimal point (\.), then followed by one or more digits([0-9]+), and nothing else after it ($). Below you will learn a simpler method that can be used.
Special Characters?and{0,1}are equal; they both mean:0 or 1 of the preceding contentorthe preceding content is optional. So the previous example can be simplified to:
^\-?[0-9]{1,}\.?[0-9]{1,}$
Special Characters*and{0,}are equal; they both mean0 or more of the preceding content. Finally, the characters+and{1,}are equal and mean1 or more of the preceding content, so the four examples above can be written as:
^[a-zA-Z0-9_]+$ // 所有包含一个以上的字母、数字或下划线的字符串 ^[1-9][0-9]*$ // 所有的正整数 ^\-?[0-9]+$ // 所有的整数 ^[-]?[0-9]+(\.[0-9]+)?$ // 所有的浮点数
Of course, this does not technically reduce the complexity of regular expressions, but it can make them easier to read.
Other Extensions