Ruby Regular Expressions
Regular Expressionsis a special sequence of characters that matches or searches a collection of strings by using a pattern with specialized syntax.
Regular expressions use a predefined set of specific characters and combinations of these characters to form a "rule string", which is used to express a filtering logic for strings.
Syntax
Regular ExpressionsLiterally, it is a pattern between slashes or between arbitrary delimiters following %r, as shown below:
Examples
Try it »
The output of the above example is:
Line1 contains Cats
Regular Expression Modifiers
A regular expression literally may contain an optional modifier to control various aspects of matching. The modifier is specified after the second slash character, as shown in the example above. The following table lists the possible modifiers:
| Modifier | Description |
|---|---|
| i | Ignore case when matching text. |
| o | Perform #{} interpolation only once; the regular expression is evaluated the first time. |
| x | Ignore spaces, allowing whitespace and comments to be placed within the entire expression. |
| m | Match multiple lines, recognizing newline characters as normal characters. |
| u,e,s,n | Interpret the regular expression as Unicode (UTF-8), EUC, SJIS, or ASCII. If no modifier is specified, the regular expression is considered to use the source encoding. |
Just as strings are delimited by %Q, Ruby allows you to use %r as the start of a regular expression, followed by any delimiter. This is very useful when describing patterns that contain many slash characters you don't want to escape.
Regular Expression Patterns
Except for control characters,(+ ? . * ^ $ ( ) [ ] { } | \)all other characters match themselves. You can escape control characters by placing a backslash before them.
The following table lists the regular expression syntax available in Ruby.
| Pattern | Description |
|---|---|
| ^ | Matches the beginning of a line. |
| $ | Matches the end of a line. |
| . | Matches any single character except newline. With the m option, it also matches newline. |
| [...] | Matches any single character in square brackets. |
| [^...] | Matches any single character not in square brackets. |
| re* | Matches the preceding subexpression zero or more times. |
| re+ | Matches the preceding subexpression one or more times. |
| re? | Matches the preceding subexpression zero or one time. |
| re{ n} | Matches the preceding subexpression exactly n times. |
| re{ n,} | Matches the preceding subexpression n times or more. |
| re{ n, m} | Matches the preceding subexpression at least n times and at most m times. |
| a| b | Matches a or b. |
| (re) | Groups a regular expression and remembers the matched text. |
| (?imx) | Temporarily turns on the i, m, or x option within the regular expression. If inside parentheses, it only affects the part within the parentheses. |
| (?-imx) | Temporarily turns off the i, m, or x option within the regular expression. If inside parentheses, it only affects the part within the parentheses. |
| (?: re) | Groups a regular expression but does not remember the matched text. |
| (?imx: re) | Temporarily turns on i, m, or x options within parentheses. |
| (?-imx: re) | Temporarily turns off i, m, or x options within parentheses. |
| (?#...) | Comment. |
| (?= re) | Specifies a position using a pattern. Has no range. |
| (?! re) | Specifies a position using a negated pattern. Has no range. |
| (?> re) | Matches an independent pattern without backtracking. |
| \w | Matches a word character. |
| \W | Matches a non-word character. |
| \s | Matches a whitespace character. Equivalent to [\t\n\r\f]. |
| \S | Matches a non-whitespace character. |
| \d | Matches a digit. Equivalent to [0-9]. |
| \D | Matches a non-digit. |
| \A | Matches the beginning of a string. |
| \Z | Matches the end of the string. If a newline exists, it matches only before the newline. |
| \z | Matches the end of the string. |
| \G | Matches the point where the last match completed. |
| \b | Matches a word boundary when outside brackets; matches a backspace (0x08) when inside brackets. |
| \B | Matches a non-word boundary. |
| \n, \t, etc. | Matches newline, carriage return, tab, etc. |
| \1...\9 | Matches the nth grouped subexpression. |
| \10 | If it has been matched, matches the nth grouped subexpression. Otherwise, it refers to the octal representation of a character encoding. |
Regular Expression Examples
Characters
| Examples | Description |
|---|---|
| /ruby/ | Matches "ruby" |
| ¥ | Matches the Yen sign. Ruby 1.9 and Ruby 1.8 support multiple characters. |
Character Classes
| Examples | Description |
|---|---|
| /[Rr]uby/ | Matches "Ruby" or "ruby" |
| /rub[ye]/ | Matches "ruby" or "rube" |
| /[aeiou]/ | Matches any lowercase vowel |
| /[0-9]/ | Matches any digit, same as / |
| /[a-z]/ | Matches any lowercase ASCII letter |
| /[A-Z]/ | Matches any uppercase ASCII letter |
| /[a-zA-Z0-9]/ | Matches any character in the brackets |
| /[^aeiou]/ | Matches any character that is not a lowercase vowel |
| /[^0-9]/ | Matches any non-digit character |
Special Character Classes
| Examples | Description |
|---|---|
| /./ | Matches any character except newline |
| /./m | Matches newline as well in multiline mode |
| /\d/ | Matches a digit, equivalent to /[0-9]/ |
| /\D/ | Matches a non-digit, equivalent to /[^0-9]/ |
| /\s/ | Matches a whitespace character, equivalent to /[ \t\r\n\f]/ |
| /\S/ | Matches a non-whitespace character, equivalent to /[^ \t\r\n\f]/ |
| /\w/ | Matches a word character, equivalent to /[A-Za-z0-9_]/ |
| /\W/ | Matches a non-word character, equivalent to /[^A-Za-z0-9_]/ |
Repetition
| Examples | Description |
|---|---|
| /ruby?/ | Matches "rub" or "ruby". The y is optional. |
| /ruby*/ | Matches "rub" plus zero or more y's. |
| /ruby+/ | Matches "rub" plus one or more y's. |
| /\d{3}/ | Matches exactly 3 digits. |
| /\d{3,}/ | Matches 3 or more digits. |
| /\d{3,5}/ | Matches 3, 4, or 5 digits. |
Non-greedy Repetition
This matches the smallest number of repetitions.
| Examples | Description |
|---|---|
| /<.*>/ | Greedy repetition: matches "<ruby>perl>" |
| /<.*?>/ | Non-greedy repetition: matches "<ruby>" in "<ruby>perl>" |
Grouping with Parentheses
| Examples | Description |
|---|---|
| /\D\d+/ | Without grouping: + repeats \d |
| /(\D\d)+/ | With grouping: + repeats the \D\d pair |
| /([Rr]uby(, )?)+/ | Matches "Ruby", "Ruby, ruby, ruby", etc. |
Backreferences
This matches a previously matched group again.
| Examples | Description |
|---|---|
| /([Rr])uby&\1ails/ | Matches ruby&rails or Ruby&Rails |
| /(['"])(?:(?!\1).)*\1/ | Single-quoted or double-quoted strings. \1 matches the characters matched by the first group, \2 matches those matched by the second group, and so on. |
Alternation
| Examples | Description |
|---|---|
| /ruby|rube/ | Matches "ruby" or "rube" |
| /rub(y|le)/ | Matches "ruby" or "ruble" |
| /ruby(!+|\?)/ | "ruby" followed by one or more ! or a ? |
Anchor
This requires specifying a match position.
| Examples | Description |
|---|---|
| /^Ruby/ | Matches a string or line beginning with "Ruby" |
| /Ruby$/ | Match a string or line ending with "Ruby" |
| /\ARuby/ | Match a string beginning with "Ruby" |
| /Ruby\Z/ | Match a string ending with "Ruby" |
| /\bRuby\b/ | Match "Ruby" at a word boundary |
| /\brub\B/ | \B is a non-word boundary: matches "rub" in "rube" and "ruby", but not a standalone "rub" |
| /Ruby(?=!)/ | Match "Ruby" if followed by an exclamation mark |
| /Ruby(?!!)/ | Match "Ruby" if not followed by an exclamation mark |
Special Syntax for Parentheses
| Examples | Description |
|---|---|
| /R(?#comment)/ | Match "R". All remaining characters are comments. |
| /R(?i)uby/ | Case-insensitive when matching "uby". |
| /R(?i:uby)/ | Same as above. |
| /rub(?:y|le))/ | Only group, no \1 backreference |
Search and Replace
subandgsuband their substitution variablessub!andgsub!are important string methods when using regular expressions.
All these methods perform search and replace operations using regular expression patterns.subandsub!Replaces the first occurrence of the pattern,gsubandgsub!Replaces all occurrences of the pattern.
subandgsubReturns a new string, leaving the original string unmodified, whilesub!andgsub!modifies the string on which they are called.
Examples
Try it »
The output of the above example is:
电话号码 : 138-3453-1111 电话号码 : 13834531111
Examples
Try it »
The output of the above example is:
Rails 是 Rails, Ruby on Rails 非常好的 Ruby 框架Other extensions