Scala Regular Expressions
Scala supports regular expressions through the scala.util.matching package'sRegexRegex class. The following example demonstrates using regular expressions to find words.Scala :
Example
object Test {
def main(args: Array[String]) {
val pattern = "Scala".r
val str = "Scala is Scalable and cool"
println(pattern findFirstIn str)
}
}
Execute the above code and the output result is:
$ scalac Test.scala $ scala Test Some(Scala)
In the example, the r() method of the String class is used to construct a Regex object.
Then use the findFirstIn method to find the first match.
If you need to view all matches, you can use the findAllIn method.
You can use the mkString() method to concatenate the strings of regex match results, and can use the pipe (|) to set different patterns:
Example
object Test {
def main(args: Array[String]) {
val pattern = new Regex("(S|s)cala") // The first letter can be uppercase S or lowercase s
val str = "Scala is scalable and cool"
println((pattern findAllIn str).mkString(",")) // Use comma , to concatenate returned results
}
}
Execute the above code and the output result is:
$ scalac Test.scala $ scala Test Scala,scala
If you need to replace the matched text with a specified keyword, you can usereplaceFirstIn( )method to replace the first match, usereplaceAllIn( )method to replace all matches, the example is as follows:
Example
def main(args: Array[String]) {
val pattern = "(S|s)cala".r
val str = "Scala is scalable and cool"
println(pattern replaceFirstIn(str, "Java"))
}
}
Execute the above code and the output result is:
$ scalac Test.scala $ scala Test Java is scalable and cool
Regular Expressions
Scala's regular expressions inherit Java's syntax rules, while Java largely uses Perl language rules.
The following table gives some commonly used regular expression rules:
| Expression | Matching rule |
|---|---|
| ^ | Matches the position at the start of the input string. |
| $ | Matches the position at the end of the input string. |
| . | Matches any single character except "\r\n". |
| [...] | Character set. Matches any one character included. For example, "[abc]" matches "a" in "plain". |
| [^...] | Negated character set. Matches any character not included. For example, "[^abc]" matches "p", "l", "i", "n" in "plain". |
| \\A | Matches the start position of the input string (without multiline support) |
| \\z | End of string (like $, but not affected by multiline processing options) |
| \\Z | End of string or line (not affected by multiline processing options) |
| re* | Repeat zero or more times |
| re+ | Repeat one or more times |
| re? | Repeat zero or one time |
| re{ n} | Repeat exactly n times |
| re{ n,} | |
| re{ n, m} | Repeat between n and m times |
| a|b | Match a or b |
| (re) | Match re and capture the text into an automatically named group |
| (?: re) | Match re, do not capture the matched text, and do not assign a group number to this group |
| (?> re) | Greedy subexpression |
| \\w | Matches letters, digits, or underscores |
| \\W | Matches any character that is not a letter, digit, underscore, or Chinese character |
| \\s | Matches any whitespace character, equivalent to [\t\n\r\f] |
| \\S | Matches any character that is not a whitespace character |
| \\d | Matches a digit, like [0-9] |
| \\D | Matches any non-digit character |
| \\G | Start of current search |
| \\n | Newline character |
| \\b | Usually a word boundary position, but if used in a character class, represents backspace |
| \\B | Matches a position that is not the start or end of a word |
| \\t | Tab character |
| \\Q | Beginning quote:\Q(a+b)*3\ECan match the text "(a+b)*3". |
| \\E | Closing quote:\Q(a+b)*3\ECan match the text "(a+b)*3". |
Regular Expression Examples
| Example | Description |
|---|---|
| . | Matches any single character except "\r\n". |
| [Rr]uby | Matches "Ruby" or "ruby" |
| rub[ye] | Matches "ruby" or "rube" |
| [aeiou] | Matches lowercase letters: aeiou |
| [0-9] | Matches any digit, like |
| [a-z] | Matches any ASCII lowercase letter |
| [A-Z] | Matches any ASCII uppercase letter |
| [a-zA-Z0-9] | Matches digits, uppercase and lowercase letters |
| [^aeiou] | Matches characters other than aeiou |
| [^0-9] | Matches characters other than digits |
| \\d | Matches a digit, like: [0-9] |
| \\D | Matches a non-digit, like: [^0-9] |
| \\s | Matches whitespace, like: [ \t\r\n\f] |
| \\S | Matches non-whitespace, like: [^ \t\r\n\f] |
| \\w | Matches letters, digits, underscores, like: [A-Za-z0-9_] |
| \\W | Matches non-letters, digits, underscores, like: [^A-Za-z0-9_] |
| ruby? | Matches "rub" or "ruby": y is optional |
| ruby* | Matches "rub" followed by zero or more y's. |
| ruby+ | Matches "rub" followed by one or more y's. |
| \\d{3} | Matches exactly 3 digits. |
| \\d{3,} | Matches 3 or more digits. |
| \\d{3,5} | Matches 3, 4, or 5 digits. |
| \\D\\d+ | No group: + repeats \d |
| (\\D\\d)+/ | Group: + repeats the \D\d pair |
| ([Rr]uby(, )?)+ | Matches "Ruby", "Ruby, ruby, ruby", etc. |
Note that each character in the above table uses two backslashes. This is because in Java and Scala, the backslash in strings is an escape character. So if you want to output\you need to write it in the string as\\to get a backslash. See the following example:
Example
object Test {
def main(args: Array[String]) {
val pattern = new Regex("abl[ae]\\d+")
val str = "ablaw is able1 and cool"
println((pattern findAllIn str).mkString(","))
}
}
Execute the above code and the output result is:
$ scalac Test.scala $ scala Test able1Other extensions