Julia Regular Expressions

A regular expression describes a pattern for string matching. It can be used to check whether a string contains a certain substring, replace matched substrings, or extract substrings that meet a certain condition from a string.

Julia has regular expressions (regexes) that are compatible with Perl.

The three forms of regular expressions in Julia are matching, substitution, and transformation:

  • Matching: m// (can also be abbreviated as //, omitting m)

  • Substitution: s///

  • Transformation: tr///

These three forms are generally used with=~or!~together; =~ indicates a match, and !~ indicates no match.

In Julia, regular expression input uses prefixes that allrstart with:

Example

julia> re = r"^\s*(?:#|$)"
r"^\s*(?:#|$)"


julia> typeof(re)
Regex

To check whether a regular expression matches a string, useoccursin:

Example

julia> occursin(r"^\s*(?:#|$)", "not a comment")
false

julia> occursin(r"^\s*(?:#|$)", "# a comment")
true

As you can see,occursinit returns only true or false, indicating whether the given regex appears in the string. However, usually we don't just want to know whether the string matches, but also how it matches. To capture the matching information, you can use thematchfunction:

Example

julia> match(r"^\s*(?:#|$)", "not a comment")

julia> match(r"^\s*(?:#|$)", "# a comment")
RegexMatch("#")

If the regex does not match the given string, match returns nothing — a special value that prints nothing in the interactive prompt. Apart from not printing, it is a perfectly normal value, which can be tested programmatically:

Example

m = match(r"^\s*(?:#|$)", line)
if m === nothing
    println("not a comment")
else
    println("blank or comment")
end

If the regex matches, match returns a RegexMatch object. These objects record how the expression matched, including the substring matched by the pattern and any substrings that may have been captured. The example above only captured part of the matching substring, but perhaps we want to capture any non-empty text after the comment character. We can do this:

Example

julia> m = match(r"^\s*(?:#\s*(.*?)\s*$|$)", "# a comment ")
RegexMatch("# a comment ", 1="a comment")

When calling match, you can optionally specify the index at which to start searching. For example:

Example

julia> m = match(r"[0-9]","aaaa1aaaa2aaaa3",1)
RegexMatch("1")

julia> m = match(r"[0-9]","aaaa1aaaa2aaaa3",6)
RegexMatch("2")

julia> m = match(r"[0-9]","aaaa1aaaa2aaaa3",11)
RegexMatch("3")

You can extract the following information from a RegexMatch object:

  • The entire matched substring:m.match
  • The captured substrings as an array of strings:m.captures
  • The offset at which the entire match begins:m.offset
  • The offsets of the captured substrings as a vector:m.offsets

When a capture does not match,m.capturesit no longer contains a substring at that position, but contains nothing; in addition,m.offsetsthe offset is 0 (remember that Julia's indexing starts at 1, so a zero offset is invalid for a string). Here are two somewhat contrived examples:

Example

julia> m = match(r"(a|b)(c)?(d)", "acd")
RegexMatch("acd", 1="a", 2="c", 3="d")

julia> m.match
"acd"

julia> m.captures
3-element Vector{Union{Nothing, SubString{String}}}:
 "a"
 "c"
 "d"

julia> m.offset
1

julia> m.offsets
3-element Vector{Int64}:
 1
 2
 3

julia> m = match(r"(a|b)(c)?(d)", "ad")
RegexMatch("ad", 1="a", 2=nothing, 3="d")

julia> m.match
"ad"

julia> m.captures
3-element Vector{Union{Nothing, SubString{String}}}:
 "a"
 nothing
 "d"

julia> m.offset
1

julia> m.offsets
3-element Vector{Int64}:
 1
 0
 2

Having captures returned as an array is convenient, because you can bind them to local variables using destructuring syntax. For convenience, RegexMatch objects implement an iterator method that delegates to the captures field, so you can directly destructure the match object:

Example

julia> first, second, third = m; first
"a"

Captures can also be accessed by indexing the RegexMatch object with the number or name of the capture group:

Example

julia> m=match(r"(?<hour>\d+):(?<minute>\d+)","12:45")
RegexMatch("12:45", hour="12", minute="45")

julia> m[:minute]
"45"

julia> m[2]
"45"

When using replace, you can reference captures in the replacement string by using \n to refer to the nth capture group and prefixing the replacement string with s. Capture group 0 refers to the entire match. In replacements, \gcan be used to reference named capture groups. For example:

julia> replace("first second", r"(\w+) (?<agroup>\w+)" => s"\g<agroup> \1")
"second first"

For clarity, numbered capture groups can also be referenced with \gfor example:

julia> replace("a", r"." => s"\g<0>1")
"a1"

You can modify a regex by adding flags such as i, m, s, and x after the closing double quote.

For more on regular expressions, see:Regular Expressions - Tutorial

Other Extensions