C# Regular expression
Regular expressionIt is a search formula used to describe and match text patterns. Simply put, you can understand it as a special wildcard language used to precisely find, validate, or replace content in strings.
For example, if you want to determine whether user input is a valid email address or phone number, or want to extract all dates from a piece of text—these are typical use cases for regular expressions.
The .NET framework has a built-in fully-featured regular expression engine, throughSystem.Text.RegularExpressionsThe namespace provides support.
A regular expression pattern consists of one or more characters, operators, and structures that together describe the text rules to match.
If you don't yet understand regular expressions, you can first read ourRegular Expressions - Tutorial。
Define a regular expression
Regular expressions are composed of the following types of building blocks, each responsible for different matching functions:
- Character escape: treat special characters as ordinary characters (e.g.,
\.match a literal period) - character class: matches a "certain type" of character (e.g.
\dMatch any digit) - Anchor point: specifies the "position" where the match occurs (e.g.
^represents the start of a line) - Grouping construct: group subpatterns together for easy capture or reference
- Qualifiers: controls the number of times an element appears (e.g.
+indicates one or more times) - Backreference construct: reference previously captured content for re-matching
- Alternation construct: implements "or" logic (e.g.
cat|dog) - replace: references captured content in replacement operations
- Miscellaneous construct: auxiliary functions such as inline options, comments, etc.
Character escape
In regular expressions, the backslash character (\) has two functions: first, it escapes the followingOrdinary characters take on special meaning(such as\nrepresents a newline), and second, it turnsSpecial characters are escaped to literal characters(such as\.match a real dot instead of "any character").
Common beginner misconception: in C# strings,\is itself an escape character, so writing in regex\d, in C# strings you need to write"\\d", or use a verbatim string@"\d"(Recommended).
The following table lists common escape characters:
| Escape characters | Description | Mode | match |
|---|---|---|---|
| \a | Matches the alarm (bell) character \u0007. | \a | The "\u0007" in "Warning!" + '\u0007' |
| \b | In character classes matches the backspace key \u0008. (Note: outside character classes, \b represents a word boundary, see "Anchors" section) | [\b]{3,} | "\b\b\b\b" in "\b\b\b\b" |
| \t | Matches tab character \u0009. Commonly used to match Tab-separated text. | (\w+)\t | The "Name " and "Addr " in "Name Addr " |
| \r | Matches carriage return \u000D. (\r is not equivalent to linefeed \n. Windows line breaks are usually \r\n) | \r\n(\w+) | "\r\nHello" in "\r\nHello\nWorld." |
| \v | Matches vertical tab \u000B. | [\v]{2,} | "\v\v\v" in "\v\v\v" |
| \f | Matches form feed \u000C. | [\f]{2,} | "\f\f\f" in "\f\f\f" |
| \n | Matches linefeed \u000A. Unix/Linux line endings usually have only \n. | \r\n(\w+) | "\r\nHello" in "\r\nHello\nWorld." |
| \e | Matches escape character \u001B. | \e | "\x001B" in "\x001B" |
| \ nnn | Specifies a character using octal representation (nnn consists of two or three digits). | \w\040\w | "a b" and "c d" in "a bc d" |
| \x nn | Specifies a character using hexadecimal representation (nn consists of exactly two digits). | \w\x20\w | "a b" and "c d" in "a bc d" |
| \c X \c x | Matches the ASCII control character specified by X or x, where X or x is the letter of the control character. | \cC | The "\x0003" inside "\x0003" (Ctrl-C) |
| \u nnnn | Matches a Unicode character using hexadecimal representation (four digits represented by nnnn). | \w\u0020\w | "a b" and "c d" in "a bc d" |
| \ | When followed by an unrecognized escape character, matches that character. | \d+[\+-x\*]\d+\d+[\+-x\*\d+ | The "2+2" and "3*9" in "(2+2) * 3*9" |
character class
Character classes are used to match a character among "a certain type" of charactersAny one. For example[aeiou]matches any one vowel letter,\dMatches any single digit. This is one of the most commonly used basic functions in regular expressions.
The following table lists character classes:
| character class | Description | Mode | match |
|---|---|---|---|
| [character_group] | Matches any single character in character_group. By default, matching is case-sensitive. | [mn] | the 'm' in 'mat', the 'm' and 'n' in 'moon' |
| [^character_group] | Negation: Matches any single character not in character_group. By default, characters in character_group are case-sensitive. | [^aei] | "v" and "l" in "avail" |
| [ first - last ] | Character range: Matches any single character in the range from first to last. | [b-d] | [b-d]irds can match Birds, Cirds, Dirds |
| . | Wildcard: Matches any single character except \n. To match a literal period character (. or \u002E), you must precede it with an escape character (\.). | a.e | "ave" in "have", "ate" in "mate" |
| \p{ name } | andnameMatches any single character in the specified Unicode general category or named block. | \p{Lu} | "C" and "L" in "City Lights" |
| \P{ name } | and not innameMatches any single character in the specified Unicode general category or named block. | \P{Lu} | "i", "t" and "y" in "City" |
| \w | Matches any word character (letters, digits, underscore). Equivalent to [a-zA-Z0-9_] (in ASCII range). | \w | The "R", "o", "m", and "1" in "Room#1" |
| \W | Matches any non-word character. This is the negation of \w. | \W | "#" in "Room#1" |
| \s | Matches any whitespace character (space, tab, newline, etc.). | \w\s | The "D " in "ID A1.3" |
| \S | Matches any non-whitespace character. This is the negation of \s. | \s\S | " _" in "int __ctr" |
| \d | Matches any decimal digit. Equivalent to [0-9]. | \d | "4" in "4 = IV" |
| \D | Matches any character that is not a decimal digit. This is the negation of \d. | \D | The " ", "=", " ", "I", and "V" in "4 = IV" |
Anchor point
Anchors (also called "anchors") do not match any specific character, but match a certain position in the string.Position. They are "zero-width", do not consume any characters, just assert that the current position satisfies a certain condition.
For example^\d{3}means "three digits at the start of the string",\bmeans "boundary between word and non-word characters".
The following table lists anchors:
| Assertion | Description | Mode | match |
|---|---|---|---|
| ^ | The match must start at the beginning of the string or line. | ^\d{3} | "567" in "567-777-" |
| $ | The match must appear at the end of the string, or before the\nbefore. | -\d{4}$ | "-2012" in "8-12-2012" |
| \A | The match must appear at the beginning of the string (not affected by multiline mode; always the very beginning of the entire string). | \A\w{4} | "Code" in "Code-007-" |
| \Z | The match must appear at the end of the string, or before the\nbefore. | -\d{3}\Z | "-007" in "Bond-901-007" |
| \z | The match must appear at the end of the string (strict end, no \n allowed at the end). | -\d{3}\z | "-333" in "-901-333" |
| \G | The match must appear where the previous match ended. Commonly used in consecutive matching scenarios. | \G\(\d\) | "(1)", "(3)", and "(5)" in "(1)(3)(5)[7](9)" |
| \b | Matches a word boundary, that is, the position between a word and a space. | er\b | Matches "er" in "never", but cannot match "er" in "verb". |
| \B | Matches a non-word boundary. | er\B | Matches "er" in "verb", but cannot match "er" in "never". |
Grouping construct
Grouping constructs use parentheses( )Enclose a part of the regular expression to form a subexpression. Grouping has two main purposes:
- catch: "save" the matched substring for later extraction or reference in replacement.
- Scope restriction: allow quantifiers or alternation constructs to only apply to the subexpression within the group.
This part is relatively difficult to understand; you can readRegular Expression - Selection 、Lookahead and lookbehind assertions in regular expressionsHelps understanding.
The following table lists grouping constructs:
| Grouping construct | Description | Mode | match |
|---|---|---|---|
| ( subexpression ) | Captures the matched subexpression and assigns it to a zero-based ordinal number. | (\w)\1 | "ee" in "deep" |
| (?< name >subexpression) | Captures the matched subexpression into a named group. Named groups are more readable than numeric numbering and are recommended for use in complex regular expressions. | (?< double>\w)\k< double> | "ee" in "deep" |
| (?< name1 -name2 >subexpression) | Defines a balancing group definition. Used to match nested structures (such as paired parentheses), an advanced feature. | (((?'Open'\()[^\(\)]*)+((?'Close-Open'\))[^\(\)]*)+)*(?(Open)(?!))$ | "((1-3)*(3-1))" in "3+2^((1-3)*(3-1))" |
| (?: subexpression) | Defines a non-capturing group. Used when only grouping (limiting scope) is needed without saving matched content; performance is slightly better than capturing groups. | Write(?:Line)? | The "WriteLine" in "Console.WriteLine()" |
| (?imnsx-imnsx:subexpression) | Apply or disablesubexpressionOptions specified in. | A\d{2}(?i:\w+)\b | "A12xl" and "A12XL" in "A12xl A12XL a12xl" |
| (?= subexpression) | Zero-width positive lookahead assertion. Matches a position that is followed by subexpression, but does not consume characters. | \w+(?=\.) | "He is. The dog ran. The sun is out." in "is"、 "ran" and "out" |
| (?! subexpression) | Zero-width negative lookahead assertion. Matches a position not followed by subexpression. | \b(?!un)\w+\b | "sure" and "used" in "unsure sure unity used" |
| (?<=subexpression) | Zero-width positive lookbehind assertion. Matches a position that is immediately preceded by subexpression. | (?<=19)\d{2}\b | "99", "50", and "05" in "1851 1999 1950 1905 2003" |
| (?<! subexpression) | Zero-width negative lookbehind assertion. Matches a position not preceded by subexpression. | (?<!wo)man\b | "man" in "Hi woman Hi man" |
| (?> subexpression) | Non-backtracking (atomic group) subexpression. Once matched, no backtracking is allowed, which can improve performance in some scenarios. | [13579](?>A+B+) | "1ABB", "3ABB", and "5AB" in "1ABB 3ABBC 5AB 5AC" |
Example
using System.Text.RegularExpressions;
public class Example
{
public static void Main()
{
string input = "1851 1999 1950 1905 2003";
string pattern = @"(?<=19)\d{2}\b";
foreach (Match match in Regex.Matches(input, pattern))
Console.WriteLine(match.Value);
}
}
Run Example »
Qualifiers
Quantifiers specify that the preceding element (character, character class, or group) must appearhow many timesOnly then is the match considered successful.
Quantifiers are by defaultgreedy, that is, match as many characters as possible. Adding after a quantifier?can becomeLazy (non-greedy)pattern, that is, match as few characters as possible. Beginners can first master*、+、?、{n}These four basic quantifiers.
The following table lists quantifiers:
| Qualifiers | Description | Mode | match |
|---|---|---|---|
| * | Matches the previous element zero or more times. (Zero occurrences also count as a match) | \d*\.\d | ".0"、 "19.9"、 "219.9" |
| + | Matches the previous element one or more times. (At least once) | "be+" | The "bee" in "been", and the "be" in "bent" |
| ? | Matches the previous element zero or one time. (i.e., the element is optional) | "rai?n" | "ran"、 "rain" |
| { n } | Matches the previous element exactly n times. | ",\d{3}" | "1,043.6" in ",043", "9,876,543,210" in ",876"、 ",543" and ",210" |
| { n ,} | Matches the previous element at least n times. | "\d{2,}" | "166"、 "29"、 "1930" |
| { n , m } | Matches the previous element at least n times, but no more than m times. | "\d{3,5}" | The "19302" in "166", "17668", "193024" |
| *? | Matches the previous element zero or more times, but as few times as possible (lazy mode). | \d*?\.\d | ".0"、 "19.9"、 "219.9" |
| +? | Matches the previous element one or more times, but as few times as possible (lazy mode). | "be+?" | "be" in "been", "be" in "bent" |
| ?? | Matches the previous element zero or one time, but as few times as possible (lazy mode). | "rai??n" | "ran"、 "rain" |
| { n }? | Matches the preceding element exactly n times. | ",\d{3}?" | "1,043.6" in ",043", "9,876,543,210" in ",876"、 ",543" and ",210" |
| { n ,}? | Matches the previous element at least n times, but as few times as possible. | "\d{2,}?" | "166", "29" and "1930" |
| { n , m }? | Matches the previous element between n and m times, but as few times as possible. | "\d{3,5}?" | "166", "17668", and the "193" and "024" in "193024" |
Backreference constructs
Backreferences allow, within the same regular expression, referencingContent of previously captured groupsperforms a match again. For example, using(\w)\1can find consecutive repeated characters like "ee", "ll".
The following table lists backreference constructs:
| Backreference constructs | Description | Mode | match |
|---|---|---|---|
| \ number | Backreference. Matches the value of a numbered subexpression. | (\w)\1 | "ee" in "seek" |
| \k< name > | Named backreference. Matches the value of a named expression. More readable than numeric references, recommended. | (?< char>\w)\k< char> | "ee" in "seek" |
Alternation construct
Alternation constructs use the vertical bar|Implements "or" logic, allowing the regular expression to match any one of multiple candidate patterns. Similar to the||Operators.
The following table lists alternation constructs:
| Alternation construct | Description | Mode | match |
|---|---|---|---|
| | | Matches any one of the elements separated by the vertical bar (|) character. | th(e|is|at) | The "the" and "this" in "this is the day. " |
| (?( expression )yes | no ) | If the regular expression pattern is matched by the specified expression, then matches theyes; otherwise match the optionalnopart. expression is interpreted as a zero-width assertion. | (?(A)A\d{2}\b|\b\d{3}\b) | The "A10" and "910" in "A10 C103 910" |
| (?( name )yes | no ) | If name or a named or numbered capturing group has a match, then matches theyes; otherwise match the optionalno。 | (?< quoted>")?(?(quoted).+?"|\S+\s) | "Dogs.jpg "Yiska playing.jpg"" in Dogs.jpg and "Yiska playing.jpg" |
replace
Replacement syntax is used forRegex.Replace()method'sReplacement Pattern StringIn, you can use$Use numbers or names to reference previously captured group content, enabling flexible text reorganization.
The following table lists characters used for substitution:
| Character | Description | Mode | Replacement pattern | Input string | result string |
|---|---|---|---|---|---|
| $number | Replacement by groupnumberThe matched substring. | \b(\w+)(\s)(\w+)\b | $3$2$1 | "one two" | "two one" |
| ${name} | Replace by named groupnameThe matched substring. | \b(?< word1>\w+)(\s)(?< word2>\w+)\b | ${word2} ${word1} | "one two" | "two one" |
| $$ | The replacement character "$". | \b(\d+)\s?USD | $$$1 | "103 USD" | "$103" |
| $& | Replaces a copy of the entire match. | (\$*(\d*(\.+\d+)?){1}) | **$& | "$1.30" | "**$1.30" |
| $` | Replaces all text of the input string before the match. | B+ | $` | "AABBCC" | "AAAACC" |
| $' | Replaces all text of the input string after the match. | B+ | $' | "AABBCC" | "AACCCC" |
| $+ | Replaces the last captured group. | B+(C+) | $+ | "AABBCCDD" | AACCDD |
| $_ | Replaces the entire input string. | B+ | $_ | "AABBCC" | "AAAABBCCCC" |
Miscellaneous construct
The following table lists miscellaneous constructs:
| Constructor | Description | Example |
|---|---|---|
| (?imnsx-imnsx) | Sets or disables options such as case-insensitivity in the middle of a pattern. | \bA(?i)b\w+\b matches "ABA" and "Able" in "ABA Able Act" |
| (?#comment) | Inline comment. This comment terminates at the first closing parenthesis. | \bA(?#matches a word starting with A)\w+\b |
| # [End of line] | This comment starts with an unescaped # and continues to the end of the line. | (?x)\bA\w+\b#Matches words starting with A |
Regex Class
The Regex class is the core class for using regular expressions in .NET, located inSystem.Text.RegularExpressionsunder the namespace. Before using it, you need to add at the top of the file:
using System.Text.RegularExpressions;
The following table lists some commonly used methods in the Regex class:
| Serial Number | Method & Description |
|---|---|
| 1 | public bool IsMatch(
string input
)
Determines whether the input string contains content that matches the regular expression pattern. Commonly used in form validation, such as validating phone number or email formats. |
| 2 | public bool IsMatch(
string input,
int startat
)
Performs match determination starting from the specified position in the string. |
| 3 | public static bool IsMatch(
string input,
string pattern
)
Static method; no need to create a Regex object first. Pass the pattern string directly for match determination. Suitable for one-off use cases. |
| 4 | public MatchCollection Matches(
string input
)
Searches the input string for all matches, returns a MatchCollection, which can be iterated with foreach over each match result. |
| 5 | public string Replace(
string input,
string replacement
)
Replaces all content in the input string that matches the regular expression pattern with the specified string. |
| 6 | public string[] Split(
string input
)
Splits the input string into an array of substrings according to the delimiters defined by the regular expression pattern. More flexible than string.Split(), supporting complex delimiter rules. |
For the complete list of properties of the Regex class, please refer to the Microsoft C# documentation.
Example 1
The following example matches words starting with 'S':
Analysis:\bis a word boundary,Smatches the uppercase letter S,\S*Matches zero or more non-whitespace characters; combined, this means "complete words starting with S".
Example
using System.Text.RegularExpressions;
namespace RegExApplication
{
class Program
{
private static void showMatch(string text, string expr)
{
Console.WriteLine("The Expression: " + expr);
MatchCollection mc = Regex.Matches(text, expr);
foreach (Match m in mc)
{
Console.WriteLine(m);
}
}
static void Main(string[] args)
{
string str = "A Thousand Splendid Suns";
Console.WriteLine("Matching words that start with 'S': ");
showMatch(str, @"\bS\S*");
Console.ReadKey();
}
}
}
When the above code is compiled and executed, it produces the following result:
Matching words that start with 'S': The Expression: \bS\S* Splendid Suns
Example 2
The following example matches words starting with 'm' and ending with 'e':
Analysis:\bmmatches words starting with m,\S*matches any non-whitespace character in the middle,e\bRequires ending with e followed by a word boundary; together, it finds all words starting with m and ending with e.
Example
using System.Text.RegularExpressions;
namespace RegExApplication
{
class Program
{
private static void showMatch(string text, string expr)
{
Console.WriteLine("The Expression: " + expr);
MatchCollection mc = Regex.Matches(text, expr);
foreach (Match m in mc)
{
Console.WriteLine(m);
}
}
static void Main(string[] args)
{
string str = "make maze and manage to measure it";
Console.WriteLine("Matching words start with 'm' and ends with 'e':");
showMatch(str, @"\bm\S*e\b");
Console.ReadKey();
}
}
}
When the above code is compiled and executed, it produces the following result:
Matching words start with 'm' and ends with 'e': The Expression: \bm\S*e\b make maze manage measure
Example 3
The following example replaces extra spaces:
Analysis:\\s+Matches one or more consecutive whitespace characters (spaces, tabs, etc.) and replaces them all with a single space, thereby merging extra whitespace.
Example
using System.Text.RegularExpressions;
namespace RegExApplication
{
class Program
{
static void Main(string[] args)
{
string input = "Hello World ";
string pattern = "\\s+";
string replacement = " ";
Regex rgx = new Regex(pattern);
string result = rgx.Replace(input, replacement);
Console.WriteLine("Original String: {0}", input);
Console.WriteLine("Replacement String: {0}", result);
Console.ReadKey();
}
}
}
When the above code is compiled and executed, it produces the following result:
Original String: Hello World Replacement String: Hello World
GeneratedRegex: source generator optimization (.NET 7+)
Beginning with .NET 7, C# introduced[GeneratedRegex]attribute (source generator), which is a major upgrade to the traditionalRegexan important upgrade to the class, specifically forPerformance-sensitiveorRepeated useregex scenarios.
Problems with traditional approaches
Use traditionalnew Regex(pattern)When ..., the parsing and compilation of the regular expression occurs atruntime, each creation has overhead. Although you can useRegexOptions.CompiledImproves runtime speed, but compilation still occurs at runtime and increases startup time and memory usage.
Traditional way
// Method 1: Re-parse every time (slowest)
bool isMatch = Regex.IsMatch(input, @"\d{4}-\d{2}-\d{2}");
// Method 2: static field + Compiled (common optimization practice)
private static readonly Regex DateRegex = new Regex(@"\d{4}-\d{2}-\d{2}", RegexOptions.Compiled);
Usage of GeneratedRegex
Usage[GeneratedRegex]With the attribute, the compiler willCompilation phasegenerate the regular expression directly as efficient C# code, without runtime parsing or compilation.
Usage requirements: .NET 7 or higher is required, and the method must bepartialmethod, the containing class must also bepartialclass.
GeneratedRegex basic usage
using System.Text.RegularExpressions;
namespace RegExApplication
{
// The class must be declared as partial
partial class Program
{
// Use the [GeneratedRegex] attribute, the compiler automatically generates the implementation
[GeneratedRegex(@"\d{4}-\d{2}-\d{2}")]
private static partial Regex DateRegex();
// You can also add options, such as ignoring case
[GeneratedRegex(@"\bS\S*", RegexOptions.IgnoreCase)]
private static partial Regex StartsWithSRegex();
static void Main(string[] args)
{
string input = "Today is 2024-06-18, next event: 2025-01-01";
// Use it like calling a normal method, returns a Regex instance
foreach (Match m in DateRegex().Matches(input))
{
Console.WriteLine("Found date:" + m.Value);
}
// Verify if it matches
Console.WriteLine(StartsWithSRegex().IsMatch("Splendid")); // True
}
}
}
When the above code is compiled and executed, it produces the following result:
找到日期:2024-06-18 找到日期:2025-01-01 True
GeneratedRegex vs Traditional Regex
The following table compares the two approaches from multiple dimensions to help you choose the appropriate usage:
| Comparison Dimension | Traditional Regex | GeneratedRegex (source generator) | Recommended |
|---|---|---|---|
| Minimum version requirement | .NET Framework / .NET Core All Versions Supported | .NET 7 and above | — |
| Compile timing | Runtime parsing and compilation | At compile time (build time), the Roslyn source generator generates code | GeneratedRegex |
| Execution performance | Normal: slower; with Compiled option: faster, but high startup overhead | Fastest, close to the performance of hand-written code, and no startup overhead | GeneratedRegex |
| Memory usage | Compiled mode generates IL code and has higher memory usage | It generates ordinary C# code, which is more memory-friendly | GeneratedRegex |
| AOT Compatibility | RegexOptions.Compiled is incompatible with AOT (ahead-of-time compilation) | Fully compatible with Native AOT, suitable for publishing standalone applications | GeneratedRegex |
| Code readability | The pattern string is written directly in the code, which is more intuitive | Requires defining a partial method, and the code structure is slightly more complex | Traditional Regex (simple scenarios) |
| Dynamic mode | Supported, the pattern can be a runtime variable | Not supported, the pattern must be a compile-time constant | Traditional Regex (dynamic scenarios) |
| Debugging support | Ordinary | The generated code can be viewed and debugged directly, making it more transparent | GeneratedRegex |
| Applicable scenarios | One-off use, dynamic patterns, legacy .NET projects | High-frequency calls, performance-sensitive, AOT publishing, new .NET projects | Depends on the context |
Usage suggestions
- If you use.NET 7+and the regex pattern is fixed,Prefer to use
[GeneratedRegex]。 - If the regex pattern needs to be dynamically built at runtime (e.g., generated from user input), only the traditional
new Regex(pattern)。 - If the project needs to supportNative AOTWhen publishing (e.g., .NET 8's AOT mode), be sure to use
[GeneratedRegex]rather thanRegexOptions.Compiled。 - For beginners, you can first use the traditional method to learn and practice. After understanding the principles of regex, gradually migrate to
[GeneratedRegex]。
Example: Using GeneratedRegex to validate email format
using System.Text.RegularExpressions;
partial class EmailValidator
{
// Define the email validation regex, ignoring case
[GeneratedRegex(@"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$", RegexOptions.IgnoreCase)]
private static partial Regex EmailRegex();
public static bool IsValidEmail(string email)
{
return EmailRegex().IsMatch(email);
}
static void Main(string[] args)
{
string[] emails = { "[email protected]", "bad-email", "[email protected]", "missing@dot" };
foreach (var email in emails)
{
Console.WriteLine($"{email}: {(IsValidEmail(email) ? "valid" : "Invalid")}");
}
}
}
When the above code is compiled and executed, it produces the following result:
[email protected]: 有效 bad-email: 无效 [email protected]: 有效 missing@dot: 无效other extensions