C# Regular expression

Regular expressionIt is a search formula used to describe and match text patterns. Simply put, you can understand it as a special wildcard language used to precisely find, validate, or replace content in strings.

For example, if you want to determine whether user input is a valid email address or phone number, or want to extract all dates from a piece of text—these are typical use cases for regular expressions.

The .NET framework has a built-in fully-featured regular expression engine, throughSystem.Text.RegularExpressionsThe namespace provides support.

A regular expression pattern consists of one or more characters, operators, and structures that together describe the text rules to match.

If you don't yet understand regular expressions, you can first read ourRegular Expressions - Tutorial。

Define a regular expression

Regular expressions are composed of the following types of building blocks, each responsible for different matching functions:

  • Character escape: treat special characters as ordinary characters (e.g.,\.match a literal period)
  • character class: matches a "certain type" of character (e.g.\dMatch any digit)
  • Anchor point: specifies the "position" where the match occurs (e.g.^represents the start of a line)
  • Grouping construct: group subpatterns together for easy capture or reference
  • Qualifiers: controls the number of times an element appears (e.g.+indicates one or more times)
  • Backreference construct: reference previously captured content for re-matching
  • Alternation construct: implements "or" logic (e.g.cat|dog)
  • replace: references captured content in replacement operations
  • Miscellaneous construct: auxiliary functions such as inline options, comments, etc.

Character escape

In regular expressions, the backslash character (\) has two functions: first, it escapes the followingOrdinary characters take on special meaning(such as\nrepresents a newline), and second, it turnsSpecial characters are escaped to literal characters(such as\.match a real dot instead of "any character").

Common beginner misconception: in C# strings,\is itself an escape character, so writing in regex\d, in C# strings you need to write"\\d", or use a verbatim string@"\d"(Recommended).

The following table lists common escape characters:

Escape charactersDescriptionModematch
\aMatches the alarm (bell) character \u0007.\aThe "\u0007" in "Warning!" + '\u0007'
\bIn character classes matches the backspace key \u0008. (Note: outside character classes, \b represents a word boundary, see "Anchors" section)[\b]{3,}"\b\b\b\b" in "\b\b\b\b"
\tMatches tab character \u0009. Commonly used to match Tab-separated text.(\w+)\tThe "Name " and "Addr " in "Name Addr "
\rMatches carriage return \u000D. (\r is not equivalent to linefeed \n. Windows line breaks are usually \r\n)\r\n(\w+)"\r\nHello" in "\r\nHello\nWorld."
\vMatches vertical tab \u000B.[\v]{2,}"\v\v\v" in "\v\v\v"
\fMatches form feed \u000C.[\f]{2,}"\f\f\f" in "\f\f\f"
\nMatches linefeed \u000A. Unix/Linux line endings usually have only \n.\r\n(\w+)"\r\nHello" in "\r\nHello\nWorld."
\eMatches escape character \u001B.\e"\x001B" in "\x001B"
\ nnnSpecifies a character using octal representation (nnn consists of two or three digits).\w\040\w"a b" and "c d" in "a bc d"
\x nnSpecifies a character using hexadecimal representation (nn consists of exactly two digits).\w\x20\w"a b" and "c d" in "a bc d"
\c X \c x Matches the ASCII control character specified by X or x, where X or x is the letter of the control character.\cCThe "\x0003" inside "\x0003" (Ctrl-C)
\u nnnnMatches a Unicode character using hexadecimal representation (four digits represented by nnnn).\w\u0020\w"a b" and "c d" in "a bc d"
\When followed by an unrecognized escape character, matches that character.\d+[\+-x\*]\d+\d+[\+-x\*\d+The "2+2" and "3*9" in "(2+2) * 3*9"

character class

Character classes are used to match a character among "a certain type" of charactersAny one. For example[aeiou]matches any one vowel letter,\dMatches any single digit. This is one of the most commonly used basic functions in regular expressions.

The following table lists character classes:

character classDescriptionModematch
[character_group]Matches any single character in character_group. By default, matching is case-sensitive.[mn]the 'm' in 'mat', the 'm' and 'n' in 'moon'
[^character_group]Negation: Matches any single character not in character_group. By default, characters in character_group are case-sensitive.[^aei]"v" and "l" in "avail"
[ first - last ]Character range: Matches any single character in the range from first to last.[b-d][b-d]irds can match Birds, Cirds, Dirds
.Wildcard: Matches any single character except \n.
To match a literal period character (. or \u002E), you must precede it with an escape character (\.).
a.e"ave" in "have", "ate" in "mate"
\p{ name }andnameMatches any single character in the specified Unicode general category or named block.\p{Lu} "C" and "L" in "City Lights"
\P{ name }and not innameMatches any single character in the specified Unicode general category or named block.\P{Lu}"i", "t" and "y" in "City"
\wMatches any word character (letters, digits, underscore). Equivalent to [a-zA-Z0-9_] (in ASCII range).\wThe "R", "o", "m", and "1" in "Room#1"
\WMatches any non-word character. This is the negation of \w.\W"#" in "Room#1"
\sMatches any whitespace character (space, tab, newline, etc.).\w\sThe "D " in "ID A1.3"
\SMatches any non-whitespace character. This is the negation of \s.\s\S" _" in "int __ctr"
\dMatches any decimal digit. Equivalent to [0-9].\d"4" in "4 = IV"
\DMatches any character that is not a decimal digit. This is the negation of \d.\DThe " ", "=", " ", "I", and "V" in "4 = IV"

Anchor point

Anchors (also called "anchors") do not match any specific character, but match a certain position in the string.Position. They are "zero-width", do not consume any characters, just assert that the current position satisfies a certain condition.

For example^\d{3}means "three digits at the start of the string",\bmeans "boundary between word and non-word characters".

The following table lists anchors:

AssertionDescriptionModematch
^The match must start at the beginning of the string or line.^\d{3}"567" in "567-777-"
$The match must appear at the end of the string, or before the\nbefore.-\d{4}$"-2012" in "8-12-2012"
\AThe match must appear at the beginning of the string (not affected by multiline mode; always the very beginning of the entire string).\A\w{4}"Code" in "Code-007-"
\ZThe match must appear at the end of the string, or before the\nbefore.-\d{3}\Z"-007" in "Bond-901-007"
\zThe match must appear at the end of the string (strict end, no \n allowed at the end).-\d{3}\z"-333" in "-901-333"
\GThe match must appear where the previous match ended. Commonly used in consecutive matching scenarios.\G\(\d\)"(1)", "(3)", and "(5)" in "(1)(3)(5)[7](9)"
\bMatches a word boundary, that is, the position between a word and a space.er\bMatches "er" in "never", but cannot match "er" in "verb".
\BMatches a non-word boundary.er\BMatches "er" in "verb", but cannot match "er" in "never".

Grouping construct

Grouping constructs use parentheses( )Enclose a part of the regular expression to form a subexpression. Grouping has two main purposes:

  • catch: "save" the matched substring for later extraction or reference in replacement.
  • Scope restriction: allow quantifiers or alternation constructs to only apply to the subexpression within the group.

This part is relatively difficult to understand; you can readRegular Expression - Selection 、Lookahead and lookbehind assertions in regular expressionsHelps understanding.

The following table lists grouping constructs:

Grouping constructDescriptionModematch
( subexpression )Captures the matched subexpression and assigns it to a zero-based ordinal number.(\w)\1"ee" in "deep"
(?< name >subexpression)Captures the matched subexpression into a named group. Named groups are more readable than numeric numbering and are recommended for use in complex regular expressions.(?< double>\w)\k< double>"ee" in "deep"
(?< name1 -name2 >subexpression)Defines a balancing group definition. Used to match nested structures (such as paired parentheses), an advanced feature.(((?'Open'\()[^\(\)]*)+((?'Close-Open'\))[^\(\)]*)+)*(?(Open)(?!))$"((1-3)*(3-1))" in "3+2^((1-3)*(3-1))"
(?: subexpression)Defines a non-capturing group. Used when only grouping (limiting scope) is needed without saving matched content; performance is slightly better than capturing groups.Write(?:Line)?The "WriteLine" in "Console.WriteLine()"
(?imnsx-imnsx:subexpression)Apply or disablesubexpressionOptions specified in.A\d{2}(?i:\w+)\b"A12xl" and "A12XL" in "A12xl A12XL a12xl"
(?= subexpression)Zero-width positive lookahead assertion. Matches a position that is followed by subexpression, but does not consume characters.\w+(?=\.)"He is. The dog ran. The sun is out." in "is"、 "ran" and "out"
(?! subexpression)Zero-width negative lookahead assertion. Matches a position not followed by subexpression.\b(?!un)\w+\b"sure" and "used" in "unsure sure unity used"
(?<=subexpression)Zero-width positive lookbehind assertion. Matches a position that is immediately preceded by subexpression.(?<=19)\d{2}\b"99", "50", and "05" in "1851 1999 1950 1905 2003"
(?<! subexpression)Zero-width negative lookbehind assertion. Matches a position not preceded by subexpression.(?<!wo)man\b"man" in "Hi woman Hi man"
(?> subexpression)Non-backtracking (atomic group) subexpression. Once matched, no backtracking is allowed, which can improve performance in some scenarios.[13579](?>A+B+)"1ABB", "3ABB", and "5AB" in "1ABB 3ABBC 5AB 5AC"

Example

using System;
using System.Text.RegularExpressions;

public class Example
{
   public static void Main()
   {
      string input = "1851 1999 1950 1905 2003";
      string pattern = @"(?<=19)\d{2}\b";

      foreach (Match match in Regex.Matches(input, pattern))
         Console.WriteLine(match.Value);
   }
}


Run Example »

Qualifiers

Quantifiers specify that the preceding element (character, character class, or group) must appearhow many timesOnly then is the match considered successful.

Quantifiers are by defaultgreedy, that is, match as many characters as possible. Adding after a quantifier?can becomeLazy (non-greedy)pattern, that is, match as few characters as possible. Beginners can first master*、+、?、{n}These four basic quantifiers.

The following table lists quantifiers:

QualifiersDescriptionModematch
*Matches the previous element zero or more times. (Zero occurrences also count as a match)\d*\.\d".0"、 "19.9"、 "219.9"
+Matches the previous element one or more times. (At least once)"be+"The "bee" in "been", and the "be" in "bent"
?Matches the previous element zero or one time. (i.e., the element is optional)"rai?n""ran"、 "rain"
{ n }Matches the previous element exactly n times.",\d{3}""1,043.6" in ",043", "9,876,543,210" in ",876"、 ",543" and ",210"
{ n ,}Matches the previous element at least n times."\d{2,}""166"、 "29"、 "1930"
{ n , m }Matches the previous element at least n times, but no more than m times."\d{3,5}"The "19302" in "166", "17668", "193024"
*?Matches the previous element zero or more times, but as few times as possible (lazy mode).\d*?\.\d".0"、 "19.9"、 "219.9"
+?Matches the previous element one or more times, but as few times as possible (lazy mode)."be+?""be" in "been", "be" in "bent"
??Matches the previous element zero or one time, but as few times as possible (lazy mode)."rai??n""ran"、 "rain"
{ n }?Matches the preceding element exactly n times.",\d{3}?""1,043.6" in ",043", "9,876,543,210" in ",876"、 ",543" and ",210"
{ n ,}?Matches the previous element at least n times, but as few times as possible."\d{2,}?""166", "29" and "1930"
{ n , m }?Matches the previous element between n and m times, but as few times as possible."\d{3,5}?""166", "17668", and the "193" and "024" in "193024"

Backreference constructs

Backreferences allow, within the same regular expression, referencingContent of previously captured groupsperforms a match again. For example, using(\w)\1can find consecutive repeated characters like "ee", "ll".

The following table lists backreference constructs:

Backreference constructsDescriptionModematch
\ numberBackreference. Matches the value of a numbered subexpression.(\w)\1"ee" in "seek"
\k< name >Named backreference. Matches the value of a named expression. More readable than numeric references, recommended.(?< char>\w)\k< char>"ee" in "seek"

Alternation construct

Alternation constructs use the vertical bar|Implements "or" logic, allowing the regular expression to match any one of multiple candidate patterns. Similar to the||Operators.

The following table lists alternation constructs:

Alternation constructDescriptionModematch
|Matches any one of the elements separated by the vertical bar (|) character.th(e|is|at)The "the" and "this" in "this is the day. "
(?( expression )yes | no )If the regular expression pattern is matched by the specified expression, then matches theyes; otherwise match the optionalnopart. expression is interpreted as a zero-width assertion.(?(A)A\d{2}\b|\b\d{3}\b)The "A10" and "910" in "A10 C103 910"
(?( name )yes | no )If name or a named or numbered capturing group has a match, then matches theyes; otherwise match the optionalno。(?< quoted>")?(?(quoted).+?"|\S+\s)"Dogs.jpg "Yiska playing.jpg"" in Dogs.jpg and "Yiska playing.jpg"

replace

Replacement syntax is used forRegex.Replace()method'sReplacement Pattern StringIn, you can use$Use numbers or names to reference previously captured group content, enabling flexible text reorganization.

The following table lists characters used for substitution:

CharacterDescriptionModeReplacement patternInput stringresult string
$numberReplacement by groupnumberThe matched substring.\b(\w+)(\s)(\w+)\b$3$2$1"one two""two one"
${name}Replace by named groupnameThe matched substring.\b(?< word1>\w+)(\s)(?< word2>\w+)\b${word2} ${word1}"one two""two one"
$$The replacement character "$".\b(\d+)\s?USD$$$1"103 USD""$103"
$&Replaces a copy of the entire match.(\$*(\d*(\.+\d+)?){1})**$&"$1.30""**$1.30"
$`Replaces all text of the input string before the match.B+$`"AABBCC""AAAACC"
$'Replaces all text of the input string after the match.B+$'"AABBCC""AACCCC"
$+Replaces the last captured group.B+(C+)$+"AABBCCDD"AACCDD
$_Replaces the entire input string.B+$_"AABBCC""AAAABBCCCC"

Miscellaneous construct

The following table lists miscellaneous constructs:

ConstructorDescriptionExample
(?imnsx-imnsx)Sets or disables options such as case-insensitivity in the middle of a pattern.\bA(?i)b\w+\b matches "ABA" and "Able" in "ABA Able Act"
(?#comment)Inline comment. This comment terminates at the first closing parenthesis.\bA(?#matches a word starting with A)\w+\b
# [End of line]This comment starts with an unescaped # and continues to the end of the line.(?x)\bA\w+\b#Matches words starting with A

Regex Class

The Regex class is the core class for using regular expressions in .NET, located inSystem.Text.RegularExpressionsunder the namespace. Before using it, you need to add at the top of the file:

using System.Text.RegularExpressions;

The following table lists some commonly used methods in the Regex class:

Serial NumberMethod & Description
1public bool IsMatch( string input )
Determines whether the input string contains content that matches the regular expression pattern. Commonly used in form validation, such as validating phone number or email formats.
2public bool IsMatch( string input, int startat )
Performs match determination starting from the specified position in the string.
3public static bool IsMatch( string input, string pattern )
Static method; no need to create a Regex object first. Pass the pattern string directly for match determination. Suitable for one-off use cases.
4public MatchCollection Matches( string input )
Searches the input string for all matches, returns a MatchCollection, which can be iterated with foreach over each match result.
5public string Replace( string input, string replacement )
Replaces all content in the input string that matches the regular expression pattern with the specified string.
6public string[] Split( string input )
Splits the input string into an array of substrings according to the delimiters defined by the regular expression pattern. More flexible than string.Split(), supporting complex delimiter rules.

For the complete list of properties of the Regex class, please refer to the Microsoft C# documentation.

Example 1

The following example matches words starting with 'S':

Analysis:\bis a word boundary,Smatches the uppercase letter S,\S*Matches zero or more non-whitespace characters; combined, this means "complete words starting with S".

Example

using System;
using System.Text.RegularExpressions;

namespace RegExApplication
{
   class Program
   {
      private static void showMatch(string text, string expr)
      {
         Console.WriteLine("The Expression: " + expr);
         MatchCollection mc = Regex.Matches(text, expr);
         foreach (Match m in mc)
         {
            Console.WriteLine(m);
         }
      }
      static void Main(string[] args)
      {
         string str = "A Thousand Splendid Suns";

         Console.WriteLine("Matching words that start with 'S': ");
         showMatch(str, @"\bS\S*");
         Console.ReadKey();
      }
   }
}

When the above code is compiled and executed, it produces the following result:

Matching words that start with 'S':
The Expression: \bS\S*
Splendid
Suns

Example 2

The following example matches words starting with 'm' and ending with 'e':

Analysis:\bmmatches words starting with m,\S*matches any non-whitespace character in the middle,e\bRequires ending with e followed by a word boundary; together, it finds all words starting with m and ending with e.

Example

using System;
using System.Text.RegularExpressions;

namespace RegExApplication
{
   class Program
   {
      private static void showMatch(string text, string expr)
      {
         Console.WriteLine("The Expression: " + expr);
         MatchCollection mc = Regex.Matches(text, expr);
         foreach (Match m in mc)
         {
            Console.WriteLine(m);
         }
      }
      static void Main(string[] args)
      {
         string str = "make maze and manage to measure it";

         Console.WriteLine("Matching words start with 'm' and ends with 'e':");
         showMatch(str, @"\bm\S*e\b");
         Console.ReadKey();
      }
   }
}

When the above code is compiled and executed, it produces the following result:

Matching words start with 'm' and ends with 'e':
The Expression: \bm\S*e\b
make
maze
manage
measure

Example 3

The following example replaces extra spaces:

Analysis:\\s+Matches one or more consecutive whitespace characters (spaces, tabs, etc.) and replaces them all with a single space, thereby merging extra whitespace.

Example

using System;
using System.Text.RegularExpressions;

namespace RegExApplication
{
   class Program
   {
      static void Main(string[] args)
      {
         string input = "Hello   World   ";
         string pattern = "\\s+";
         string replacement = " ";
         Regex rgx = new Regex(pattern);
         string result = rgx.Replace(input, replacement);

         Console.WriteLine("Original String: {0}", input);
         Console.WriteLine("Replacement String: {0}", result);    
         Console.ReadKey();
      }
   }
}

When the above code is compiled and executed, it produces the following result:

Original String: Hello   World   
Replacement String: Hello World   

GeneratedRegex: source generator optimization (.NET 7+)

Beginning with .NET 7, C# introduced[GeneratedRegex]attribute (source generator), which is a major upgrade to the traditionalRegexan important upgrade to the class, specifically forPerformance-sensitiveorRepeated useregex scenarios.

Problems with traditional approaches

Use traditionalnew Regex(pattern)When ..., the parsing and compilation of the regular expression occurs atruntime, each creation has overhead. Although you can useRegexOptions.CompiledImproves runtime speed, but compilation still occurs at runtime and increases startup time and memory usage.

Traditional way

using System.Text.RegularExpressions;

// Method 1: Re-parse every time (slowest)
bool isMatch = Regex.IsMatch(input, @"\d{4}-\d{2}-\d{2}");

// Method 2: static field + Compiled (common optimization practice)
private static readonly Regex DateRegex = new Regex(@"\d{4}-\d{2}-\d{2}", RegexOptions.Compiled);

Usage of GeneratedRegex

Usage[GeneratedRegex]With the attribute, the compiler willCompilation phasegenerate the regular expression directly as efficient C# code, without runtime parsing or compilation.

Usage requirements: .NET 7 or higher is required, and the method must bepartialmethod, the containing class must also bepartialclass.

GeneratedRegex basic usage

using System;
using System.Text.RegularExpressions;

namespace RegExApplication
{
   // The class must be declared as partial
   partial class Program
   {
      // Use the [GeneratedRegex] attribute, the compiler automatically generates the implementation
      [GeneratedRegex(@"\d{4}-\d{2}-\d{2}")]
      private static partial Regex DateRegex();

      // You can also add options, such as ignoring case
      [GeneratedRegex(@"\bS\S*", RegexOptions.IgnoreCase)]
      private static partial Regex StartsWithSRegex();

      static void Main(string[] args)
      {
         string input = "Today is 2024-06-18, next event: 2025-01-01";

         // Use it like calling a normal method, returns a Regex instance
         foreach (Match m in DateRegex().Matches(input))
         {
            Console.WriteLine("Found date:" + m.Value);
         }

         // Verify if it matches
         Console.WriteLine(StartsWithSRegex().IsMatch("Splendid")); // True
      }
   }
}

When the above code is compiled and executed, it produces the following result:

找到日期:2024-06-18
找到日期:2025-01-01
True

GeneratedRegex vs Traditional Regex

The following table compares the two approaches from multiple dimensions to help you choose the appropriate usage:

Comparison DimensionTraditional RegexGeneratedRegex (source generator)Recommended
Minimum version requirement.NET Framework / .NET Core All Versions Supported.NET 7 and above—
Compile timingRuntime parsing and compilationAt compile time (build time), the Roslyn source generator generates codeGeneratedRegex
Execution performanceNormal: slower; with Compiled option: faster, but high startup overheadFastest, close to the performance of hand-written code, and no startup overheadGeneratedRegex
Memory usageCompiled mode generates IL code and has higher memory usageIt generates ordinary C# code, which is more memory-friendlyGeneratedRegex
AOT CompatibilityRegexOptions.Compiled is incompatible with AOT (ahead-of-time compilation)Fully compatible with Native AOT, suitable for publishing standalone applicationsGeneratedRegex
Code readabilityThe pattern string is written directly in the code, which is more intuitiveRequires defining a partial method, and the code structure is slightly more complexTraditional Regex (simple scenarios)
Dynamic modeSupported, the pattern can be a runtime variableNot supported, the pattern must be a compile-time constantTraditional Regex (dynamic scenarios)
Debugging supportOrdinaryThe generated code can be viewed and debugged directly, making it more transparentGeneratedRegex
Applicable scenariosOne-off use, dynamic patterns, legacy .NET projectsHigh-frequency calls, performance-sensitive, AOT publishing, new .NET projectsDepends on the context

Usage suggestions

  • If you use.NET 7+and the regex pattern is fixed,Prefer to use[GeneratedRegex]。
  • If the regex pattern needs to be dynamically built at runtime (e.g., generated from user input), only the traditionalnew Regex(pattern)。
  • If the project needs to supportNative AOTWhen publishing (e.g., .NET 8's AOT mode), be sure to use[GeneratedRegex]rather thanRegexOptions.Compiled。
  • For beginners, you can first use the traditional method to learn and practice. After understanding the principles of regex, gradually migrate to[GeneratedRegex]。

Example: Using GeneratedRegex to validate email format

using System;
using System.Text.RegularExpressions;

partial class EmailValidator
{
   // Define the email validation regex, ignoring case
   [GeneratedRegex(@"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$", RegexOptions.IgnoreCase)]
   private static partial Regex EmailRegex();

   public static bool IsValidEmail(string email)
   {
      return EmailRegex().IsMatch(email);
   }

   static void Main(string[] args)
   {
      string[] emails = { "[email protected]", "bad-email", "[email protected]", "missing@dot" };
      foreach (var email in emails)
      {
         Console.WriteLine($"{email}: {(IsValidEmail(email) ? "valid" : "Invalid")}");
      }
   }
}

When the above code is compiled and executed, it produces the following result:

[email protected]: 有效
bad-email: 无效
[email protected]: 有效
missing@dot: 无效
other extensions