Regular Expressions - Position Matching

Position matching (also known as anchoring or boundary matching) refers to matching specific positions in a string, rather than actual characters. Unlike ordinary character matching, position matching does not consume any characters; it only specifies where a match must occur.

Why Position Matching Is Needed

  1. Precise positioning: allows you to precisely specify where a match occurs
  2. Efficiency improvement: avoids unnecessary full-text searches
  3. Pattern validation: checks whether a string meets specific format requirements

Common Position Matching Metacharacters

1. Line Start and Line End Matching

^- Matches the start of a line

Example

// Matches a line starting with "Hello"
const pattern = /^Hello/;
console.log(pattern.test("Hello World")); // true
console.log(pattern.test("Say Hello"));   // false

$- Matches the end of a line

Example

// Matches a line ending with "World"
const pattern = /World$/;
console.log(pattern.test("Hello World")); // true
console.log(pattern.test("World Peace")); // false

2. Word Boundary Matching

\b- Matches a word boundary

A word boundary is the position between\w([a-zA-Z0-9_]) and\Wnon-word characters, or the start/end position of a string.

Example

// Matches the standalone word "cat"
const pattern = /\bcat\b/;
console.log(pattern.test("cat"));        // true
console.log(pattern.test("concatenate")); // false
console.log(pattern.test("a cat"));      // true

\B- Matches a non-word boundary

Example

// Matches "cat" that is not at a word boundary
const pattern = /\Bcat\B/;
console.log(pattern.test("concatenate")); // true
console.log(pattern.test("cat"));        // false

3. Other Position Matching

\Aand\Z(Supported in some languages)

  • \A: Matches the start of a string (unlike^, not affected by multiline mode)
  • \Z: Matches the end of a string or before the newline at the end

Practical Applications of Position Matching

1. Validating Input Format

Example

// Validates a mobile number (starts with 1, 11 digits in total)
const phonePattern = /^1\d{10}$/;
console.log(phonePattern.test("13800138000")); // true
console.log(phonePattern.test("a13800138000")); // false

2. Extracting Words at Specific Positions

Example

// Extracts the first word of each line
const text = "Apple Banana\nCherry Date";
const firstWords = text.match(/^\w+/gm);
console.log(firstWords); // ["Apple", "Cherry"]

3. Replacing Text at Specific Positions

Example

// Adds "> " at the beginning of each paragraph
const text = "First line\nSecond line";
const result = text.replace(/^/gm, "> ");
console.log(result);
// > First line
// > Second line

Advanced Position Matching Techniques

1. Position Matching in Multiline Mode

Using themflag (multiline mode) changes^and$the behavior of:

Example

const text = "Line 1\nLine 2\nLine 3";
// Normal mode
console.log(text.match(/^Line \d/g)); // ["Line 1"]
// Multiline mode
console.log(text.match(/^Line \d/gm)); // ["Line 1", "Line 2", "Line 3"]

2. Lookarounds

Although not strictly position matching, lookarounds can help implement more complex positional conditions:

Positive lookahead (?=)

Example

// Matches numbers followed by "px"
const pattern = /\d+(?=px)/;
console.log(pattern.exec("12px")); // ["12"]

Negative lookahead (?!)

Example

// Matches numbers not followed by "px"
const pattern = /\d+(?!px)/;
console.log(pattern.exec("12em")); // ["12"]

Common Errors and Considerations

  1. Confusing^the use of ^ within a character class:

    • [^abc]Represents "a character that is not a, b, or c"
    • ^abcRepresents "a string starting with abc"
  2. Ignoring the effect of multiline mode:

    • By default^and$^ and $ match the beginning and end of the entire string
    • In multiline mode, they match the beginning and end of each line
  3. Boundary matching does not consume characters:

    // 这个模式不会匹配任何内容,因为\b不消耗字符
    const pattern = /\b\b/;

Practice Challenges

  1. Write a regular expression that matches all strings starting with "Chapter" followed by 1-2 digits
  2. Create a pattern that matches all standalone "the" words in a string (not matching "the" in "there")
  3. Write a regular expression to verify whether a string is a valid URL (starting with http:// or https://)

Summary of Key Points

Metacharacter Description Example
^ Matches the start of a line/string ^Start
$ Matches the end of a line/string end$
\b Matches a word boundary \bword\b
\B Matches a non-word boundary \Bword\B
\A Matches the start of a string (some languages) \AStart
\Z Matches the end of a string (some languages) end\Z

Mastering position matching can significantly improve the accuracy and efficiency of regular expressions. Remember:

  • Position matching does not consume characters; it only specifies where a match occurs
  • Using boundary matching appropriately can avoid unnecessary full-text searches
  • Multiline mode changes^and$the behavior of ^ and $
Other Extensions