Skip to main content
ToolsBay

DEVELOPER TOOLS

Regex Explained: Where the Syntax Quietly Misleads You

8 min read · ToolsBay editorial · Published · Updated

Just want to do it now?

Test regular expressions live with match highlighting and groups.

Open Regex Tester

Regex has a reputation for being unreadable, and it is earned — but not for the reason people usually give. The syntax is small. You can learn every character class and quantifier in an afternoon and forget none of it. What makes regular expressions hard is a behaviour you never see: when a pattern fails partway through, the engine walks back and tries a different split of the input. Nearly every regex bug worth writing about — the pattern that swallows the whole line, the pattern that freezes a tab — is backtracking doing exactly what you told it to.

So this is not a tour of the syntax. It is the places where the syntax quietly means something other than what it reads like.

The pieces, in one pass

[aeiou]   one character from the set
[^aeiou]  one character not in the set
[a-z]     one character in the range
\d        one digit — exactly [0-9]
\w        one word character — exactly [A-Za-z0-9_]
\s        one whitespace character
.         any character except a line terminator
a*        zero or more    a+  one or more
a?        zero or one     a{3}  exactly three    a{2,}  two or more

One correction while we are here, because it bites people who write validators for real names: \w in JavaScript is ASCII only. It includes the underscore and excludes every accented letter.

/^\w+$/.test('abc_1')  // true
/^\w+$/.test('naïve')  // false

If your "letters only" pattern is \w, it rejects a large share of the world's names. That is a product decision, not a syntax detail, and it is worth making on purpose.

Anchors bind to the string, not the sentence

^ and $ are usually described as start-of-line and end-of-line. That is only true with a flag set. By default they anchor to the start and end of the entire subject, newlines included:

'one\ntwo\nthree'.match(/^\w+$/g)   // null
'one\ntwo\nthree'.match(/^\w+$/gm)  // ['one', 'two', 'three']

The first returns nothing because there is no point at which the whole string is one run of word characters. Add m and the anchors start matching at every line break, which is what most people meant all along. This is the most common reason a pattern that works in a one-line test box returns nothing against a pasted log file.

The flags change what the pattern means

Four of them do real work, and none is cosmetic.

  • g — return every match instead of only the first.
  • m — ^ and $ match at each line break, as above.
  • s — let . match a newline. Without it, /a.b/ fails on "a\nb"; with it, it matches.
  • i — ignore case.

g has a side effect that surprises people the first time it costs them an afternoon. A global regex object carries a lastIndex, and each test resumes from it:

const r = /a/g;
r.test('banana');  // true  — lastIndex is now 2
r.test('banana');  // true  — lastIndex is now 4
r.test('banana');  // true  — lastIndex is now 6
r.test('banana');  // false — and lastIndex resets to 0

Same regex, same string, four different answers. If you keep a compiled global pattern in a module-level constant and call test on it in a loop, this is your bug. Either drop the g for a boolean check, or build the regex fresh each time.

Greedy is the default, and the default is usually wrong

* and + take as much as they can, and give characters back only when forced. Ask for "an angle bracket, anything, an angle bracket" and you do not get a tag:

const s = '<b>bold</b> and <i>italic</i>';

s.match(/<.*>/)[0]    // '<b>bold</b> and <i>italic</i>' — all 29 characters
s.match(/<.*?>/g)     // ['<b>', '</b>', '<i>', '</i>']
s.match(/<[^>]*>/g)   // ['<b>', '</b>', '<i>', '</i>']

.* ran to the end of the string, then backed up to the last > it could find. It matched, so the engine stopped. Adding ? makes the quantifier lazy — take as little as possible, extend only when the rest of the pattern fails — and you get the four tags.

The third line does the same job without lazy matching at all. [^>]* cannot cross a >, so there is nothing to give back and no backtracking to do. When you have a natural terminator, saying "anything but the terminator" is clearer than "anything, reluctantly" and does less work. Keep .*? for the cases where the terminator is more than one character.

Groups, and a bug you can watch happen

Parentheses capture. Numbered groups are (...), counted left to right by opening bracket; named groups are (?<name>...), which show up in the match's groups object and as $<name> in a JavaScript replacement string.

Here is the part worth doing yourself. Open the regex tester and press Example. It loads this pattern, with the g and m flags, over two lines of sample text:

(?<user>[\w.+-]+)@(?<domain>[\w-]+\.[\w.]+)

Two matches highlight. Look at the second one's domain group. It reads navy.mil. — with a trailing dot, because the source text is grace.hopper+work@navy.mil. at the end of a sentence and [\w.]+ is greedy about dots. The match looks right. The captured group is wrong, and the group is the thing you would go on to store.

That is the whole argument for a tester that lists groups separately instead of only highlighting matches: the highlight was fine. The fix is to stop letting a dot be an ordinary member of the class, and require a label after each one:

(?<user>[\w.+-]+)@(?<domain>[\w-]+(?:\.[\w-]+)+)

(?:...) is a non-capturing group. It groups for the quantifier without adding a numbered capture, which keeps your existing group numbers stable when you refactor a pattern.

The email regex everyone pastes

This one, or something close to it, is in a great many codebases:

^[\w.-]+@[\w.-]+\.[a-zA-Z]{2,}$

It is often explained as accepting co.uk at the end via [a-zA-Z]{2,}. It does not. There is no dot in that character class, so that piece can only ever match uk. john@example.co.uk passes the pattern as a whole, but the part before the final dot — the part people think of as the domain — is example.co. Split on it and you get the wrong answer on every multi-label domain.

Here is what the pattern actually does, tested:

ada@example.com              accepted
grace.hopper+work@navy.mil   REJECTED — no + in [\w.-]
.a@c.com                     accepted — leading dot in the local part
a@-.com                      accepted — domain is a single hyphen
a@b..com                     accepted — empty label between the dots
a@b.c                        rejected — one-letter TLDs do exist

The rejection is the expensive one. Plus-addressing is how a lot of people file their mail, and a signup form that refuses you+shop@gmail.com looks broken to a user doing nothing unusual. The false accepts are cheap by comparison: a bad row you find later.

No short pattern fixes this. Valid addresses under RFC 5322 include quoted local parts and comments, and that grammar is not something to reproduce on one line. The workable answer is to keep the pattern deliberately loose — treat it as a typo check, not a validator — and prove the address by sending mail to it. <input type="email"> applies a defined, permissive check for free, and it accepts the plus.

Catastrophic backtracking, measured

Nested quantifiers are the failure that takes production with it. (a+)+$ reads as redundant rather than dangerous. Feed it a run of as that almost matches — ending in one character that cannot — and the engine must try every way of dividing those as between the inner + and the outer + before it can conclude failure.

Timing new RegExp('^(a+)+$').test('a'.repeat(n) + '!') on Node 25:

n = 20      10 ms
n = 22      39 ms
n = 24     165 ms
n = 26     774 ms
n = 28   2,800 ms
n = 30  13,003 ms

Roughly double per character added. Thirty characters of input is thirteen seconds; forty would be measured in hours. The flat version, ^a+$, clears a million characters in about two milliseconds — the cost is the nesting, not the length.

The absolute numbers depend on the machine and the engine. The shape does not. On a server this is a denial-of-service vector with a name, ReDoS, and the input that fires it is short enough to fit in a URL.

The tell is a quantifier applied to a group that already quantifies the same characters: (a+)+, (\d*)*, (\s|\s)+. Rewrite the inner one so there is only one way to split the input.

Where to stop

A regular expression cannot correctly parse a format that nests. HTML, JSON and CSV all have quoting and nesting rules that need a real parser — not because regex is weak in general, but because a fixed pattern cannot count arbitrarily deep. Every attempt becomes a pattern that handles your sample and breaks on a comment, an escaped quote, or an attribute containing >.

For those, use something that understands the grammar: the JSON formatter and HTML beautifier parse the document rather than pattern-matching over it. Save the regex for what it is genuinely good at, which is finding and pulling things out of flat text.

And test the pattern against text you did not write. The failure is almost never a syntax error. It is a pattern that works on the three examples you tried and quietly mishandles the fourth — and that is the one that ships.

Tools covered in this guide