Skip to main content
Text Processing

Test Your Regex Before It Ships

A regex that matches your test string can still choke on the one a user pastes in. I learned this the obvious way: an email validator worked fine until somebody entered a name with an apostrophe in it, and the pattern had never seen one.

The usual failure modes

Most broken regexes fail in one of three ways. They match too much, so a pattern meant to grab one value swallows a whole line. They match too little, because the author tested against input shaped exactly like the happy path. Or they take too long, which is the nastiest one: a nested quantifier like (a+)+ can hang a request on certain inputs, and the test suite never contains that input.

None of these are exotic. They happen because regex is hard to read, so people change it without fully understanding it, then check whether the tests pass. The tests pass. The bug ships.

Test against real input, not toy input

The fix is boring and it works. Paste the actual data into a regex tester before you write the pattern into code. Not a clean sample. The messy one: the log line with two timestamps, the address with the umlaut, the comment field with six emojis in it. Look at what matched, not just whether something did.

Two habits cover most cases:

  • Test the negative. Add a few strings that should NOT match and check they do not. Most people only ever confirm the positive.
  • Check the capture groups. A pattern can "work" while silently capturing the wrong slice, and you will not notice until the database fills with half-values.

When the pattern gets too clever

If a regex needs a comment above it to explain what it does, consider whether it should be a regex at all. Parsing an email address is a famous example. The RFC-compliant pattern is famously unreadable, and a simple /^[^@\s]+@[^@\s]+\.[^@\s]+$/ plus a confirmation email does the real job better. Readability is a feature. A colleague who can read your pattern is worth more than one that handles every edge case of an address format nobody uses.

A quick routine

Open the pattern in a tester, paste five real inputs and two inputs designed to break it, eyeball every match and capture group, then move on. Two minutes. It catches the over-matching and the wrong-capture problems almost every time, and it costs far less than a bug report three weeks later.