Programming and IT

Regular expressions

A regular expression is a compact description of a set of strings. The same syntax works in grep, in your code editor, in JavaScript, PHP, Java and Python, so you only need to learn it once. This page covers the pattern language itself, not the library of any particular language.

Updated
In this article

What a pattern actually does#

The matching engine takes a string and walks along it from left to right, trying to fit the pattern at each position. If it fits, the engine reports the piece it found and where it starts and ends; if not, it moves one character to the right and tries again. An important consequence: by default a pattern looks for a substring anywhere, it does not check the whole string. "Validate the entire input" is a separate task, solved with anchors.

Pattern characters come in two kinds. Ordinary ones mean themselves: cat finds the letters "c", "a", "t" in a row. Metacharacters — . ^ $ * + ? ( ) [ ] { } | \ — have special meaning, and to find them literally you put a backslash in front: \. is a dot, \$ a dollar sign, \\ the backslash itself.

Character sets#

Square brackets describe a single character from the listed ones. Inside brackets almost all metacharacters lose their power: [.+] is a dot or a plus, no escaping needed.

Notation Matches
[abc] one letter a, b or c
[a-z] any lowercase Latin letter
[A-Za-z] a Latin letter in either case
[0-9] a digit
[^0-9] anything except a digit
[-+*/] one of four signs (a hyphen placed first is literal)

A caret ^ negates the set only right after the opening bracket; in the middle — [a^b] — it is an ordinary character.

Common sets have short names. They exist almost everywhere except classic POSIX syntax:

Shorthand Equivalent Negation
\d a digit \D
\w a letter, digit or underscore \W
\s space, tab, newline \S
. any character except a newline —

How many times to repeat#

A quantifier applies to whatever is to its left: one character, a set in brackets or a whole group.

Notation Meaning
? zero or one time
* zero or more times
+ one or more times
{3} exactly three times
{2,5} from two to five
{2,} two or more

The difference between * and + catches beginners: -* also matches nothing at all, so such a pattern is "found" in any string. If the element is required, use +.

A quantifier after a group repeats the whole group: (ab)+ finds "abababab", while ab+ finds "abbbb".

Anchors and boundaries#

Notation Anchors to
^ the start of the string
$ the end of the string
\b a word boundary
\B a position that is not a boundary

The ^…$ pair turns a substring search into a check of the whole value — this is how form validation is written. The pattern ^\d{5}$ accepts a five-digit US ZIP code and rejects "1234567", whereas \d{5} without anchors would find a match in it.

\b sits between a word character and a non-word character. So \bcat\b finds the separate word and skips "category" and "concat". It is the quickest way to fix a search that "finds too much".

Multiline mode changes what ^ and $ mean — from the edges of the whole text to the edges of each line. It is turned on with a flag (usually m), and without it the pattern ^Error in a multi-line log finds one match at best.

Groups and alternation#

Round brackets do two things at once: they bind part of a pattern into one unit and remember what matched inside.

  • (...) — a capturing group; its contents are available by number: 1, 2, 3 — left to right by opening bracket;
  • (?:...) — grouping without capturing, when you need the brackets only for structure;
  • (?<name>...) — a named group, referenced by name instead of number;
  • | — alternation, "or": cat|dog|ferret.

Alternation has the lowest precedence, so it almost always goes inside brackets. The pattern ^yes|no$ reads as "the string starts with yes OR ends with no", while the intended "the whole string is yes or no" is written ^(yes|no)$.

Inside the pattern itself you refer to a captured piece with \1, \2 — a backreference. The pattern \b(\w+) \1\b finds an accidentally doubled word, and (["']).*?\1 finds a quoted string closed with the same quote mark it was opened with.

Greediness#

By default quantifiers grab as much as they can and only give back under pressure. The classic mistake is matching a pair of parentheses:

string:   (one) and (two)
pattern:  \(.*\)     → finds  (one) and (two)
pattern:  \(.*?\)    → finds  (one)  and separately  (two)
pattern:  \([^)]*\)  → the same, but faster

A ? right after a quantifier makes it lazy: *?, +?, {2,5}?. The third option — with a negated set — needs no backtracking at all and is noticeably faster on long strings; whenever "up to the first such character" can be expressed with a set, choose it.

Ready-made patterns#

Task Pattern
An integer, possibly signed ^-?\d+$
A number with a fractional part ^-?\d+([.,]\d+)?$
A date in YYYY-MM-DD format ^\d{4}-\d{2}-\d{2}$
An HTML colour like #a1b2c3 ^#[0-9a-fA-F]{6}$
A phone number in international E.164 format ^\+[1-9]\d{6,14}$
A web address https?://\S+
Two or more spaces in a row \s{2,}
An empty line or only spaces ^\s*$
A whole word \b\w+\b

Time of day deserves its own example, because the range 00:00–23:59 cannot be described with one set:

^([01]\d|2[0-3]):[0-5]\d$

Here [01]\d covers hours 00 to 19, the alternative 2[0-3] adds 20–23, and the minutes [0-5]\d will not let ":60" through. Any numeric range is built the same way — digit by digit, not with one expression.

The checks for email and addresses here are deliberately rough. A full standards-compliant email regex runs to several hundred characters and still does not prove the mailbox exists; in practice it is enough to filter obvious junk with ^\S+@\S+\.\S+$ and do the real check by sending a code by email.

Where patterns differ#

There is no single standard, only families of dialects.

  • grep without options understands basic POSIX, where +, ?, | and group brackets must be escaped. The -E option turns on extended syntax and the pattern looks familiar: grep -E "^[0-9]{4}-" file.
  • The \d and \w shorthands come from Perl. In POSIX you write [[:digit:]] and [[:alnum:]_] instead. In GNU grep they are available with -P.
  • Languages treat Unicode differently: in some \w is Latin only, in others it includes accented and non-Latin letters. When in doubt, write the set explicitly.
  • Named groups are written (?<name>…) in most implementations and (?P<name>…) in Python.

How to call these patterns from Python code is covered in the article on Python regex: search, replace and flags. Searching files from the terminal is in the Linux commands cheat sheet, in the section on grep.

How to debug, and when not to use a regex#

Build a pattern piece by piece. First check that the simplest part matches, then add one element at a time, testing on a sample string each time. You need two test strings: one that must match, and a similar one that must not — otherwise it is easy to write a pattern that matches everything.

Long patterns are split across lines with comments: verbose mode (x) allows whitespace and comments inside the expression and exists in most languages. Readability beats brevity — in a month even the author will not recognise their own one-liner.

There are tasks where a regex is the wrong tool. HTML and XML are parsed with a parser: nesting of arbitrary depth cannot, in principle, be described by a regular expression. JSON is read with a library. A string with a fixed delimiter is easier to cut with a split function than with a pattern.

Remember the cost, too: nested quantifiers like (a+)+$ on an unsuitable string force the engine to try an exponential number of combinations, and a form check hangs the server. The warning sign is a quantifier inside a group that has another quantifier applied to it.

Step-by-step plan

  1. Sets and quantifiersBuild a pattern for an integer, then for a signed number with a fractional part, testing it on five strings.
  2. AnchorsTurn a substring search into a whole-value check with ^ and $; confirm it on a string with extra characters at the edges.
  3. GroupsSplit a date into three groups and rebuild it in reverse order with the references \1 and \3.
  4. GreedinessRun the greedy, lazy and negated-set versions on a string with two pairs of parentheses.
  5. Your own pattern for a taskWrite a time-of-day check and test it on 23:59, 24:00 and 12:60.

Start learning this in your own space

The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.

Start the plan

Check yourself

1.How many non-overlapping matches does the pattern \d{2} give in the string 12345?

2.Which quantifier means “zero or one time”?

3.What does [^0-9] mean?

4.How many matches does the greedy pattern a+ give in the string aaabaa?

Sources

Was this helpful?

More articles

Programming and IT Python: regular expressions A regular expression is a pattern that describes a set of strings. In Python you work with them through the `re` module from the standard library. You write the pattern as a string with an `r` prefix, search with `re.search` and pull the result out of a match object. Everything else is details of the pattern language. Programming and IT Git commands Git has well over a hundred commands, but on an ordinary day you use about twenty. Here they are in the order you need them — from your first copy of a repository to undoing a bad step — with one line of explanation each. Programming and IT Linux commands The Linux terminal works the same way in every distribution: command name, options, arguments. Below are six tables by area of work, each with one line per command and an example you can type to see the result straight away. Programming and IT Markdown tables A Markdown table is built from pipes and a separator line under the header. Below is the table syntax with column alignment, plus a short cheat sheet for the rest of the markup — headings, lists, links, images, code, quotes and task lists. Programming and IT Arduino for beginners Arduino is a microcontroller board you can tell when to light an LED and when to read a sensor. You write the program in simplified C++, upload it over USB, and it runs on its own, without a computer. Your first blinking LED takes about twenty minutes. Programming and IT How to learn Python from scratch Python is a good first programming language: code reads almost like text, and the standard library covers most everyday tasks. This plan takes you from installing the interpreter to your own scripts covered by tests in about four months, at roughly an hour a day.

More solutions