Programming and IT
Regular expressions
A regular expression is a compact description of a set of strings. The same syntax works in grep, in your code editor, in JavaScript, PHP, Java and Python, so you only need to learn it once. This page covers the pattern language itself, not the library of any particular language.
In this article
What a pattern actually does#
The matching engine takes a string and walks along it from left to right, trying to fit the pattern at each position. If it fits, the engine reports the piece it found and where it starts and ends; if not, it moves one character to the right and tries again. An important consequence: by default a pattern looks for a substring anywhere, it does not check the whole string. "Validate the entire input" is a separate task, solved with anchors.
Pattern characters come in two kinds. Ordinary ones mean themselves: cat finds
the letters "c", "a", "t" in a row. Metacharacters — . ^ $ * + ? ( ) [ ] { } | \
— have special meaning, and to find them literally you put a backslash in front:
\. is a dot, \$ a dollar sign, \\ the backslash itself.
Character sets#
Square brackets describe a single character from the listed ones. Inside brackets
almost all metacharacters lose their power: [.+] is a dot or a plus, no escaping
needed.
| Notation | Matches |
|---|---|
[abc] |
one letter a, b or c |
[a-z] |
any lowercase Latin letter |
[A-Za-z] |
a Latin letter in either case |
[0-9] |
a digit |
[^0-9] |
anything except a digit |
[-+*/] |
one of four signs (a hyphen placed first is literal) |
A caret ^ negates the set only right after the opening bracket; in the middle —
[a^b] — it is an ordinary character.
Common sets have short names. They exist almost everywhere except classic POSIX syntax:
| Shorthand | Equivalent | Negation |
|---|---|---|
\d |
a digit | \D |
\w |
a letter, digit or underscore | \W |
\s |
space, tab, newline | \S |
. |
any character except a newline | — |
How many times to repeat#
A quantifier applies to whatever is to its left: one character, a set in brackets or a whole group.
| Notation | Meaning |
|---|---|
? |
zero or one time |
* |
zero or more times |
+ |
one or more times |
{3} |
exactly three times |
{2,5} |
from two to five |
{2,} |
two or more |
The difference between * and + catches beginners: -* also matches nothing at
all, so such a pattern is "found" in any string. If the element is required, use
+.
A quantifier after a group repeats the whole group: (ab)+ finds "abababab",
while ab+ finds "abbbb".
Anchors and boundaries#
| Notation | Anchors to |
|---|---|
^ |
the start of the string |
$ |
the end of the string |
\b |
a word boundary |
\B |
a position that is not a boundary |
The ^…$ pair turns a substring search into a check of the whole value — this is
how form validation is written. The pattern ^\d{5}$ accepts a five-digit US ZIP
code and rejects "1234567", whereas \d{5} without anchors would find a match in
it.
\b sits between a word character and a non-word character. So \bcat\b finds the
separate word and skips "category" and "concat". It is the quickest way to fix a
search that "finds too much".
Multiline mode changes what ^ and $ mean — from the edges of the whole text to
the edges of each line. It is turned on with a flag (usually m), and without it the
pattern ^Error in a multi-line log finds one match at best.
Groups and alternation#
Round brackets do two things at once: they bind part of a pattern into one unit and remember what matched inside.
(...)— a capturing group; its contents are available by number: 1, 2, 3 — left to right by opening bracket;(?:...)— grouping without capturing, when you need the brackets only for structure;(?<name>...)— a named group, referenced by name instead of number;|— alternation, "or":cat|dog|ferret.
Alternation has the lowest precedence, so it almost always goes inside brackets.
The pattern ^yes|no$ reads as "the string starts with yes OR ends with no", while
the intended "the whole string is yes or no" is written ^(yes|no)$.
Inside the pattern itself you refer to a captured piece with \1, \2 — a
backreference. The pattern \b(\w+) \1\b finds an accidentally doubled word, and
(["']).*?\1 finds a quoted string closed with the same quote mark it was opened
with.
Greediness#
By default quantifiers grab as much as they can and only give back under pressure. The classic mistake is matching a pair of parentheses:
string: (one) and (two)
pattern: \(.*\) → finds (one) and (two)
pattern: \(.*?\) → finds (one) and separately (two)
pattern: \([^)]*\) → the same, but faster
A ? right after a quantifier makes it lazy: *?, +?, {2,5}?. The third
option — with a negated set — needs no backtracking at all and is noticeably
faster on long strings; whenever "up to the first such character" can be expressed
with a set, choose it.
Ready-made patterns#
| Task | Pattern |
|---|---|
| An integer, possibly signed | ^-?\d+$ |
| A number with a fractional part | ^-?\d+([.,]\d+)?$ |
| A date in YYYY-MM-DD format | ^\d{4}-\d{2}-\d{2}$ |
An HTML colour like #a1b2c3 |
^#[0-9a-fA-F]{6}$ |
| A phone number in international E.164 format | ^\+[1-9]\d{6,14}$ |
| A web address | https?://\S+ |
| Two or more spaces in a row | \s{2,} |
| An empty line or only spaces | ^\s*$ |
| A whole word | \b\w+\b |
Time of day deserves its own example, because the range 00:00–23:59 cannot be described with one set:
^([01]\d|2[0-3]):[0-5]\d$
Here [01]\d covers hours 00 to 19, the alternative 2[0-3] adds 20–23, and the
minutes [0-5]\d will not let ":60" through. Any numeric range is built the same
way — digit by digit, not with one expression.
The checks for email and addresses here are deliberately rough. A full
standards-compliant email regex runs to several hundred characters and still does
not prove the mailbox exists; in practice it is enough to filter obvious junk with
^\S+@\S+\.\S+$ and do the real check by sending a code by email.
Where patterns differ#
There is no single standard, only families of dialects.
grepwithout options understands basic POSIX, where+,?,|and group brackets must be escaped. The-Eoption turns on extended syntax and the pattern looks familiar:grep -E "^[0-9]{4}-" file.- The
\dand\wshorthands come from Perl. In POSIX you write[[:digit:]]and[[:alnum:]_]instead. In GNU grep they are available with-P. - Languages treat Unicode differently: in some
\wis Latin only, in others it includes accented and non-Latin letters. When in doubt, write the set explicitly. - Named groups are written
(?<name>…)in most implementations and(?P<name>…)in Python.
How to call these patterns from Python code is covered in the article on
Python regex: search, replace and flags. Searching files
from the terminal is in the Linux commands
cheat sheet, in the section on grep.
How to debug, and when not to use a regex#
Build a pattern piece by piece. First check that the simplest part matches, then add one element at a time, testing on a sample string each time. You need two test strings: one that must match, and a similar one that must not — otherwise it is easy to write a pattern that matches everything.
Long patterns are split across lines with comments: verbose mode (x) allows
whitespace and comments inside the expression and exists in most languages.
Readability beats brevity — in a month even the author will not recognise their own
one-liner.
There are tasks where a regex is the wrong tool. HTML and XML are parsed with a parser: nesting of arbitrary depth cannot, in principle, be described by a regular expression. JSON is read with a library. A string with a fixed delimiter is easier to cut with a split function than with a pattern.
Remember the cost, too: nested quantifiers like (a+)+$ on an unsuitable string
force the engine to try an exponential number of combinations, and a form check
hangs the server. The warning sign is a quantifier inside a group that has another
quantifier applied to it.
Step-by-step plan
- Sets and quantifiersBuild a pattern for an integer, then for a signed number with a fractional part, testing it on five strings.
- AnchorsTurn a substring search into a whole-value check with ^ and $; confirm it on a string with extra characters at the edges.
- GroupsSplit a date into three groups and rebuild it in reverse order with the references \1 and \3.
- GreedinessRun the greedy, lazy and negated-set versions on a string with two pairs of parentheses.
- Your own pattern for a taskWrite a time-of-day check and test it on 23:59, 24:00 and 12:60.
Start learning this in your own space
The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.
Check yourself
1.How many non-overlapping matches does the pattern \d{2} give in the string 12345?
2.Which quantifier means “zero or one time”?
3.What does [^0-9] mean?
4.How many matches does the greedy pattern a+ give in the string aaabaa?
Sources
-
Regular expressions on MDNA complete syntax guide with examples and tables of special charactersfree
-
GNU grep manualThe differences between basic, extended and Perl syntax from the sourcefree
-
Regular Expression HOWTOA step-by-step introduction that is useful beyond Pythonfree
-
regex101An online tester that explains every part of a patternfree
Was this helpful?