Programming and IT
Python: regular expressions
A regular expression is a pattern that describes a set of strings. In Python you work with them through the `re` module from the standard library. You write the pattern as a string with an `r` prefix, search with `re.search` and pull the result out of a match object. Everything else is details of the pattern language.
In this article
When a pattern is the right tool#
A regex wins where the structure of the text is fuzzy: pulling every date out of a
paragraph, checking the format of a code, finding repeated spaces. If a string is
separated by commas or tabs, a plain split is faster, reads better and does not
break because of one stray bracket — see the overview of
Python string methods. Do not parse HTML or JSON with
regexes at all: there are ready-made libraries for that.
Your first search#
=
=
# 1207
# (6, 10)
# ['1207', '21', '09', '2026']
The r prefix turns off Python's own processing of backslashes: without it you
would have to write "\d" as "\\d". Make a habit of always using r.
Four functions cover almost everything:
re.search— the first match anywhere in the string, otherwiseNone;re.match— a match only at the start of the string;re.fullmatch— the whole string;re.findall— a list of all matches.
# None — the string starts with a letter
# True
Mixing up search and match is the most common reason for "the pattern is
right, but it finds nothing".
What a pattern is made of#
| Syntax | Meaning |
|---|---|
. |
any character except a newline |
\d \w \s |
a digit; a letter, digit or _; a whitespace character |
\D \W \S |
the same, negated |
[abc] [a-z] |
any character from a set or range |
[^abc] |
any character outside the set |
* + ? |
zero or more, one or more, zero or one |
{2} {2,5} |
exactly two, from two to five |
^ $ |
start and end of the string |
| |
or |
(...) |
a group whose contents you can extract separately |
In Python 3 patterns work with Unicode by default, so \w also matches accented
and non-Latin letters: re.findall(r"\w+", "café naïve") returns both words.
Groups give you access to parts of a match:
=
# 21.09.2026
# 21
# ('21', '09', '2026')
=
# 02
Greediness is the main trap#
Quantifiers are greedy by default: they take as much as possible and then give back exactly as much as the match needs.
=
# ['<b>bold</b> and <i>italic</i>']
# ['<b>', '</b>', '<i>', '</i>']
# ['<b>', '</b>', '<i>', '</i>']
The first pattern grabbed the whole string from the first < to the last >. A
? after a quantifier makes it lazy — it takes the minimum. The third version,
with a negated set, is faster than the lazy one and usually preferable.
The second mine is right next to it: when a pattern has groups, findall returns
the contents of the groups, not the whole matches.
# [('1', '2'), ('3', '4')]
# ['2', '4']
If you need a group only for the parentheses, make it non-capturing — (?:...).
Replacing, splitting, flags#
# lots of spaces and newlines
# 2026-09-21
# ['a', 'b', 'c']
# ['one', 'two']
# ['Cat', 'cat']
In a replacement, \1 and \2 refer to groups — write the replacement string
with r too. The re.MULTILINE flag makes ^ and $ match on every line,
re.IGNORECASE ignores case, and re.DOTALL lets the dot match a newline.
If you apply one pattern in a loop, compile it once:
pattern = re.compile(r"\d+"), then pattern.findall(text).
Common mistakes#
Calling a method on None. re.search(...).group() fails with
AttributeError: 'NoneType' object has no attribute 'group' when there is no
match. Always check: m = re.search(...), then if m:. More on careful error
handling in the article on Python exceptions.
A forgotten r. Without it, "\b" is the backspace character, not a word
boundary, and the pattern silently stops matching.
An unescaped dot. \d+.\d+ matches 12x34, because a dot is any character.
For a literal dot write \., and to escape a whole substring use re.escape.
Practice: write a pattern that extracts every address of the form
word@word.domain from a text, and another one that replaces two or more spaces
in a row with a single space. Test both on a string that definitely contains both
matching pieces and similar-looking ones that should not match.
Step-by-step plan
- Find digitsFind the first number in a string with re.search and all numbers with re.findall.
- Compare search and matchApply both functions to a string where the match is not at the start and explain the None.
- Extract groupsParse a date with a three-group pattern and print the day, month and year separately.
- Greedy versus lazyRun <.*> and <.*?> on the same string with tags and write down the difference.
- Replace with backreferencesReorder the parts of a date with re.sub using \1 and \3.
Start learning this in your own space
The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.
Check yourself
1.How many items does re.findall(r"\d+", "order 1207 on 21.09.2026") return?
2.What does re.match(r"\d+", "order 1207") return?
3.How many items are in the list returned by re.findall(r"<.*>", "<b>x</b>")?
4.Which quantifier makes a search lazy — matching as little as possible?
Sources
-
The re module in the Python docsFull pattern syntax, functions and flagsfree
-
Regular Expression HOWTOThe official step-by-step introduction with examplesfree
-
regex101Test patterns online with a Python flavour and an explanation of each partfree
Was this helpful?