Programming and IT

Python: regular expressions

A regular expression is a pattern that describes a set of strings. In Python you work with them through the `re` module from the standard library. You write the pattern as a string with an `r` prefix, search with `re.search` and pull the result out of a match object. Everything else is details of the pattern language.

Updated
In this article

When a pattern is the right tool#

A regex wins where the structure of the text is fuzzy: pulling every date out of a paragraph, checking the format of a code, finding repeated spaces. If a string is separated by commas or tabs, a plain split is faster, reads better and does not break because of one stray bracket — see the overview of Python string methods. Do not parse HTML or JSON with regexes at all: there are ready-made libraries for that.

import re

text = "order 1207 on 21.09.2026"
m = re.search(r"\d+", text)
print(m.group())        # 1207
print(m.span())         # (6, 10)
print(re.findall(r"\d+", text))   # ['1207', '21', '09', '2026']

The r prefix turns off Python's own processing of backslashes: without it you would have to write "\d" as "\\d". Make a habit of always using r.

Four functions cover almost everything:

  • re.search — the first match anywhere in the string, otherwise None;
  • re.match — a match only at the start of the string;
  • re.fullmatch — the whole string;
  • re.findall — a list of all matches.
print(re.match(r"\d+", text))              # None — the string starts with a letter
print(bool(re.fullmatch(r"\d{4}", "2026")))  # True

Mixing up search and match is the most common reason for "the pattern is right, but it finds nothing".

What a pattern is made of#

Syntax Meaning
. any character except a newline
\d \w \s a digit; a letter, digit or _; a whitespace character
\D \W \S the same, negated
[abc] [a-z] any character from a set or range
[^abc] any character outside the set
* + ? zero or more, one or more, zero or one
{2} {2,5} exactly two, from two to five
^ $ start and end of the string
| or
(...) a group whose contents you can extract separately

In Python 3 patterns work with Unicode by default, so \w also matches accented and non-Latin letters: re.findall(r"\w+", "café naïve") returns both words.

Groups give you access to parts of a match:

m = re.search(r"(\d{2})\.(\d{2})\.(\d{4})", "due by 21.09.2026 inclusive")
print(m.group(0))    # 21.09.2026
print(m.group(1))    # 21
print(m.groups())    # ('21', '09', '2026')

named = re.search(r"(?P<day>\d{2})\.(?P<month>\d{2})", "01.02")
print(named.group("month"))   # 02

Greediness is the main trap#

Quantifiers are greedy by default: they take as much as possible and then give back exactly as much as the match needs.

s = "<b>bold</b> and <i>italic</i>"
print(re.findall(r"<.*>", s))    # ['<b>bold</b> and <i>italic</i>']
print(re.findall(r"<.*?>", s))   # ['<b>', '</b>', '<i>', '</i>']
print(re.findall(r"<[^>]*>", s)) # ['<b>', '</b>', '<i>', '</i>']

The first pattern grabbed the whole string from the first < to the last >. A ? after a quantifier makes it lazy — it takes the minimum. The third version, with a negated set, is faster than the lazy one and usually preferable.

The second mine is right next to it: when a pattern has groups, findall returns the contents of the groups, not the whole matches.

print(re.findall(r"(\d)(\d)", "12 34"))   # [('1', '2'), ('3', '4')]
print(re.findall(r"(?:\d)(\d)", "12 34")) # ['2', '4']

If you need a group only for the parentheses, make it non-capturing — (?:...).

Replacing, splitting, flags#

print(re.sub(r"\s+", " ", "lots   of\tspaces\nand newlines"))
  # lots of spaces and newlines
print(re.sub(r"(\d{2})\.(\d{2})\.(\d{4})", r"\3-\2-\1", "21.09.2026"))
  # 2026-09-21
print(re.split(r"[;,]\s*", "a, b; c"))   # ['a', 'b', 'c']
print(re.findall(r"^\w+", "one\ntwo", re.MULTILINE))   # ['one', 'two']
print(re.findall(r"cat", "Cat and cat", re.IGNORECASE))  # ['Cat', 'cat']

In a replacement, \1 and \2 refer to groups — write the replacement string with r too. The re.MULTILINE flag makes ^ and $ match on every line, re.IGNORECASE ignores case, and re.DOTALL lets the dot match a newline.

If you apply one pattern in a loop, compile it once: pattern = re.compile(r"\d+"), then pattern.findall(text).

Common mistakes#

Calling a method on None. re.search(...).group() fails with AttributeError: 'NoneType' object has no attribute 'group' when there is no match. Always check: m = re.search(...), then if m:. More on careful error handling in the article on Python exceptions.

A forgotten r. Without it, "\b" is the backspace character, not a word boundary, and the pattern silently stops matching.

An unescaped dot. \d+.\d+ matches 12x34, because a dot is any character. For a literal dot write \., and to escape a whole substring use re.escape.

Practice: write a pattern that extracts every address of the form word@word.domain from a text, and another one that replaces two or more spaces in a row with a single space. Test both on a string that definitely contains both matching pieces and similar-looking ones that should not match.

Step-by-step plan

  1. Find digitsFind the first number in a string with re.search and all numbers with re.findall.
  2. Compare search and matchApply both functions to a string where the match is not at the start and explain the None.
  3. Extract groupsParse a date with a three-group pattern and print the day, month and year separately.
  4. Greedy versus lazyRun <.*> and <.*?> on the same string with tags and write down the difference.
  5. Replace with backreferencesReorder the parts of a date with re.sub using \1 and \3.

Start learning this in your own space

The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.

Start the plan

Check yourself

1.How many items does re.findall(r"\d+", "order 1207 on 21.09.2026") return?

2.What does re.match(r"\d+", "order 1207") return?

3.How many items are in the list returned by re.findall(r"<.*>", "<b>x</b>")?

4.Which quantifier makes a search lazy — matching as little as possible?

Sources

Was this helpful?

More articles

Programming and IT Regular expressions A regular expression is a compact description of a set of strings. The same syntax works in grep, in your code editor, in JavaScript, PHP, Java and Python, so you only need to learn it once. This page covers the pattern language itself, not the library of any particular language. Programming and IT How to learn Python from scratch Python is a good first programming language: code reads almost like text, and the standard library covers most everyday tasks. This plan takes you from installing the interpreter to your own scripts covered by tests in about four months, at roughly an hour a day. Programming and IT Python: slicing A slice cuts a piece out of a sequence by the rule "from which index, up to which, with what step". The notation is compact, but it has two places where almost everyone slips: the right bound is not included, and with a negative step the bounds swap places. Let us go through both. Programming and IT Python: generators A generator hands out values one at a time, and only when someone asks for them. It does not keep the whole sequence in memory, so it suits large files and endless streams. The price is that a generator is single-use: you cannot loop over it twice. Programming and IT Python: lists A list is a mutable, ordered collection of values of any type. It is also the main source of surprises for beginners: assignment does not copy a list, sort returns nothing, and removing items inside a loop skips half of them. Let us go through it in order. Programming and IT Python: classes A class describes what data an object is made of and what it can do. The `__init__` constructor fills in a new object, and `self` is a reference to the object itself. Below: a minimal working class, the difference between class and instance attributes, and ways to make the code shorter.

More solutions