Python · Lesson 16 of 21

Regular Expressions

Learn Python regular expressions with the re module: patterns, character classes, groups, named groups, findall, sub, flags and common pitfalls.

  • Intermediate
  • 18 min read
  • 4 objectives

Before this lessonLesson 15: Error Handling

What you will learn

  • Write patterns with character classes, quantifiers and anchors
  • Choose between search, match, fullmatch, findall and finditer
  • Extract data with groups and named groups
  • Clean and transform text with re.sub and flags

Your Progress

0 of 21 lessons 0%

  • Lessons0 / 21
  • Completed0
  • Est. time left~ 5 hours

Create a free account to keep your progress on every device.

Tip: pressing Next marks this lesson complete automatically.

String methods like split, replace and startswith handle fixed text well. But what about "find every order ID that looks like ORD- followed by digits" or "check this looks like a date"? For patterns rather than exact text, you use regular expressions (regex), a small language for describing text shapes. Python supports them through the built-in re module.

Regex has a reputation for being cryptic. The trick is to learn a handful of building blocks and always test patterns on real examples, which is exactly what this lesson does.

Your first pattern and raw strings

re.search(pattern, text) scans the text and returns a Match object for the first hit, or None. \d means "any digit" and + means "one or more". Always write patterns as raw strings (r"...") so Python does not interpret the backslashes before re sees them.

import re

text = "Your order ORD-4821 ships Tuesday."
m = re.search(r"ORD-\d+", text)
print(m)
print(m.group(), m.start(), m.end())

print(re.search(r"INV-\d+", text))   # no match gives None
Output
<re.Match object; span=(11, 19), match='ORD-4821'>
ORD-4821 11 19
None

The building blocks

Nearly every pattern is built from these pieces:

  • . any character except newline; \d digit, \w letter, digit or underscore, \s whitespace (uppercase \D \W \S mean the opposite).
  • [abc] one of these characters, [a-z0-9] ranges, [^,] anything except a comma.
  • Quantifiers: * zero or more, + one or more, ? optional, {3} exactly three, {2,4} two to four.
  • Anchors: ^ start, $ end, \b word boundary.
  • a|b either side; ( ) groups; \. a literal dot (escape special characters).
import re

samples = ["2026-09-23", "26-9-23", "on 2026-01-05 we shipped", "2026-13-99"]
date = r"\b\d{4}-\d{2}-\d{2}\b"
for s in samples:
    print(f"{s!r:30} ->", bool(re.search(date, s)))

print(re.findall(r"[A-Z]{2,}", "Use SQL, not CSV or Excel, for BIG data"))
print(re.findall(r"colou?r", "color colour colr"))
Output
'2026-09-23'                   -> True
'26-9-23'                      -> False
'on 2026-01-05 we shipped'     -> True
'2026-13-99'                   -> True
['SQL', 'CSV', 'BIG']
['color', 'colour']

Notice that 2026-13-99 matches: regex checks shape, not meaning. To validate a real date, parse it with datetime.strptime after the regex check.

search, match, fullmatch, findall, finditer

Pick the function by the question you are asking. match only looks at the start of the string, fullmatch requires the whole string to fit (perfect for validation), findall returns all matches as a list of strings, and finditer yields Match objects lazily when you need positions or groups.

import re

code = "A7-2026"
print(bool(re.search(r"\d{4}", code)))       # anywhere
print(bool(re.match(r"\d{4}", code)))        # only at the start
print(bool(re.fullmatch(r"[A-Z]\d-\d{4}", code)))

log = "GET /learn 200 12ms, GET /api 500 340ms, GET /blog 200 8ms"
print(re.findall(r"\d+ms", log))
for m in re.finditer(r"/\w+", log):
    print(m.group(), "at", m.start())
Output
True
False
True
['12ms', '340ms', '8ms']
/learn at 4
/api at 25
/blog at 45

Groups: extracting parts of a match

Parentheses capture part of the match so you can pull it out. Numbered groups are fine for short patterns, but named groups (?P<name>...) make code far more readable. When a pattern has groups, findall returns tuples of the groups instead of whole matches.

import re

line = "2026-09-23 14:05:11 ERROR payment failed for user=42"
pattern = r"(?P<date>\d{4}-\d{2}-\d{2}) (?P<time>[\d:]+) (?P<level>[A-Z]+) (?P<msg>.*)"
m = re.match(pattern, line)
print(m.group("level"))
print(m.group(1), m.group(2))
print(m.groupdict())

pairs = re.findall(r"(\w+)=(\w+)", "user=42 plan=pro region=eu")
print(pairs)
print(dict(pairs))
Output
ERROR
2026-09-23 14:05:11
{'date': '2026-09-23', 'time': '14:05:11', 'level': 'ERROR', 'msg': 'payment failed for user=42'}
[('user', '42'), ('plan', 'pro'), ('region', 'eu')]
{'user': '42', 'plan': 'pro', 'region': 'eu'}

Greedy versus lazy

Quantifiers are greedy: they grab as much as possible and only give back what is needed. That often surprises people when extracting text between delimiters. Add ? after a quantifier (*?, +?) to make it lazy, matching as little as possible. A negated class like [^"]* is often clearer still.

import re

html = '<a href="/learn">Learn</a> <a href="/blog">Blog</a>'
print(re.findall(r'href="(.*)"', html))      # greedy: too much
print(re.findall(r'href="(.*?)"', html))     # lazy
print(re.findall(r'href="([^"]*)"', html))   # negated class
Output
['/learn">Learn</a> <a href="/blog']
['/learn', '/blog']
['/learn', '/blog']

Replacing text with re.sub

re.sub(pattern, replacement, text) replaces every match. The replacement can refer to groups with \1 or \g<name>, or it can be a function that receives each Match and returns the new text. re.split splits on a pattern instead of a fixed string.

import re

print(re.sub(r"\s+", " ", "too    many\n\t spaces"))
print(re.sub(r"(\d{4})-(\d{2})-(\d{2})", r"\3/\2/\1", "due 2026-09-23"))

def mask(m):
    return m.group()[:2] + "*" * (len(m.group()) - 2)
print(re.sub(r"\b\d{6,}\b", mask, "card 4111222233334444, ref 12"))

def slugify(title):
    return re.sub(r"[^a-z0-9]+", "-", title.lower()).strip("-")
print(slugify("Closures & Decorators: A Guide!"))

print(re.split(r"[;,]\s*", "python, java;sql,  go"))
Output
too many spaces
due 23/09/2026
card 41**************, ref 12
closures-decorators-a-guide
['python', 'java', 'sql', 'go']

Flags and compiled patterns

Flags change how matching works: re.IGNORECASE ignores case, re.MULTILINE makes ^ and $ match at each line, and re.VERBOSE lets you spread a pattern over several lines with comments. If you reuse a pattern, re.compile gives you a pattern object with the same methods (Python caches recent patterns anyway, so this is mainly about readability).

import re

email = re.compile(r"""
    ^[\w.+-]+        # user part
    @
    [\w-]+           # domain
    (\.[\w-]+)+$     # one or more .tld parts
""", re.VERBOSE)

for addr in ["ada@example.com", "bad@", "linus.t+dev@mail.co.uk"]:
    print(addr, bool(email.match(addr)))

notes = "todo: write tests\nTODO: ship\ndone: docs"
print(re.findall(r"^todo: (.*)$", notes, re.IGNORECASE | re.MULTILINE))
Output
ada@example.com True
bad@ False
linus.t+dev@mail.co.uk True
['write tests', 'ship']

Recap

  • Always write patterns as raw strings: r"\d+".
  • Learn the core pieces: character classes, quantifiers, anchors, groups and alternation.
  • Use fullmatch to validate, search to find one, findall/finditer to find many.
  • Named groups and groupdict() make extraction readable; add ? for lazy matching.
  • re.sub with group references or a function handles most text clean-up.
# Write your solution here

Finished reading? Mark this lesson complete to track your progress.

Up next · Lesson 17Modules, Packages and pipSplit code across files, import the standard library, and install packages in a virtual environment.