Python · Lesson 16 of 21
Regular Expressions
Learn Python regular expressions with the re module: patterns, character classes, groups, named groups, findall, sub, flags and common pitfalls.
- Intermediate
- 18 min read
- 4 objectives
Before this lessonLesson 15: Error Handling
What you will learn
- Write patterns with character classes, quantifiers and anchors
- Choose between search, match, fullmatch, findall and finditer
- Extract data with groups and named groups
- Clean and transform text with re.sub and flags
Your Progress
0 of 21 lessons 0%
- Lessons0 / 21
- Completed0
- Est. time left~ 5 hours
Create a free account to keep your progress on every device.
Tip: pressing Next marks this lesson complete automatically.
String methods like split, replace and startswith handle fixed text well. But what about "find every order ID that looks like ORD- followed by digits" or "check this looks like a date"? For patterns rather than exact text, you use regular expressions (regex), a small language for describing text shapes. Python supports them through the built-in re module.
Regex has a reputation for being cryptic. The trick is to learn a handful of building blocks and always test patterns on real examples, which is exactly what this lesson does.
Your first pattern and raw strings
re.search(pattern, text) scans the text and returns a Match object for the first hit, or None. \d means "any digit" and + means "one or more". Always write patterns as raw strings (r"...") so Python does not interpret the backslashes before re sees them.
import re
text = "Your order ORD-4821 ships Tuesday."
m = re.search(r"ORD-\d+", text)
print(m)
print(m.group(), m.start(), m.end())
print(re.search(r"INV-\d+", text)) # no match gives None<re.Match object; span=(11, 19), match='ORD-4821'> ORD-4821 11 19 None
The building blocks
Nearly every pattern is built from these pieces:
.any character except newline;\ddigit,\wletter, digit or underscore,\swhitespace (uppercase\D \W \Smean the opposite).[abc]one of these characters,[a-z0-9]ranges,[^,]anything except a comma.- Quantifiers:
*zero or more,+one or more,?optional,{3}exactly three,{2,4}two to four. - Anchors:
^start,$end,\bword boundary. a|beither side;( )groups;\.a literal dot (escape special characters).
import re
samples = ["2026-09-23", "26-9-23", "on 2026-01-05 we shipped", "2026-13-99"]
date = r"\b\d{4}-\d{2}-\d{2}\b"
for s in samples:
print(f"{s!r:30} ->", bool(re.search(date, s)))
print(re.findall(r"[A-Z]{2,}", "Use SQL, not CSV or Excel, for BIG data"))
print(re.findall(r"colou?r", "color colour colr"))'2026-09-23' -> True '26-9-23' -> False 'on 2026-01-05 we shipped' -> True '2026-13-99' -> True ['SQL', 'CSV', 'BIG'] ['color', 'colour']
Notice that 2026-13-99 matches: regex checks shape, not meaning. To validate a real date, parse it with datetime.strptime after the regex check.
search, match, fullmatch, findall, finditer
Pick the function by the question you are asking. match only looks at the start of the string, fullmatch requires the whole string to fit (perfect for validation), findall returns all matches as a list of strings, and finditer yields Match objects lazily when you need positions or groups.
import re
code = "A7-2026"
print(bool(re.search(r"\d{4}", code))) # anywhere
print(bool(re.match(r"\d{4}", code))) # only at the start
print(bool(re.fullmatch(r"[A-Z]\d-\d{4}", code)))
log = "GET /learn 200 12ms, GET /api 500 340ms, GET /blog 200 8ms"
print(re.findall(r"\d+ms", log))
for m in re.finditer(r"/\w+", log):
print(m.group(), "at", m.start())True False True ['12ms', '340ms', '8ms'] /learn at 4 /api at 25 /blog at 45
Groups: extracting parts of a match
Parentheses capture part of the match so you can pull it out. Numbered groups are fine for short patterns, but named groups (?P<name>...) make code far more readable. When a pattern has groups, findall returns tuples of the groups instead of whole matches.
import re
line = "2026-09-23 14:05:11 ERROR payment failed for user=42"
pattern = r"(?P<date>\d{4}-\d{2}-\d{2}) (?P<time>[\d:]+) (?P<level>[A-Z]+) (?P<msg>.*)"
m = re.match(pattern, line)
print(m.group("level"))
print(m.group(1), m.group(2))
print(m.groupdict())
pairs = re.findall(r"(\w+)=(\w+)", "user=42 plan=pro region=eu")
print(pairs)
print(dict(pairs))ERROR
2026-09-23 14:05:11
{'date': '2026-09-23', 'time': '14:05:11', 'level': 'ERROR', 'msg': 'payment failed for user=42'}
[('user', '42'), ('plan', 'pro'), ('region', 'eu')]
{'user': '42', 'plan': 'pro', 'region': 'eu'}Greedy versus lazy
Quantifiers are greedy: they grab as much as possible and only give back what is needed. That often surprises people when extracting text between delimiters. Add ? after a quantifier (*?, +?) to make it lazy, matching as little as possible. A negated class like [^"]* is often clearer still.
import re
html = '<a href="/learn">Learn</a> <a href="/blog">Blog</a>'
print(re.findall(r'href="(.*)"', html)) # greedy: too much
print(re.findall(r'href="(.*?)"', html)) # lazy
print(re.findall(r'href="([^"]*)"', html)) # negated class['/learn">Learn</a> <a href="/blog'] ['/learn', '/blog'] ['/learn', '/blog']
Replacing text with re.sub
re.sub(pattern, replacement, text) replaces every match. The replacement can refer to groups with \1 or \g<name>, or it can be a function that receives each Match and returns the new text. re.split splits on a pattern instead of a fixed string.
import re
print(re.sub(r"\s+", " ", "too many\n\t spaces"))
print(re.sub(r"(\d{4})-(\d{2})-(\d{2})", r"\3/\2/\1", "due 2026-09-23"))
def mask(m):
return m.group()[:2] + "*" * (len(m.group()) - 2)
print(re.sub(r"\b\d{6,}\b", mask, "card 4111222233334444, ref 12"))
def slugify(title):
return re.sub(r"[^a-z0-9]+", "-", title.lower()).strip("-")
print(slugify("Closures & Decorators: A Guide!"))
print(re.split(r"[;,]\s*", "python, java;sql, go"))too many spaces due 23/09/2026 card 41**************, ref 12 closures-decorators-a-guide ['python', 'java', 'sql', 'go']
Flags and compiled patterns
Flags change how matching works: re.IGNORECASE ignores case, re.MULTILINE makes ^ and $ match at each line, and re.VERBOSE lets you spread a pattern over several lines with comments. If you reuse a pattern, re.compile gives you a pattern object with the same methods (Python caches recent patterns anyway, so this is mainly about readability).
import re
email = re.compile(r"""
^[\w.+-]+ # user part
@
[\w-]+ # domain
(\.[\w-]+)+$ # one or more .tld parts
""", re.VERBOSE)
for addr in ["ada@example.com", "bad@", "linus.t+dev@mail.co.uk"]:
print(addr, bool(email.match(addr)))
notes = "todo: write tests\nTODO: ship\ndone: docs"
print(re.findall(r"^todo: (.*)$", notes, re.IGNORECASE | re.MULTILINE))ada@example.com True bad@ False linus.t+dev@mail.co.uk True ['write tests', 'ship']
Recap
- Always write patterns as raw strings:
r"\d+". - Learn the core pieces: character classes, quantifiers, anchors, groups and alternation.
- Use
fullmatchto validate,searchto find one,findall/finditerto find many. - Named groups and
groupdict()make extraction readable; add?for lazy matching. re.subwith group references or a function handles most text clean-up.
# Write your solution here
Finished reading? Mark this lesson complete to track your progress.
