The re Module
The re Module
Section titled “The re Module”Introduction
Section titled “Introduction”Python’s re module provides comprehensive regular expression support with functions for searching, matching, replacing, and splitting strings.
Core Functions
Section titled “Core Functions”re.search() — Find First Match
Section titled “re.search() — Find First Match”import re
text = "Contact: support@example.com or admin@test.org"match = re.search(r"[\w.+-]+@[\w-]+\.[\w.-]+", text)if match: print(match.group()) # support@example.com print(match.start()) # 9 (start index) print(match.end()) # 29 (end index) print(match.span()) # (9, 29)re.findall() — Find All Matches
Section titled “re.findall() — Find All Matches”emails = re.findall(r"[\w.+-]+@[\w-]+\.[\w.-]+", text)print(emails) # ['support@example.com', 'admin@test.org']re.match() — Match at Start
Section titled “re.match() — Match at Start”text = "Python is great"print(re.match(r"Python", text)) # Match objectprint(re.match(r"great", text)) # None (must match at start!)re.fullmatch() — Match Entire String
Section titled “re.fullmatch() — Match Entire String”print(re.fullmatch(r"\d{3}-\d{4}", "555-1234")) # Matchprint(re.fullmatch(r"\d{3}-\d{4}", "ext 555-1234")) # Nonere.sub() — Search and Replace
Section titled “re.sub() — Search and Replace”text = "My number is 555-1234. Call 555-5678 too."result = re.sub(r"\d{3}-\d{4}", "[REDACTED]", text)print(result) # My number is [REDACTED]. Call [REDACTED] too.
# With callbackdef mask_number(match): return "***-****"result = re.sub(r"\d{3}-\d{4}", mask_number, text)re.split() — Split by Pattern
Section titled “re.split() — Split by Pattern”text = "apple,banana;cherry|date"parts = re.split(r"[,;|]", text)print(parts) # ['apple', 'banana', 'cherry', 'date']Compiled Patterns
Section titled “Compiled Patterns”# Compile for reuse (faster for repeated use)pattern = re.compile(r"\d{3}-\d{4}")
# Same methods on compiled patternpattern.search(text)pattern.findall(text)pattern.sub("[REDACTED]", text)
# Flags during compilationpattern = re.compile(r"^hello", re.IGNORECASE)print(pattern.search("HELLO World")) # Match!# Common flags (can be combined with |)re.IGNORECASE # Case-insensitive matchingre.MULTILINE # ^ and $ match line boundariesre.DOTALL # . matches newlines toore.VERBOSE # Allow comments in patterns
# Examplesre.findall(r"^hello", "HELLO\nhello", re.IGNORECASE | re.MULTILINE)
# VERBOSE allows readable patternspattern = re.compile(r""" \b # Word boundary [A-Z][a-z]+ # First name \s+ # Whitespace [A-Z][a-z]+ # Last name \b # Word boundary""", re.VERBOSE)Greedy vs Non-greedy
Section titled “Greedy vs Non-greedy”text = "<div>Hello</div><span>World</span>"
# Greedy (default) — matches as much as possibleprint(re.findall(r"<.+>", text)) # ['<div>Hello</div><span>World</span>']
# Non-greedy — matches as little as possibleprint(re.findall(r"<.+?>", text)) # ['<div>', '</div>', '<span>', '</span>']Best Practices
Section titled “Best Practices”- Always use raw strings
r"pattern"to avoid escaping issues - Compile patterns when using the same regex multiple times
- Use
re.VERBOSEfor complex patterns to add comments - Test your regex with various inputs — edge cases matter!
- Avoid overly complex regex — sometimes string methods are simpler
Practice Exercises
Section titled “Practice Exercises”Exercise 1: Write a function that validates email addresses using regex.
Exercise 2: Create a function that extracts all hashtags from a tweet.
Exercise 3: Implement a simple template engine that replaces {{variable}} placeholders with values from a dictionary.