A regular expression (regex) describes a text pattern. \d matches a digit, \w a letter, digit or underscore, \s whitespace, . any character; + means one or more, * zero or more, {8} exactly eight, and [67] either 6 or 7. Write patterns as raw strings (r"...") so backslashes stay as typed.
re.fullmatch checks that a whole string fits a pattern (validation); re.search finds the first match anywhere; re.findall returns all matches; re.sub replaces matches. Parentheses create groups you can read back — name them with (?P<name>...) for clarity.
Compile patterns you reuse with re.compile. Keep regexes simple: for structured formats such as CSV or JSON, use the proper module instead.
import re
TZ_PHONE = re.compile(r"(?:\+255|0)([67]\d{8})")
for raw in ["0712345678", "+255754000111", "0812345678", "071234567"]:
m = TZ_PHONE.fullmatch(raw)
print(f"{raw:<15}", f"valid -> 0{m.group(1)}" if m else "invalid")
text = "Amina scored 88 in Maths, Juma 42 in Maths and Neema 71 in Biology."
print(re.findall(r"\d+", text)) # ['88', '42', '71']
pattern = re.compile(r"(?P<name>[A-Z][a-z]+)(?: scored)? (?P<score>\d+) in (?P<subject>\w+)")
for m in pattern.finditer(text):
print(m.group("name"), m.group("subject"), int(m.group("score")))
messy = "Juma Said , Form 4"
print(re.sub(r"\s+", " ", messy).replace(" ,", ",")) # Juma Said, Form 4
print(re.sub(r"(\d{4})(\d{3})(\d{3})", r"\1 \2 \3", "0712345678")) # 0712 345 678
print(bool(re.search(r"\bexam\b", "Mock exams start Monday"))) # False: 'exams' is a different wordKey points
- Use raw strings for patterns;
\d,\w,\s,+,*,{n}and[...]are the core pieces. fullmatchvalidates,searchfinds,findall/finditerextract,subreplaces.- Named groups
(?P<name>...)make extraction readable; compile patterns you reuse.
Exercise
Write valid_email(text) using re.fullmatch (something@something.something, no spaces). Then extract every date written like 14/03/2026 from a paragraph with re.findall and convert each to ISO format (2026-03-14) with re.sub and groups.
Show solution
Try the exercise yourself first — then compare your approach with this one.
The email pattern requires non-space characters, an @, more non-space characters, a dot, and a final part. For the dates, three capture groups (day, month, year) are rearranged in the replacement with \\3-\\2-\\1.
import re
EMAIL = re.compile(r"[^@\s]+@[^@\s]+\.[^@\s]+")
def valid_email(text: str) -> bool:
return EMAIL.fullmatch(text) is not None
for e in ["amina@example.com", "juma@school", "bad email@x.com", "neema.k@shule.ac.tz"]:
print(f"{e:<22} {valid_email(e)}")
paragraph = "Registration closes on 14/03/2026, mocks start 02/11/2026 and results come 15/12/2026."
dates = re.findall(r"\b\d{2}/\d{2}/\d{4}\b", paragraph)
iso = [re.sub(r"(\d{2})/(\d{2})/(\d{4})", r"\3-\2-\1", d) for d in dates]
print(iso) # ['2026-03-14', '2026-11-02', '2026-12-15']