IntermediatePython · Lesson 4 of 9

Regular Expressions

Match, validate, extract and replace text patterns with the re module.

A regular expression (regex) describes a text pattern. \d matches a digit, \w a letter, digit or underscore, \s whitespace, . any character; + means one or more, * zero or more, {8} exactly eight, and [67] either 6 or 7. Write patterns as raw strings (r"...") so backslashes stay as typed.

re.fullmatch checks that a whole string fits a pattern (validation); re.search finds the first match anywhere; re.findall returns all matches; re.sub replaces matches. Parentheses create groups you can read back — name them with (?P<name>...) for clarity.

Compile patterns you reuse with re.compile. Keep regexes simple: for structured formats such as CSV or JSON, use the proper module instead.

patterns.pyPython
import re

TZ_PHONE = re.compile(r"(?:\+255|0)([67]\d{8})")

for raw in ["0712345678", "+255754000111", "0812345678", "071234567"]:
    m = TZ_PHONE.fullmatch(raw)
    print(f"{raw:<15}", f"valid -> 0{m.group(1)}" if m else "invalid")

text = "Amina scored 88 in Maths, Juma 42 in Maths and Neema 71 in Biology."
print(re.findall(r"\d+", text))                                  # ['88', '42', '71']

pattern = re.compile(r"(?P<name>[A-Z][a-z]+)(?: scored)? (?P<score>\d+) in (?P<subject>\w+)")
for m in pattern.finditer(text):
    print(m.group("name"), m.group("subject"), int(m.group("score")))

messy = "Juma   Said ,   Form   4"
print(re.sub(r"\s+", " ", messy).replace(" ,", ","))              # Juma Said, Form 4

print(re.sub(r"(\d{4})(\d{3})(\d{3})", r"\1 \2 \3", "0712345678"))   # 0712 345 678
print(bool(re.search(r"\bexam\b", "Mock exams start Monday")))  # False: 'exams' is a different word
Runs in your browser · Python

Key points

  • Use raw strings for patterns; \d, \w, \s, +, *, {n} and [...] are the core pieces.
  • fullmatch validates, search finds, findall/finditer extract, sub replaces.
  • Named groups (?P<name>...) make extraction readable; compile patterns you reuse.

Exercise

Write valid_email(text) using re.fullmatch (something@something.something, no spaces). Then extract every date written like 14/03/2026 from a paragraph with re.findall and convert each to ISO format (2026-03-14) with re.sub and groups.

Show solution

Try the exercise yourself first — then compare your approach with this one.

The email pattern requires non-space characters, an @, more non-space characters, a dot, and a final part. For the dates, three capture groups (day, month, year) are rearranged in the replacement with \\3-\\2-\\1.

regex_solution.pyPython
import re

EMAIL = re.compile(r"[^@\s]+@[^@\s]+\.[^@\s]+")


def valid_email(text: str) -> bool:
    return EMAIL.fullmatch(text) is not None


for e in ["amina@example.com", "juma@school", "bad email@x.com", "neema.k@shule.ac.tz"]:
    print(f"{e:<22} {valid_email(e)}")

paragraph = "Registration closes on 14/03/2026, mocks start 02/11/2026 and results come 15/12/2026."
dates = re.findall(r"\b\d{2}/\d{2}/\d{4}\b", paragraph)
iso = [re.sub(r"(\d{2})/(\d{2})/(\d{4})", r"\3-\2-\1", d) for d in dates]
print(iso)   # ['2026-03-14', '2026-11-02', '2026-12-15']
Runs in your browser · Python

Check your understanding

  1. What does \d{8} match?

  2. Which function checks that the WHOLE string matches a pattern?

  3. Why write patterns as raw strings, e.g. r"\d+"?

  4. In re.sub(r"(\d{4})(\d{3})", r"\2-\1", s), what does \2 insert?

Ask AI