In the standard CPython interpreter, the Global Interpreter Lock (GIL) lets only one thread run Python code at a time. Threads still help for I/O-bound work (downloads, file and database access) because the GIL is released while waiting.
For CPU-bound work (number crunching, image processing, parsing big files) use multiple processes: each has its own interpreter and GIL, so they truly run in parallel on multiple cores.
concurrent.futures gives one simple API for both: ThreadPoolExecutor and ProcessPoolExecutor. Code that starts processes must sit under if __name__ == "__main__":.
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def count_primes(limit: int) -> int:
count = 0
for n in range(2, limit):
if all(n % d for d in range(2, int(n ** 0.5) + 1)):
count += 1
return count
def fake_download(i: int) -> int:
time.sleep(0.2) # I/O wait: the GIL is released
return i
def timed(label: str, fn) -> None:
start = time.perf_counter()
result = fn()
print(f"{label:<22}{time.perf_counter() - start:6.2f}s {result}")
if __name__ == "__main__":
jobs = [150_000] * 4
timed("CPU, sequential", lambda: [count_primes(n) for n in jobs])
with ProcessPoolExecutor() as pool:
timed("CPU, processes", lambda: list(pool.map(count_primes, jobs)))
timed("I/O, sequential", lambda: len([fake_download(i) for i in range(10)]))
with ThreadPoolExecutor(max_workers=10) as pool:
timed("I/O, threads", lambda: len(list(pool.map(fake_download, range(10)))))Key points
- I/O-bound → threads or asyncio. CPU-bound → processes.
concurrent.futuresoffers the same API for thread and process pools.- Guard process-pool code with
if __name__ == "__main__":.
Exercise
Write a script that resizes or hashes (with hashlib.sha256) every file in a folder. Time it sequentially, with a thread pool and with a process pool, and explain the results.
Show solution
Try the exercise yourself first — then compare your approach with this one.
Hashing reads each file (I/O) and then computes SHA-256 (CPU). hashlib releases the GIL while hashing large buffers, so threads already help here; processes also help but pay a start-up and data-transfer cost. For pure-Python CPU work (like the prime counting in the lesson) only processes give a real speed-up.
import hashlib
import os
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
from pathlib import Path
def sha256_of(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()[:12]
def timed(label: str, fn) -> None:
start = time.perf_counter()
results = fn()
print(f"{label:<12} {time.perf_counter() - start:6.2f}s first hash {results[0]}")
if __name__ == "__main__":
folder = Path("files")
folder.mkdir(exist_ok=True)
for i in range(8):
(folder / f"file{i}.bin").write_bytes(os.urandom(20 * 1024 * 1024)) # 8 x 20 MB
paths = sorted(folder.glob("*.bin"))
timed("sequential", lambda: [sha256_of(p) for p in paths])
with ThreadPoolExecutor() as pool:
timed("threads", lambda: list(pool.map(sha256_of, paths)))
with ProcessPoolExecutor() as pool:
timed("processes", lambda: list(pool.map(sha256_of, paths)))