IntermediatePython · Lesson 3 of 9

Files, Paths, CSV & JSON

Navigate the file system with pathlib and convert between CSV and JSON.

pathlib.Path is the modern way to work with files: join paths with /, check .exists(), list files with .glob("*.csv"), and read or write small files with .read_text() / .write_text().

The csv module handles quoting and commas inside values correctly — never split CSV lines by hand. csv.DictReader gives each row as a dictionary keyed by the header.

The json module converts between Python objects and JSON text: json.dumps/json.loads for strings, json.dump/json.load for files. JSON is the standard format for APIs and config files.

convert.pyPython
import csv
import json
from pathlib import Path

data_dir = Path("data")
data_dir.mkdir(exist_ok=True)

csv_path = data_dir / "results.csv"
csv_path.write_text(
    "name,subject,score\n"
    "Amina,Maths,88\n"
    "Amina,Biology,79\n"
    'Juma,"Book-Keeping, Paper 1",42\n',
    encoding="utf-8",
)

report: dict[str, dict[str, int]] = {}
with csv_path.open(encoding="utf-8", newline="") as f:
    for row in csv.DictReader(f):
        report.setdefault(row["name"], {})[row["subject"]] = int(row["score"])

json_path = data_dir / "results.json"
json_path.write_text(json.dumps(report, indent=2), encoding="utf-8")

loaded = json.loads(json_path.read_text(encoding="utf-8"))
print(loaded["Juma"])                         # {'Book-Keeping, Paper 1': 42}
print([p.name for p in sorted(data_dir.glob("*"))])
Runs in your browser · Python

Key points

  • Use pathlib.Path instead of string paths.
  • Use the csv module — values can contain commas and quotes.
  • json.dumps(obj, indent=2) for readable output; json.loads to parse.

Exercise

Write a script that reads every .csv file in a folder, combines them, and writes a summary.json with each student's average and best subject.

Show solution

Try the exercise yourself first — then compare your approach with this one.

Path.glob("*.csv") finds every CSV file in the folder. Each row is added to a per-student dictionary of subject → score; then the summary computes each student's average and best subject and is written as JSON.

summarise.pyPython
import csv
import json
from pathlib import Path

folder = Path("results")
folder.mkdir(exist_ok=True)
(folder / "term1.csv").write_text("name,subject,score\nAmina,Maths,88\nJuma,Maths,42\n", encoding="utf-8")
(folder / "term2.csv").write_text("name,subject,score\nAmina,Biology,79\nJuma,Biology,61\n", encoding="utf-8")

scores: dict[str, dict[str, int]] = {}
for path in sorted(folder.glob("*.csv")):
    with path.open(encoding="utf-8", newline="") as f:
        for row in csv.DictReader(f):
            scores.setdefault(row["name"], {})[row["subject"]] = int(row["score"])

summary = {
    name: {
        "average": round(sum(subjects.values()) / len(subjects), 1),
        "best_subject": max(subjects, key=subjects.get),
    }
    for name, subjects in scores.items()
}

Path("summary.json").write_text(json.dumps(summary, indent=2), encoding="utf-8")
print(Path("summary.json").read_text(encoding="utf-8"))
Runs in your browser · Python

Check your understanding

  1. How do you join a folder and a file name with pathlib?

  2. Why use the csv module instead of line.split(",")?

  3. What is the difference between json.dumps and json.dump?

  4. What does csv.DictReader give you for each row?

Ask AI