Python105 min total · 18 parts
Python Fundamentals for Interviews: Data Structures, Comprehensions, and Gotchas
Part 1 of 18 · ~4 min
Overview
Python has one particular reputation, and it's mostly earned: the syntax looks so close to pseudocode already that you can pick up a whole feature from a five-line toy example alone. list from a grocery list. dict from a phonebook. A closure from a toy counter. A decorator from a stopwatch. The mutable-default-argument bug from a shopping basket that mysteriously remembers yesterday's groceries. Every one of those makes total sense in isolation.
Then you sit down to write an actual program, where all of that machinery has to run in the same forty lines at once — a decorator wrapping a function that closes over a mutable default that lives inside an object whose class also has a mutable default — and the five tidy demos turn out to have taught you five separate tricks instead of one language.
So here's what we're going to do instead. One small program, start to finish, and every mechanism in this reference gets built into it as it comes up — never a fresh cast of throwaway variables per section.
The program: a log watcher for a SaaS company's API gateway. It reads raw access-log lines and decides which IPs are worth an ops engineer's attention. Every line looks like this:
2026-03-02T09:14:03Z 203.0.113.42 GET /api/search 200 118
2026-03-02T09:14:04Z 198.51.100.7 POST /api/login 401 42
2026-03-02T09:14:04Z 198.51.100.7 POST /api/login 401 39
Six whitespace-separated fields: timestamp, IP, HTTP method, path, status code, duration in milliseconds. 203.0.113.42 is a perfectly normal user, searching the product a lot. 198.51.100.7 is failing to log in, over and over, which is either someone who forgot their password or the first ten seconds of a credential-stuffing attempt — and figuring out which, fast, is the entire point of the tool we're building. 198.51.100.7 is going to be the thread that runs through this whole reference: flagged early, rate-limited in the middle, and finally looked up against a geolocation service in the last chapter to see where the attack is actually coming from.
Here's the first version. It works, in the way a paper map works before you notice it doesn't fold back up correctly:
def check_logs(path):
lines = open(path).readlines() # reads the ENTIRE file into memory, never closes it
seen_ips = [] # a list — checked with `in` on every single line
counts = {}
for line in lines:
parts = line.split()
ip = parts[1]
if ip not in seen_ips: # O(n) scan against a list that only grows
seen_ips.append(ip)
if ip not in counts:
counts[ip] = 0
counts[ip] += 1
for ip, n in counts.items():
if n > 50:
print(f"{ip} made {n} requests")
Nineteen lines, and it already has a small pile of real problems. It loads the whole file into memory, so it falls over on a log too big to fit in RAM. seen_ips is a list checked with in, so every line pays for a linear scan against a list that only ever grows. The threshold 50 is baked into the function, so watching /api/login (where five failed attempts is suspicious) and /api/search (where five hundred hits an hour is a Tuesday) needs two copies of this function, not one. It has no idea what a 401 even is — a user who searched fifty-one times gets flagged exactly like an attacker who failed to log in fifty-one times. One malformed line anywhere in the file crashes the whole run. And it only ever looks at one file, when a real gateway rotates its log into a new file every few hours.
By the last chapter, this same program streams a log of any size without holding more than one line in memory at a time, tells a chatty-but-harmless user apart from a credential-stuffing attempt using the status code and the path instead of a raw count, rate-limits itself using memory that survives between calls without ever becoming a global variable, survives a corrupted line without losing the rest of the batch, and — the part that needs everything before it — looks up exactly where its worst offender is connecting from, by calling a slow external API concurrently instead of one IP at a time.
Read it front to back if you can; more than one chapter collects a detail a much earlier chapter planted on purpose. If you're here for one topic, the sidebar drops you straight in — every chapter opens by naming which version of the watcher it's working on.