Python105 min total · 18 parts
Python Fundamentals for Interviews: Data Structures, Comprehensions, and Gotchas
Part 17 of 18 · ~5 min
The GIL, Threading, and Async Basics
Somewhere inside the standard CPython interpreter sits a single lock — the Global Interpreter Lock, or GIL — and holding it is a prerequisite for running any Python bytecode at all. Exactly one thread can hold it at any instant, which means exactly one thread can be executing Python code at any instant, full stop, regardless of how many CPU cores the machine underneath has sitting idle. This one fact is responsible for more confused performance debugging than almost anything else in the language.
Side one of it: try to parse several rotated log files faster by splitting them across threads.
import threading
def parse_all(paths):
for path in paths:
Report.from_file(path) # mostly CPU: string splitting, int conversion, object creation
Hand four paths to four threads expecting a four-times speedup and you won't get one. The GIL forces those four threads to take turns holding it, so at no point are two of them actually executing Python bytecode simultaneously — they're interleaved, not parallel, and the constant handoff between them costs a bit extra on top. Four threads parsing sequentially inside one process finishes around the same time as, or slightly worse than, one thread parsing the same four files back to back.
| Watcher task | Right tool | What actually changes |
|---|---|---|
| Parsing thousands of lines, scoring entries | multiprocessing | Each process gets its own interpreter and its own GIL, so the cores genuinely run at the same time |
| Waiting on a network response | threading or asyncio | Whichever thread or coroutine is waiting isn't holding the GIL, freeing it up for whatever isn't |
Side two: asking a slow, external service where each flagged IP actually connects from. Waiting on a network round trip is exactly the kind of work concurrency was built for:
import asyncio, random
async def lookup_geo(ip):
await asyncio.sleep(0.4) # the actual network round trip
if random.random() < 0.3:
raise ConnectionError(f"geo lookup timed out for {ip}")
return {"ip": ip, "country": "RO"}
It's tempting to reuse @retry from the decorators chapter directly — don't, at least not unmodified:
@retry(times=3) # the SYNC retry from the decorators chapter
async def lookup_geo(ip):
...
This silently does nothing useful. retry's wrapper calls func(*args, **kwargs) and expects either a return value or an exception — but calling an async def function doesn't run its body at all, it just builds a coroutine object and hands it back immediately, without raising anything. wrapper's try block "succeeds" on the very first attempt, no matter what, because there's nothing yet to fail. It returns that unexecuted coroutine straight through — and only once the caller actually awaits it does the real work run, and any exception surface, long after retry's own except has already exited. The retry logic never even sees it. The fix is an async-aware version:
def async_retry(times):
def decorator(func):
@functools.wraps(func)
async def wrapper(*args, **kwargs):
last_error = None
for attempt in range(times):
try:
return await func(*args, **kwargs) # await, not just call
except Exception as e:
last_error = e
raise last_error
return wrapper
return decorator
@async_retry(times=3)
async def lookup_geo(ip):
...
Now the flagged IPs, looked up one at a time:
async def enrich_sequential(ips):
results = {}
for ip in ips:
results[ip] = await lookup_geo(ip)
return results
# three IPs, ~0.4s each, one after another: roughly 1.2 seconds
and concurrently, which is the entire reason asyncio exists:
async def enrich_concurrent(ips):
results = await asyncio.gather(*(lookup_geo(ip) for ip in ips), return_exceptions=True)
return dict(zip(ips, results))
# the same three IPs, all in flight together: roughly 0.4 seconds, not 1.2
None of this needs more than one thread, because await isn't a passive pause — it's an active handoff. The moment lookup_geo hits await asyncio.sleep(0.4), it tells the event loop "nothing for me to do until this finishes, go run someone else," and the loop does exactly that, switching to whichever other coroutine is ready. That's cooperative: each coroutine chooses its own yield points, unlike an OS thread scheduler that can interrupt a thread at essentially any instruction. asyncio.gather sets all three lookup_geo calls in motion up front rather than one at a time, so their 0.4-second waits overlap almost completely — the batch finishes close to whenever the single slowest one does, not after all three durations have been added together.
One more detail, and it closes a loop from several chapters back: hammering a third party's API with every flagged IP at once is a good way to get rate-limited by them. The make_rate_limiter closure built for inbound traffic works exactly as well pointed the other direction — same tool, same deque, just checked before each outbound call this time instead of after each inbound one.
That closure counts requests over time, which is what a real rate limit is. If the actual goal is simpler — never more than N lookups in flight at once, regardless of timing — asyncio has a purpose-built tool that's worth knowing instead of reaching for a hand-rolled one:
async def enrich_limited(ips, max_concurrent=3):
sem = asyncio.Semaphore(max_concurrent)
async def bounded_lookup(ip):
async with sem: # blocks here once 3 lookups are already in flight
return await lookup_geo(ip)
results = await asyncio.gather(*(bounded_lookup(ip) for ip in ips), return_exceptions=True)
return dict(zip(ips, results))
Semaphore(3) lets three coroutines through async with sem: at once; a fourth arriving early simply waits at that line until one of the first three releases it on exit. Between the two, make_rate_limiter's deque is the right tool for "no more than N per unit of time," and Semaphore is the right tool for "no more than N happening simultaneously" — genuinely different constraints that happen to look similar from a distance.