Python · Lesson 10 of 21
Generators, Iterators and itertools
Understand Python iterators and generators: the iterator protocol, yield, yield from, lazy pipelines and the most useful itertools functions.
- Intermediate
- 18 min read
- 4 objectives
Before this lessonLesson 9: Closures and Decorators
What you will learn
- Explain the iterator protocol behind every for loop
- Write generator functions with yield and yield from
- Chain generators into memory-friendly pipelines
- Use itertools for slicing, grouping and combining
Your Progress
0 of 21 lessons 0%
- Lessons0 / 21
- Completed0
- Est. time left~ 5 hours
Create a free account to keep your progress on every device.
Tip: pressing Next marks this lesson complete automatically.
In the comprehensions lesson you met generator expressions: swap brackets for parentheses and values are produced one at a time. This lesson opens the hood. You will see what a for loop really does, how to write your own lazy sequences with yield, and how the itertools module solves common looping problems in one line.
Why care? Laziness means you can process a 10 GB log file, an endless stream of events or a huge range of numbers while only ever holding one item in memory.
What a for loop really does
Anything you can loop over is an iterable: it has an __iter__ method that hands out an iterator. An iterator has a __next__ method that returns the next item, and raises StopIteration when there are no more. The built-ins iter() and next() call those methods for you. A for loop is just this, with the StopIteration handled quietly.
langs = ["python", "java", "sql"]
it = iter(langs)
print(next(it))
print(next(it))
print(next(it))
try:
next(it)
except StopIteration:
print("done: StopIteration")
# a list is iterable but not an iterator; the iterator is used up
print(list(it))
print(list(langs))python java sql done: StopIteration [] ['python', 'java', 'sql']
Writing an iterator class
You can implement the protocol yourself. This countdown is correct, but notice how much ceremony it needs: store state on self, update it, and raise StopIteration by hand.
class Countdown:
def __init__(self, start):
self.current = start
def __iter__(self):
return self
def __next__(self):
if self.current <= 0:
raise StopIteration
value = self.current
self.current -= 1
return value
for n in Countdown(3):
print(n)3 2 1
Generator functions and yield
A generator function is any function that contains yield. Calling it does not run the body; it returns a generator object (which is an iterator). Each next() runs the body until the next yield, hands out that value, and pauses with all local variables intact. The same countdown shrinks to four lines.
def countdown(start):
print("starting")
while start > 0:
yield start
start -= 1
print("finished")
gen = countdown(3)
print(type(gen).__name__)
print(next(gen))
print(list(gen))generator starting 3 finished [2, 1]
The word starting only prints on the first next(), proving the body is lazy. When the function body ends, Python raises StopIteration for you.
Infinite sequences and yield from
Because values are produced on demand, a generator can be infinite, as long as the consumer stops asking. yield from delegates to another iterable, yielding each of its items; it is handy for flattening nested data or splitting a generator into helpers.
def ids(prefix):
n = 1
while True: # infinite, but only runs when asked
yield f"{prefix}-{n:03}"
n += 1
order_ids = ids("ORD")
print(next(order_ids), next(order_ids), next(order_ids))
def flatten(nested):
for item in nested:
if isinstance(item, list):
yield from flatten(item)
else:
yield item
print(list(flatten([1, [2, [3, 4]], 5, [[6]]])))ORD-001 ORD-002 ORD-003 [1, 2, 3, 4, 5, 6]
Lazy pipelines
Generators compose. Each stage pulls one item from the stage before it, so the whole pipeline holds only one record in memory at a time, no matter how big the input is. Here a list of strings stands in for lines read from a large log file.
log_lines = [
"200 GET /learn",
"500 GET /api/orders",
"200 GET /blog",
"404 GET /old-page",
"500 POST /api/login",
]
def parse(lines):
for line in lines:
status, method, path = line.split()
yield {"status": int(status), "method": method, "path": path}
def errors(records):
for r in records:
if r["status"] >= 500:
yield r
def paths(records):
for r in records:
yield r["path"]
pipeline = paths(errors(parse(log_lines)))
print(type(pipeline).__name__)
print(list(pipeline))generator ['/api/orders', '/api/login']
Printing the type shows a generator: nothing has run yet. Work only happens when list() starts pulling values. In a real program you would pass an open file object as lines, since files are already lazy iterators over their lines.
itertools: slicing and combining
The itertools module is a toolbox of fast, lazy building blocks. islice takes a slice of any iterator (even an infinite one), count counts forever, chain glues iterables together, batched (Python 3.12+) splits into fixed-size chunks, and accumulate produces running totals.
from itertools import islice, count, chain, batched, accumulate
print(list(islice(count(10, 5), 4))) # 10, 15, 20, 25
print(list(chain(["a", "b"], ("c",), "de")))
print(list(batched(range(1, 8), 3)))
print(list(accumulate([120, 80, 200, 50]))) # running revenue[10, 15, 20, 25] ['a', 'b', 'c', 'd', 'e'] [(1, 2, 3), (4, 5, 6), (7,)] [120, 200, 400, 450]
itertools: grouping and combinatorics
groupby groups consecutive items with the same key, so sort by that key first. product, permutations and combinations generate every pairing without nested loops, and pairwise gives overlapping neighbours.
from itertools import groupby, product, combinations, pairwise
orders = [("Ada", 30), ("Linus", 12), ("Ada", 45), ("Grace", 20), ("Linus", 8)]
orders.sort(key=lambda o: o[0])
for name, group in groupby(orders, key=lambda o: o[0]):
print(name, sum(amount for _, amount in group))
print(list(product(["S", "M"], ["red", "blue"])))
print(list(combinations(["api", "db", "cache"], 2)))
print([b - a for a, b in pairwise([3, 7, 12, 20])])Ada 75
Grace 20
Linus 20
[('S', 'red'), ('S', 'blue'), ('M', 'red'), ('M', 'blue')]
[('api', 'db'), ('api', 'cache'), ('db', 'cache')]
[4, 5, 8]Recap
- A
forloop callsiter()once, thennext()untilStopIteration. - A function with
yieldreturns a lazy generator that pauses between values. - Generators are single-use; recreate them to iterate again.
- Chained generators form pipelines that use constant memory.
itertoolsgives lazy tools likeislice,chain,batched,groupbyandcombinations.
# Write your solution here
Finished reading? Mark this lesson complete to track your progress.
