Python · Lesson 10 of 21

Generators, Iterators and itertools

Understand Python iterators and generators: the iterator protocol, yield, yield from, lazy pipelines and the most useful itertools functions.

  • Intermediate
  • 18 min read
  • 4 objectives

Before this lessonLesson 9: Closures and Decorators

What you will learn

  • Explain the iterator protocol behind every for loop
  • Write generator functions with yield and yield from
  • Chain generators into memory-friendly pipelines
  • Use itertools for slicing, grouping and combining

Your Progress

0 of 21 lessons 0%

  • Lessons0 / 21
  • Completed0
  • Est. time left~ 5 hours

Create a free account to keep your progress on every device.

Tip: pressing Next marks this lesson complete automatically.

In the comprehensions lesson you met generator expressions: swap brackets for parentheses and values are produced one at a time. This lesson opens the hood. You will see what a for loop really does, how to write your own lazy sequences with yield, and how the itertools module solves common looping problems in one line.

Why care? Laziness means you can process a 10 GB log file, an endless stream of events or a huge range of numbers while only ever holding one item in memory.

What a for loop really does

Anything you can loop over is an iterable: it has an __iter__ method that hands out an iterator. An iterator has a __next__ method that returns the next item, and raises StopIteration when there are no more. The built-ins iter() and next() call those methods for you. A for loop is just this, with the StopIteration handled quietly.

langs = ["python", "java", "sql"]
it = iter(langs)
print(next(it))
print(next(it))
print(next(it))
try:
    next(it)
except StopIteration:
    print("done: StopIteration")

# a list is iterable but not an iterator; the iterator is used up
print(list(it))
print(list(langs))
Output
python
java
sql
done: StopIteration
[]
['python', 'java', 'sql']

Writing an iterator class

You can implement the protocol yourself. This countdown is correct, but notice how much ceremony it needs: store state on self, update it, and raise StopIteration by hand.

class Countdown:
    def __init__(self, start):
        self.current = start

    def __iter__(self):
        return self

    def __next__(self):
        if self.current <= 0:
            raise StopIteration
        value = self.current
        self.current -= 1
        return value

for n in Countdown(3):
    print(n)
Output
3
2
1

Generator functions and yield

A generator function is any function that contains yield. Calling it does not run the body; it returns a generator object (which is an iterator). Each next() runs the body until the next yield, hands out that value, and pauses with all local variables intact. The same countdown shrinks to four lines.

def countdown(start):
    print("starting")
    while start > 0:
        yield start
        start -= 1
    print("finished")

gen = countdown(3)
print(type(gen).__name__)
print(next(gen))
print(list(gen))
Output
generator
starting
3
finished
[2, 1]

The word starting only prints on the first next(), proving the body is lazy. When the function body ends, Python raises StopIteration for you.

Infinite sequences and yield from

Because values are produced on demand, a generator can be infinite, as long as the consumer stops asking. yield from delegates to another iterable, yielding each of its items; it is handy for flattening nested data or splitting a generator into helpers.

def ids(prefix):
    n = 1
    while True:            # infinite, but only runs when asked
        yield f"{prefix}-{n:03}"
        n += 1

order_ids = ids("ORD")
print(next(order_ids), next(order_ids), next(order_ids))

def flatten(nested):
    for item in nested:
        if isinstance(item, list):
            yield from flatten(item)
        else:
            yield item

print(list(flatten([1, [2, [3, 4]], 5, [[6]]])))
Output
ORD-001 ORD-002 ORD-003
[1, 2, 3, 4, 5, 6]

Lazy pipelines

Generators compose. Each stage pulls one item from the stage before it, so the whole pipeline holds only one record in memory at a time, no matter how big the input is. Here a list of strings stands in for lines read from a large log file.

log_lines = [
    "200 GET /learn",
    "500 GET /api/orders",
    "200 GET /blog",
    "404 GET /old-page",
    "500 POST /api/login",
]

def parse(lines):
    for line in lines:
        status, method, path = line.split()
        yield {"status": int(status), "method": method, "path": path}

def errors(records):
    for r in records:
        if r["status"] >= 500:
            yield r

def paths(records):
    for r in records:
        yield r["path"]

pipeline = paths(errors(parse(log_lines)))
print(type(pipeline).__name__)
print(list(pipeline))
Output
generator
['/api/orders', '/api/login']

Printing the type shows a generator: nothing has run yet. Work only happens when list() starts pulling values. In a real program you would pass an open file object as lines, since files are already lazy iterators over their lines.

itertools: slicing and combining

The itertools module is a toolbox of fast, lazy building blocks. islice takes a slice of any iterator (even an infinite one), count counts forever, chain glues iterables together, batched (Python 3.12+) splits into fixed-size chunks, and accumulate produces running totals.

from itertools import islice, count, chain, batched, accumulate

print(list(islice(count(10, 5), 4)))           # 10, 15, 20, 25
print(list(chain(["a", "b"], ("c",), "de")))
print(list(batched(range(1, 8), 3)))
print(list(accumulate([120, 80, 200, 50])))      # running revenue
Output
[10, 15, 20, 25]
['a', 'b', 'c', 'd', 'e']
[(1, 2, 3), (4, 5, 6), (7,)]
[120, 200, 400, 450]

itertools: grouping and combinatorics

groupby groups consecutive items with the same key, so sort by that key first. product, permutations and combinations generate every pairing without nested loops, and pairwise gives overlapping neighbours.

from itertools import groupby, product, combinations, pairwise

orders = [("Ada", 30), ("Linus", 12), ("Ada", 45), ("Grace", 20), ("Linus", 8)]
orders.sort(key=lambda o: o[0])
for name, group in groupby(orders, key=lambda o: o[0]):
    print(name, sum(amount for _, amount in group))

print(list(product(["S", "M"], ["red", "blue"])))
print(list(combinations(["api", "db", "cache"], 2)))
print([b - a for a, b in pairwise([3, 7, 12, 20])])
Output
Ada 75
Grace 20
Linus 20
[('S', 'red'), ('S', 'blue'), ('M', 'red'), ('M', 'blue')]
[('api', 'db'), ('api', 'cache'), ('db', 'cache')]
[4, 5, 8]

Recap

  • A for loop calls iter() once, then next() until StopIteration.
  • A function with yield returns a lazy generator that pauses between values.
  • Generators are single-use; recreate them to iterate again.
  • Chained generators form pipelines that use constant memory.
  • itertools gives lazy tools like islice, chain, batched, groupby and combinations.
# Write your solution here

Finished reading? Mark this lesson complete to track your progress.

Up next · Lesson 11Classes and ObjectsModel data with classes, __init__, methods, and self.