Multiprocessing in Python Tutorial: Processes, Pools and Queues with Examples

Build three complete multiprocessing programs, trace how results move between processes, and learn when process startup and data transfer outweigh parallel work.

KnowledgeGate Team

Exam prep & CS education

Updated 16 Sep 20266 min read

A CPU-heavy loop does not become faster just because its function runs in a worker. A first script can repeat work or fail when its entry point and data flow are unclear. Use Process for custom workers, Pool for repeated calls, Queue for messages and results, and a lock for shared state.

Related reading: Python multithreading and threads and process creation.

What multiprocessing changes and when it is worth using

Python's official multiprocessing documentation describes separate processes that can use multiple processors. Each worker has its own interpreter and memory, while threads share a process. Threads and process creation and process states and lifecycle teach operating-system concepts, not Python syntax.

Keep a 20-item arithmetic demo sequential; startup dominates. Consider processes for CPU-heavy chunks, such as rendering eight images or evaluating four large ranges. For 100 waiting network requests, threads or asynchronous I/O may fit better. Workload size, data transfer, and available processors decide whether processes help.

Before writing worker code:

  • Put worker functions at module level.

  • Pass serialisable arguments.

  • Create processes only inside if __name__ == "__main__":.

Start-method defaults differ by Python version and platform, so do not assume one default. For a wider path, explore the Coding & DSA courses.

Start three processes and collect deterministic results with a Queue

Save this program as a .py file and run it:

python
from multiprocessing import Process, Queue

def count_primes(start, end, results):
    count = 0
    for number in range(start, end + 1):
        if number >= 2 and all(
            number % divisor != 0
            for divisor in range(2, int(number**0.5) + 1)
        ):
            count += 1
    results.put(((start, end), count))

if __name__ == "__main__":
    results = Queue()
    ranges = [(2, 20), (21, 40), (41, 60)]
    workers = [
        Process(target=count_primes, args=(*bounds, results))
        for bounds in ranges
    ]

    for worker in workers:
        worker.start()

    answers = [results.get() for _ in workers]

    for worker in workers:
        worker.join()

    print(sorted(answers))

The output is:

Code
[((2, 20), 8), ((21, 40), 4), ((41, 60), 5)]

The parent creates three workers. They find 8 primes from 2 through 20, 4 from 21 through 40, and 5 from 41 through 60. Each worker puts one range-and-count tuple on the process-safe queue. The parent makes exactly three get() calls. Arrival order can vary, but sorted(answers) makes the display deterministic. These small ranges keep the trace checkable; enlarge them only when measuring process overhead against useful CPU work.

start() launches a worker, while join() waits for it. join() does not return the target's value, so the queue transports results. A parent list does not receive mutations made to a child's separate copy.

Reuse workers with Pool and split one calculation into chunks

A pool keeps workers available for repeated calls:

python
from multiprocessing import Pool

def sum_squares(start, end):
    return sum(number * number for number in range(start, end + 1))

if __name__ == "__main__":
    ranges = [(1, 5), (6, 10), (11, 15), (16, 20)]

    with Pool(processes=4) as pool:
        subtotals = pool.starmap(sum_squares, ranges)

    print(subtotals)
    print(sum(subtotals))

The exact output is:

Code
[55, 330, 855, 1630]
2870

Squares from 1 through 5 total 55; 6 through 10 total 330; 11 through 15 total 855; and 16 through 20 total 1630. Thus, 55 + 330 + 855 + 1630 = 2870. starmap unpacks tuple arguments and collects results in input order.

This small example is traceable, not evidence of a speed-up. A pool pays for startup, dispatch, serialisation, and result transfer. Four equal independent chunks give four workers useful work; 20 one-number tasks add needless coordination here.

Four Pool workers sum the squares of ranges (1,5), (6,10), (11,15), (16,20) into [55, 330, 855, 1630], totalling 2870.

Share as little state as possible, and lock the state you must share

Prefer independent inputs and returned or queued results. Use Queue for messages, Pipe for a two-end connection, Value or Array for small primitives, and a manager only when a higher-level proxy is necessary. More sharing creates more races.

Here three workers calculate private subtotals, then lock one shared update:

python
from multiprocessing import Process, Value, Lock

def add_batch(total, lock, values):
    subtotal = sum(values)
    with lock:
        total.value += subtotal

if __name__ == "__main__":
    total = Value("i", 0)
    lock = Lock()
    batches = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]
    workers = [
        Process(target=add_batch, args=(total, lock, batch))
        for batch in batches
    ]

    for worker in workers:
        worker.start()
    for worker in workers:
        worker.join()

    print(total.value)

The locked output is 45, because the private subtotals are 6, 15, and 24, and 6 + 15 + 24 = 45. Each process sums its own batch before taking the lock, so only one shared read-add-write is protected per worker. Without the lock, two workers can read the same old total and one update can overwrite another. See process synchronisation and semaphores for the operating-system foundation; Python Lock mechanics come from the official reference.

Common multiprocessing errors and how to fix them

Five common symptoms have direct fixes:

Symptom

Cause

Fix

Repeated spawning or a bootstrapping error

Process creation is outside the main guard

Move the creation code under if __name__ == "__main__":

Can't pickle local object

The worker is nested or is a lambda

Use a module-level def

A parent list remains []

A child changed only its own memory

Return or queue (range, count) tuples

Expected work is not observed before exit

The parent neither collects nor waits

Retrieve results and join workers

A new pool appears on every loop iteration

Pool creation is inside the loop

Put one with Pool(...) around the batch

Avoid three performance traps. Time complete sequential and process versions on the real workload, not only the teaching example. Send compact immutable inputs or coarse chunks, not one huge mutable object per task. Start with at most one worker per useful chunk and measure.

For visible failures, handle errors in the parent, inspect Process exit codes, and retrieve pool results so exceptions surface. Multiprocessing does not make code deterministic, thread-safe, or fault-tolerant.

How interviews and exams test multiprocessing

First, predict the queue example. Three child processes send three messages, but raw arrival order is unknowable. Sorting must produce [((2, 20), 8), ((21, 40), 4), ((41, 60), 5)]. Why cannot join() produce the tuples? Process completion is separate from result transport.

Next, trace a chunk: sum_squares(2, 4) returns 2^2 + 3^2 + 4^2 = 4 + 9 + 16 = 29. Pool.starmap(sum_squares, [(1, 2), (3, 4)]) returns [5, 25] because it preserves input-task order.

Finally, consider a labelled measurement. If a CPU job takes 8.0 seconds sequentially and 2.8 seconds with four processes, speed-up is 8.0 / 2.8 = 2.857...x, displayed as 2.86x. Using that displayed value, efficiency is 2.86 / 4 = 71.5% (71.4% if full precision is kept). Startup, serialisation, coordination, and non-parallel work explain why it is below 4x. Use the official reference for API and platform-specific behaviour.

Exercises, short version and the next step

Try each change before reading its answer:

  1. Add the range (61, 80) as a fourth process. Check: its queued tuple is ((61, 80), 5) for primes 61, 67, 71, 73, 79.

  2. Split cubes from 1 through 12 into (1,3), (4,6), (7,9), (10,12). Check: subtotals are [36, 405, 1584, 4059], and 36 + 405 + 1584 + 4059 = 6084.

  3. Change the shared-total batches to [2, 4], [6, 8], and [10, 12]. Check: private subtotals 6 + 14 + 22 produce a locked total of 42.

Use Process for a few custom workers, Pool for many calls to one top-level function, Queue for messages and results, and Value plus Lock only for minimal shared primitive state. Stay sequential for tiny work, or use waiting-oriented concurrency for mostly I/O. Measurement decides whether processes are worthwhile.

Your next action is to build the Python foundation and practise language concepts. Save the pool example as a script, predict its four subtotals on paper, run it, then increase the workload and time both sequential and pooled versions.