Skip to content
Languages

Python asyncio TaskGroup vs gather: Failure and Cancellation in Practice

Two concurrent checks support one request: inventory fails while pricing is still pending. Should pricing continue, or does the whole operation fail together? TaskGroup and gather give different answers. The example uses events instead of timing guesses to show which task owns pending work. It requires Python 3.11 or later.

By Published

4 min readEditorial analysisUpdated
  • Python
  • asyncio
  • Concurrency
  • Backend interviews
What to remember

Choose a failure policy before choosing an API. TaskGroup scopes related tasks; gather can collect independent results. Cancellation still needs cooperative cleanup and does not reverse external side effects.

Decide what a sibling failure should mean

A complete quote may be useless without inventory; independent widgets may benefit from partial results. Choose that product behavior first. The table concerns an ordinary child failure such as ValueError; child cancellation and cancellation of the parent operation are distinct cases.

Policies for concurrent work
ChoiceChild failureUseful for
TaskGroupCancels group siblingsRelated subtasks
gather()Propagates; siblings runExplicit continuation
gather(return_exceptions=True)Collects result or errorPartial-result policy

Run a deterministic sibling-failure experiment

Save this as one Python file and run it. slow signals that it has started, then waits on a gate. fail cannot raise before that signal. TaskGroup cancels slow and waits for its cleanup before handling the grouped ValueError. In the gather version, we deliberately release and await the still-running sibling.

Expected lines: slow cleanup; group caught; group cancelled: True; gather cancelled: False; slow cleanup; slow result. The two cleanup messages have different causes: cancellation in the group, ordinary completion in gather. The explicit gate makes the distinction observable without depending on which task happens to win a sleep race.

import asyncio

async def slow(started, release):
    try:
        started.set()
        await release.wait()
        return "slow result"
    finally:
        print("slow cleanup")

async def fail(started):
    await started.wait()
    raise ValueError("stock")

async def taskgroup_demo():
    started, release = asyncio.Event(), asyncio.Event()
    try:
        async with asyncio.TaskGroup() as group:
            sibling = group.create_task(slow(started, release))
            group.create_task(fail(started))
    except* ValueError:
        print("group caught")
    print("group cancelled:", sibling.cancelled())

async def gather_demo():
    started, release = asyncio.Event(), asyncio.Event()
    sibling = asyncio.create_task(slow(started, release))
    try:
        await asyncio.gather(sibling, fail(started))
    except ValueError:
        print("gather cancelled:", sibling.cancelled())
    release.set()
    print(await sibling)

async def main():
    await taskgroup_demo()
    await gather_demo()

asyncio.run(main())

Do not confuse a raised error with cancellation

Default gather propagates an ordinary failure without canceling siblings. Canceling that already-completed gather does not stop them. Canceling a pending gather instead requests cancellation of unfinished children. Distinguish these timelines when debugging.

Ending asyncio.run immediately after catching the error can trigger shutdown cancellation of remaining tasks, concealing gather's continuation behavior. Our example explicitly awaits the sibling before main returns. If continuation is intentional in real code, keep references, observe results, and define who supervises that work beyond the request's failure.

Treat collected exceptions as decisions, not successful data

With return_exceptions=True, gather waits for results and returns failures alongside successful values in input order. For independent widgets, assign each result to its widget and decide whether to show data, an error or a retry. Do not send that mixed list to a consumer expecting only successful payloads.

For TaskGroup, except* ValueError handles matching errors within a group; unrelated errors remain unhandled. Use this when you have a defined response to that error family. Logging and swallowing every grouped failure would make an incomplete operation look successful.

Make cancellation cooperate with resource cleanup

Put resource cleanup in finally or an appropriate context manager. If you explicitly catch CancelledError to record an event or release something, normally re-raise it. It inherits from BaseException, so except Exception does not catch it. Suppressing cancellation can interfere with the surrounding operation's shutdown policy.

Cancellation is a request delivered cooperatively, not an instant kill. A coroutine that blocks the event loop prevents timely progress. Nor does cancellation undo an accepted payment or committed write; those require their own transaction, idempotency or compensation design. TaskGroup supervises tasks created through the group, not arbitrary detached tasks you create elsewhere.

An interview exercise with three failure timelines

First predict the printed lines without running the file. Then add return_exceptions=True and move release.set() before await asyncio.gather(...), so the gated sibling can finish; inspect both returned entries. Finally create a pending gather of two gated workers, cancel the gather, await its CancelledError and confirm both workers ran cleanup.

Explain which behavior suits a complete quote and which suits independent widgets. Use the Python interview overview to revisit coroutines and tasks, then state the remaining limitation: this experiment proves task ownership and cleanup, not the rollback of external operations.

Quick answers

Frequently asked questions

Does gather cancel the other tasks when one fails?

An ordinary failure under default gather is propagated while other children continue. Canceling a pending gather is different and requests cancellation of its unfinished children.

When should I use TaskGroup?

Use it when tasks belong to one scoped operation and ordinary sibling failure should stop the remaining work. Define how the resulting exception group is handled and ensure tasks cooperate with cancellation.

Can TaskGroup replace gather(return_exceptions=True)?

They express different policies. Collected errors may support deliberate partial results; group failure stops related work. Choose based on the outcome the caller needs rather than swapping APIs mechanically.

Source notes

References and review policy

Information checked on October 4, 2026. Section links identify sources for factual claims and technical explanations. Interpretations, practice scenarios and preparation recommendations are RecallDeck’s editorial work.

From reading to recall

Practice the full interview loop.

RecallDeck schedules the concepts you miss and keeps coding, design, and behavioral fundamentals available when the interviewer changes direction.

Start studying

Keep going

Languages105 min

Top 100 Go Interview Questions and Answers

Prepare with 100 Go interview questions and detailed answers on slices, interfaces, concurrency, runtime, testing, distributed systems, and reliability.

100 detailed answers