LangGraph Track  /  Module 03  /  Control & Humans
Module 3 of 5 ~2h · build along
LangGraph · Module 03 · Control & Humans

Control Flow & Human-in-the-Loop
routing, pausing, fanning out

Now the agent acts in the world, and some actions are risky enough to need a human's nod. You will route with Command, pause mid-run for approval with interrupt(), fan work out in parallel with Send, compose graphs as subgraphs, and make all of it survive failure.

~2h · build-along builds on Module 02 Command · interrupt · Send subgraphs · durability
0

Sahil's agent grows teeth

Sahil's research agent now drafts emails and can actually send them. That is useful and a little terrifying. An LLM that can hit "send" with no oversight will, eventually, email the wrong person something wrong. Sahil does not want to remove the power; he wants a checkpoint where a human approves before anything irreversible happens. He also wants the agent to research several sub-topics at once instead of one slow loop. Both needs are about control flow: who decides what runs, when execution pauses, and how work spreads out.

Sahil's problem

Two new demands. First: pause before risky actions and wait for a human, possibly minutes or hours later, without losing the run. Second: do many things in parallel and merge the results. Module 02's persistence makes the first possible; this module shows you the tools.

1

Command: update state and route in one move

In Module 01 a node returned an update, and a separate conditional edge decided where to go next. Sometimes you want a node to do both at once: change state and say where to go. That is what Command is for. A node returns a Command carrying an update (state change) and a goto (next node). It folds the router into the node itself.

command.py · one return value does the update AND the routing
from langgraph.types import Command
from typing import Literal

def triage(state) -> Command[Literal["send", "revise"]]:
    if looks_good(state):
        return Command(update={"status": "ok"}, goto="send")
    return Command(update={"status": "redo"}, goto="revise")
# No add_conditional_edges needed, the node routes itself.
# The Literal[...] hint tells LangGraph the possible destinations.
When you want to...Use
Decide the next node from a separate pure functionadd_conditional_edges(node, router)
Decide the next node inside the working node, alongside a state updatereturn a Command(update=..., goto=...)
Both are fine

Conditional edges and Command are two styles for the same job. Use conditional edges when routing logic is cleanly separable; reach for Command when the decision and the work are tangled together, or when (Module 04) you need to route and hand off to another agent in one step.

2

interrupt(): pause the graph for a human

This is the heart of the module. interrupt() stops the graph mid-node, saves everything (thanks to the checkpointer), and hands control back to your application along with whatever you want to show the human. The run is not blocked in memory waiting; it is persisted and paused. Hours later, you resume by sending the human's response back in, and execution continues from exactly where it stopped.

approval.py · stop, ask, resume
from langgraph.types import interrupt, Command

def human_approval(state):
    # execution PAUSES here. The value is surfaced to your app.
    decision = interrupt({"draft": state["draft_email"], "ask": "Send this?"})
    # ...resumes HERE with whatever the human sent back...
    if decision == "approve":
        return Command(goto="send_email")
    return Command(goto="revise")

# 1) run until the interrupt, the graph stops cleanly
graph.invoke(inputs, config)        # requires a checkpointer + thread_id
# 2) later, resume by passing a Command(resume=...) on the SAME thread
graph.invoke(Command(resume="approve"), config)
draft
human_approval
pause
your app
human replies
resume
send_email
done

The pause is durable. The process can restart, the human can take an hour, and the run still continues from the saved checkpoint.

The gotcha that bites everyone

When the graph resumes after an interrupt(), the node containing it re-runs from the start, up to the interrupt point. So any code before interrupt() in that node runs twice. Keep side effects (sending, charging, writing) after the interrupt, never before it. This trips up nearly every beginner.

3

The four human-in-the-loop patterns

Almost every "human in the loop" requirement is one of these four shapes. They are all interrupt() under the hood; the difference is what you ask for and what you do with the answer.

Approve / reject

Pause before a risky action, show it, continue only on approval. The email-send gate. Most common.

Edit state

Let the human correct the draft or fix a field, then resume with the edited value. Review-and-revise.

Review a tool call

Pause to inspect the arguments the model wants to pass to a tool, approve, tweak, or block them.

Ask for input

The agent is missing information; it pauses to ask the human a clarifying question, then resumes with the answer.

Interview gold

"How does human-in-the-loop work in LangGraph?" A node calls interrupt(payload), which durably pauses the run (state is already checkpointed) and surfaces the payload to your application. You resume by invoking with Command(resume=value) on the same thread; the node re-runs to the interrupt and continues with the human's value. It needs a checkpointer and a thread id. The four flavours, approve, edit, review-tool-call, and ask-for-input, are all the same mechanism.

4

Send: dynamic parallel fan-out (map-reduce)

Sahil's other wish was speed: research five sub-topics at once, not one at a time in a loop. The problem is he does not know how many sub-topics there will be until runtime. Send solves exactly this. From a node you can return a list of Send objects, one per item, each kicking off the same worker node with its own slice of input. The runtime runs them in parallel; a reducer on the result field merges what they each return. That is map-reduce.

fanout.py · spawn one worker per item, merge results
from langgraph.types import Send

def assign_workers(state):
    # MAP: one Send per subtopic → N parallel "research_one" runs
    return [Send("research_one", {"topic": t}) for t in state["subtopics"]]

builder.add_conditional_edges("plan", assign_workers, ["research_one"])

# research_one returns {"findings": [one_result]}; the reducer (add) on
# "findings" REDUCES all parallel results into one list. Map, then reduce.
plan
Send xN
research_one
research_one
research_one
reduce
synthesize

The number of workers is decided at runtime. The reducer on findings is what makes parallel writes safe.

Why the reducer is mandatory here

Several workers write the findings field at the same time. Without a reducer the runtime cannot combine them and raises an error. This is the moment Module 02's reducers stop being theory: parallelism requires them.

5

Subgraphs: a graph as a node

As agents grow, you want to reuse and isolate chunks of logic. A subgraph is a compiled graph used as a node inside a bigger graph. The research loop Sahil built can become a single "research" node in a larger workflow, with its own internal state, tested on its own, and dropped in wherever needed. It is the same modularity instinct that made you split a big function into small ones, applied to graphs.

subgraph.py · compile one graph, use it as a node in another
research_graph = research_builder.compile()      # a complete graph

parent = StateGraph(ParentState)
parent.add_node("research", research_graph)        # ← used as a single node
parent.add_node("write_report", write_report)
parent.add_edge("research", "write_report")
The one rule to remember

Parent and subgraph communicate through shared state keys. If they share field names, state flows straight through. If their schemas differ, you wrap the subgraph in a small function node that translates the parent's state into the subgraph's inputs and back. Keep the seams explicit and subgraphs stay easy to test in isolation.

6

Retries and durability

Real agents call flaky things: network APIs, rate-limited models, tools that occasionally time out. Two features keep a long run from dying on a transient hiccup. A retry policy on a node automatically re-runs it on failure, with backoff. And because every step is checkpointed, a crash does not lose the whole task: on restart the graph resumes from the last good checkpoint instead of from zero. That property is called durable execution, and it is a big reason teams pick LangGraph for production.

durable.py · auto-retry flaky steps; resume after crashes
from langgraph.types import RetryPolicy

builder.add_node(
    "call_api", call_api,
    retry_policy=RetryPolicy(max_attempts=3),   # transient failures → retried
)
# Crash mid-run? Re-invoke on the same thread_id; the checkpointer resumes
# from the last completed step. No checkpointer = no durability.
Tie it together

Notice how one idea, the checkpointer, keeps paying off: it gave you memory (Module 02), it makes interrupt() durable, and it makes crash-recovery automatic. When an interviewer asks "why is LangGraph good for long-running agents?", the honest one-word answer is persistence, and these are its three dividends.

7

Hands-on: build an approval gate

Give Sahil's agent the email gate. Draft, pause for a human, then either send (just print, do not actually email) or revise. Run it twice: once approving, once rejecting, on two different threads.

exercise.py · fill the TODOs
from langgraph.graph import StateGraph, START, END
from langgraph.types import interrupt, Command
from langgraph.checkpoint.memory import InMemorySaver

def draft(state):       return {"draft": "Hi, ..."}

def approve(state):
    # TODO 1: call interrupt({"draft": state["draft"]}) → decision
    # TODO 2: return Command(goto="send") if approved else Command(goto="revise")
    ...

def send(state):     print("SENT:", state["draft"]); return {}
def revise(state):   print("revising"); return {}

b = StateGraph(dict)
for n in (draft, approve, send, revise): b.add_node(n.__name__, n)
b.add_edge(START, "draft"); b.add_edge("draft", "approve")
b.add_edge("send", END); b.add_edge("revise", END)
graph = b.compile(checkpointer=InMemorySaver())   # TODO 3: why is this required?

cfg = {"configurable": {"thread_id": "a"}}
graph.invoke({}, cfg)                       # runs to the interrupt, pauses
graph.invoke(Command(resume="approve"), cfg)  # resumes → SENT
  • The first invoke stopped at the interrupt without finishing.
  • Resuming with "approve" printed SENT; resuming with anything else went to revise.
  • You can explain why removing the checkpointer breaks interrupt().
  • You kept all side effects after the interrupt and can say why.
  • Bonus: you fanned out with Send and a reducer merged the parallel results.
8

Interview check

More in Module 05. = must-know cold.

Q1How does human-in-the-loop work, end to end?
A node calls interrupt(payload); the run durably pauses (state is checkpointed) and the payload is surfaced to your app. You resume by invoking with Command(resume=value) on the same thread; the node re-runs up to the interrupt and continues with the human's value. Requires a checkpointer and thread id.
Q2What is Command and how does it differ from a conditional edge?
A Command returned by a node carries both an update (state change) and a goto (next node), letting the node route itself. A conditional edge keeps routing in a separate function. Same outcome, different placement; Command also enables routing plus agent handoff in one step.
Q3Why must side effects go after interrupt()?
On resume, the node re-executes from the top to the interrupt point, so anything before the interrupt runs again. Put sends, charges, and writes after the interrupt so they happen exactly once.
Q4What is Send for?
Dynamic parallel fan-out (map-reduce). A node returns a list of Send(node, input), one per item, and the runtime runs that node in parallel for each. A reducer on the result field merges the parallel outputs. The worker count is decided at runtime.
Q5What is a subgraph and how do parent and child share data?
A compiled graph used as a node in a larger graph, for reuse and isolation. They communicate through shared state keys; if schemas differ, wrap the subgraph in a function node that translates state in and out.
Q6What makes LangGraph execution "durable"?
Every step is checkpointed, so a crash resumes from the last good step rather than restarting; node-level RetryPolicy re-runs transient failures. Together they let long runs survive failures, the same persistence that powers memory and interrupts.