Control Flow & Human-in-the-Loop
routing, pausing, fanning out
Now the agent acts in the world, and some actions are risky enough to need a human's nod. You will route with
Command, pause mid-run for approval with interrupt(), fan work out in
parallel with Send, compose graphs as subgraphs, and make all of it survive failure.
Sahil's agent grows teeth
Sahil's research agent now drafts emails and can actually send them. That is useful and a little terrifying. An LLM that can hit "send" with no oversight will, eventually, email the wrong person something wrong. Sahil does not want to remove the power; he wants a checkpoint where a human approves before anything irreversible happens. He also wants the agent to research several sub-topics at once instead of one slow loop. Both needs are about control flow: who decides what runs, when execution pauses, and how work spreads out.
Two new demands. First: pause before risky actions and wait for a human, possibly minutes or hours later, without losing the run. Second: do many things in parallel and merge the results. Module 02's persistence makes the first possible; this module shows you the tools.
Command: update state and route in one move
In Module 01 a node returned an update, and a separate conditional edge decided where to go next. Sometimes you want a node
to do both at once: change state and say where to go. That is what Command is for.
A node returns a Command carrying an update (state change) and a
goto (next node). It folds the router into the node itself.
from langgraph.types import Command
from typing import Literal
def triage(state) -> Command[Literal["send", "revise"]]:
if looks_good(state):
return Command(update={"status": "ok"}, goto="send")
return Command(update={"status": "redo"}, goto="revise")
# No add_conditional_edges needed, the node routes itself.
# The Literal[...] hint tells LangGraph the possible destinations.
| When you want to... | Use |
|---|---|
| Decide the next node from a separate pure function | add_conditional_edges(node, router) |
| Decide the next node inside the working node, alongside a state update | return a Command(update=..., goto=...) |
Conditional edges and Command are two styles for the same job. Use conditional edges when routing
logic is cleanly separable; reach for Command when the decision and the work are tangled together, or
when (Module 04) you need to route and hand off to another agent in one step.
interrupt(): pause the graph for a human
This is the heart of the module. interrupt() stops the graph mid-node, saves everything (thanks to
the checkpointer), and hands control back to your application along with whatever you want to show the human. The run is not
blocked in memory waiting; it is persisted and paused. Hours later, you resume by sending the human's response
back in, and execution continues from exactly where it stopped.
from langgraph.types import interrupt, Command
def human_approval(state):
# execution PAUSES here. The value is surfaced to your app.
decision = interrupt({"draft": state["draft_email"], "ask": "Send this?"})
# ...resumes HERE with whatever the human sent back...
if decision == "approve":
return Command(goto="send_email")
return Command(goto="revise")
# 1) run until the interrupt, the graph stops cleanly
graph.invoke(inputs, config) # requires a checkpointer + thread_id
# 2) later, resume by passing a Command(resume=...) on the SAME thread
graph.invoke(Command(resume="approve"), config)
The pause is durable. The process can restart, the human can take an hour, and the run still continues from the saved checkpoint.
When the graph resumes after an interrupt(), the node containing it re-runs from the start,
up to the interrupt point. So any code before interrupt() in that node runs twice. Keep side
effects (sending, charging, writing) after the interrupt, never before it. This trips up nearly every beginner.
The four human-in-the-loop patterns
Almost every "human in the loop" requirement is one of these four shapes. They are all interrupt()
under the hood; the difference is what you ask for and what you do with the answer.
Approve / reject
Pause before a risky action, show it, continue only on approval. The email-send gate. Most common.
Edit state
Let the human correct the draft or fix a field, then resume with the edited value. Review-and-revise.
Review a tool call
Pause to inspect the arguments the model wants to pass to a tool, approve, tweak, or block them.
Ask for input
The agent is missing information; it pauses to ask the human a clarifying question, then resumes with the answer.
"How does human-in-the-loop work in LangGraph?" A node calls interrupt(payload), which durably
pauses the run (state is already checkpointed) and surfaces the payload to your application. You resume by invoking with
Command(resume=value) on the same thread; the node re-runs to the interrupt and continues with the
human's value. It needs a checkpointer and a thread id. The four flavours, approve, edit, review-tool-call, and ask-for-input, are all the same mechanism.
Send: dynamic parallel fan-out (map-reduce)
Sahil's other wish was speed: research five sub-topics at once, not one at a time in a loop. The problem is he does not know
how many sub-topics there will be until runtime. Send solves exactly this. From a node you can
return a list of Send objects, one per item, each kicking off the same worker node with its own slice of
input. The runtime runs them in parallel; a reducer on the result field merges what they each return. That is map-reduce.
from langgraph.types import Send
def assign_workers(state):
# MAP: one Send per subtopic → N parallel "research_one" runs
return [Send("research_one", {"topic": t}) for t in state["subtopics"]]
builder.add_conditional_edges("plan", assign_workers, ["research_one"])
# research_one returns {"findings": [one_result]}; the reducer (add) on
# "findings" REDUCES all parallel results into one list. Map, then reduce.
The number of workers is decided at runtime. The reducer on findings is what makes parallel writes safe.
Several workers write the findings field at the same time. Without a reducer the runtime cannot
combine them and raises an error. This is the moment Module 02's reducers stop being theory: parallelism requires them.
Subgraphs: a graph as a node
As agents grow, you want to reuse and isolate chunks of logic. A subgraph is a compiled graph used as a node inside a bigger graph. The research loop Sahil built can become a single "research" node in a larger workflow, with its own internal state, tested on its own, and dropped in wherever needed. It is the same modularity instinct that made you split a big function into small ones, applied to graphs.
research_graph = research_builder.compile() # a complete graph
parent = StateGraph(ParentState)
parent.add_node("research", research_graph) # ← used as a single node
parent.add_node("write_report", write_report)
parent.add_edge("research", "write_report")
Parent and subgraph communicate through shared state keys. If they share field names, state flows straight through. If their schemas differ, you wrap the subgraph in a small function node that translates the parent's state into the subgraph's inputs and back. Keep the seams explicit and subgraphs stay easy to test in isolation.
Retries and durability
Real agents call flaky things: network APIs, rate-limited models, tools that occasionally time out. Two features keep a long run from dying on a transient hiccup. A retry policy on a node automatically re-runs it on failure, with backoff. And because every step is checkpointed, a crash does not lose the whole task: on restart the graph resumes from the last good checkpoint instead of from zero. That property is called durable execution, and it is a big reason teams pick LangGraph for production.
from langgraph.types import RetryPolicy
builder.add_node(
"call_api", call_api,
retry_policy=RetryPolicy(max_attempts=3), # transient failures → retried
)
# Crash mid-run? Re-invoke on the same thread_id; the checkpointer resumes
# from the last completed step. No checkpointer = no durability.
Notice how one idea, the checkpointer, keeps paying off: it gave you memory (Module 02), it makes
interrupt() durable, and it makes crash-recovery automatic. When an interviewer asks "why is LangGraph
good for long-running agents?", the honest one-word answer is persistence, and these are its three dividends.
Hands-on: build an approval gate
Give Sahil's agent the email gate. Draft, pause for a human, then either send (just print, do not actually email) or revise. Run it twice: once approving, once rejecting, on two different threads.
from langgraph.graph import StateGraph, START, END
from langgraph.types import interrupt, Command
from langgraph.checkpoint.memory import InMemorySaver
def draft(state): return {"draft": "Hi, ..."}
def approve(state):
# TODO 1: call interrupt({"draft": state["draft"]}) → decision
# TODO 2: return Command(goto="send") if approved else Command(goto="revise")
...
def send(state): print("SENT:", state["draft"]); return {}
def revise(state): print("revising"); return {}
b = StateGraph(dict)
for n in (draft, approve, send, revise): b.add_node(n.__name__, n)
b.add_edge(START, "draft"); b.add_edge("draft", "approve")
b.add_edge("send", END); b.add_edge("revise", END)
graph = b.compile(checkpointer=InMemorySaver()) # TODO 3: why is this required?
cfg = {"configurable": {"thread_id": "a"}}
graph.invoke({}, cfg) # runs to the interrupt, pauses
graph.invoke(Command(resume="approve"), cfg) # resumes → SENT
- The first invoke stopped at the interrupt without finishing.
- Resuming with
"approve"printed SENT; resuming with anything else went to revise. - You can explain why removing the checkpointer breaks
interrupt(). - You kept all side effects after the interrupt and can say why.
- Bonus: you fanned out with
Sendand a reducer merged the parallel results.
Interview check
More in Module 05. = must-know cold.
Q1How does human-in-the-loop work, end to end?
interrupt(payload); the run durably pauses (state is checkpointed) and the payload is
surfaced to your app. You resume by invoking with Command(resume=value) on the same thread; the node
re-runs up to the interrupt and continues with the human's value. Requires a checkpointer and thread id.Q2What is Command and how does it differ from a conditional edge?
Command returned by a node carries both an update (state change) and a
goto (next node), letting the node route itself. A conditional edge keeps routing in a separate function.
Same outcome, different placement; Command also enables routing plus agent handoff in one step.Q3Why must side effects go after interrupt()?
Q4What is Send for?
Send(node, input), one per item, and
the runtime runs that node in parallel for each. A reducer on the result field merges the parallel outputs. The worker count is
decided at runtime.Q5What is a subgraph and how do parent and child share data?
Q6What makes LangGraph execution "durable"?
RetryPolicy re-runs transient failures. Together they let long runs survive failures, the same persistence
that powers memory and interrupts.