LangGraph Track  /  Module 04  /  Multi-Agent & Prod
Module 4 of 5 ~2.5h · build along
LangGraph · Module 04 · Multi-Agent & Production

Multi-Agent Systems & Production
from one agent to a team, shipped

Sahil's solo agent is doing too much. Here you give it tools, then split it into a team: a supervisor that delegates to specialists. Then the part most tutorials skip, taking it to production: seeing what it does (observability) and getting it running for real users (deployment).

~2.5h · build-along builds on Modules 01-03 tools · supervisor · handoffs LangSmith · deploy
0

One agent is doing too much

Sahil's single agent now has to plan, search, judge sources, and write, all in one prompt with one growing pile of tools. The prompt is bloated, it picks the wrong tool, and when something goes wrong he cannot tell which job failed. This is the classic moment teams reach for multiple agents: split the one overloaded brain into focused specialists, each with a small toolset and a clear job, coordinated by a graph. The reassuring part: a multi-agent system is still just a LangGraph graph. Each agent is a node (often a subgraph), and the coordination is edges. You already know the primitives.

Sahil's problem

One agent with ten tools and a paragraph of instructions is unreliable and unobservable. Sahil wants a researcher that only searches, a writer that only writes, and a supervisor that decides who works next, then he wants to ship it and watch it run.

1

Tools and the prebuilt ReAct agent

Before teams, the single most useful building block: an agent that can call tools. A tool is just a function the model is allowed to invoke, a web search, a calculator, a database lookup. The loop is "model thinks, maybe calls a tool, reads the result, thinks again", which is the ReAct pattern you built by hand in Module 01. LangGraph ships it prebuilt as create_react_agent, so you do not rewire that loop every time.

react.py · a tool-using agent in a few lines
from langgraph.prebuilt import create_react_agent

def web_search(query: str) -> str:
    """Search the web for a query and return the top result."""   # docstring = tool description the model reads
    ...

agent = create_react_agent(model=llm, tools=[web_search])
agent.invoke({"messages": [user("Who won the 2022 World Cup?")]})
# Under the hood: the same think → call tool → observe → think loop as Module 01,
# but compiled, batteries-included, and ready to drop into a bigger graph as a node.
It is still a graph

create_react_agent returns a compiled graph. It takes a checkpointer, streams, and interrupts just like everything you built by hand. The prebuilt is a convenience, not a different world. Knowing the manual version (Module 01) is what lets you debug the prebuilt when it misbehaves.

2

Why split into multiple agents

The honest answer to "should I use multiple agents?" is often "not yet". One well-prompted agent with a few tools beats a tangle of agents you cannot debug. But there is a real threshold where splitting pays off, and naming it keeps you out of trouble.

Split when

The toolset is too big for one prompt to choose well; jobs need different instructions or models; or you want each part testable and observable on its own.

Stay single when

A few tools and one coherent job. Multiple agents add coordination overhead, more tokens, more latency, and more places to fail. Do not pay that for nothing.

Interview gold

"When would you use a multi-agent architecture?" When a single agent's responsibilities or toolset grow large enough that prompt-following degrades, or when sub-tasks need genuinely different instructions, models, or independent testing. Otherwise prefer one agent: multi-agent buys modularity at the cost of coordination, latency, and tokens.

3

The supervisor pattern

The most common and most interview-relevant multi-agent shape. A supervisor agent sits at the centre. It looks at the state, decides which specialist should act next, routes to it, gets the result back, and decides again, until the job is done. The workers do not talk to each other; they all report to the supervisor. It is an org chart, and it maps cleanly onto a graph where the supervisor routes with Command (Module 03) and each worker is a node or subgraph.

researcher
supervisor
writer
delegates & collects, then
END

The supervisor is the only router. Each worker does its job and hands control back. Add a worker = add a node + a route.

supervisor.py · the supervisor routes; workers report back
from langgraph.types import Command
from typing import Literal

def supervisor(state) -> Command[Literal["researcher", "writer", "__end__"]]:
    nxt = pick_next(state)            # LLM decides: who works next, or done?
    return Command(goto=nxt)          # route, same Command from Module 03

def researcher(state):
    return Command(update={"findings": [...]}, goto="supervisor")  # report back

# Workers always goto "supervisor". The supervisor is the hub of the wheel.
# (langgraph also ships langgraph-supervisor to scaffold this for you.)
Prebuilt shortcut

You can hand-build the supervisor (best for understanding) or use the langgraph-supervisor library, which scaffolds the hub-and-workers wiring from a list of agents. Build it by hand once, then reach for the prebuilt.

4

Handoffs and the swarm pattern

The supervisor is hub-and-spoke. The other shape is the swarm: agents hand off directly to each other, peer to peer, no central boss. The researcher decides on its own to pass to the writer. The tool that makes this work is a handoff: a Command whose goto targets another agent, with graph=Command.PARENT so it can jump to a sibling in the parent graph rather than inside itself.

handoff.py · one agent passes control to another
def handoff_to_writer(state):
    return Command(
        goto="writer",
        graph=Command.PARENT,        # jump to a SIBLING in the parent graph
        update={"handoff_note": "research done, your turn"},
    )
# Supervisor = central router. Swarm = peers handing off directly.
# Both are just Command routing; the difference is who decides.
PatternWho decides routingGood when
SupervisorOne central agentYou want oversight, clear control, easy debugging. The default choice.
SwarmEach agent, peer to peerSpecialists with clear "next owner" handoffs; less central bottleneck.
Agent-as-toolThe calling agentA sub-agent is wrapped as a tool another agent can call like any function.
The unifying idea

Supervisor, swarm, and agent-as-tool are not three different technologies. They are three answers to one question: who decides what runs next? A central node, the peers themselves, or the caller. All three are expressed with the same Command routing and shared state you already learned. That framing is gold in an interview.

5

Observability: see what the agent actually did

An agent that works on your laptop and fails silently in production is worthless. You need to see every step: which node ran, what the model was prompted with, what each tool returned, how many tokens it cost, where it got slow. LangChain's tracing tool, LangSmith, captures all of that automatically. You usually turn it on with a couple of environment variables, no code change, and every run becomes a clickable trace of nodes, prompts, tool calls, latencies, and costs.

tracing · environment, not code
# set these and your existing graph is traced, no code change
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY=...      # from smith.langchain.com
# now every invoke/stream appears as a step-by-step trace you can inspect

Trace

Every node, prompt, and tool call for a single run, in order, with inputs and outputs.

Cost & latency

Tokens and time per step, so you can find the slow or expensive node and fix it.

Evals

Run datasets through the agent and score quality over time, so changes are measured, not guessed.

Do not skip this

Most "the agent is broken" problems are really "I cannot see what the agent did" problems. Tracing is the difference between debugging by guesswork and debugging by reading the actual run. Interviewers ask how you would debug a flaky agent; "I read the LangSmith trace to find which node or tool call went wrong" is the answer they want.

6

Deployment: from script to service

Your compiled graph runs fine in a script. Production means it runs as a service: an API many users hit concurrently, with persistence backed by a real database, background runs, and scaling. You have three broad options, and knowing the trade-off is enough for now.

OptionWhat it isTrade-off
LangGraph PlatformManaged hosting for LangGraph apps: API, persistence, scaling, cron, all provided.Fastest to production; you run less infrastructure yourself.
Self-hosted serverRun the LangGraph server in your own infra (containers) with your own Postgres.Full control; you own ops, scaling, and upgrades.
Embed the graphImport the compiled graph into your own FastAPI/Flask app and call it.Maximum flexibility; you rebuild the serving features yourself.

Whichever you pick, two things carry over from this whole track. First, the checkpointer becomes a real database (Postgres), so memory and durability survive restarts and scale across instances. Second, LangGraph Studio lets you visualise and step through your graph while you build, a visual debugger that draws the exact nodes and edges you wrote.

The thread you have been pulling

Notice the whole track converging: state and reducers (M2) define what persists, the checkpointer (M2) becomes production Postgres here, interrupts and durability (M3) keep long runs safe at scale, and multi-agent graphs (M4) are deployed as one service. You did not learn four disconnected topics; you built one thing, layer by layer.

7

Hands-on: a two-agent team

Build the smallest real team: a supervisor routing between a researcher (one search tool) and a writer. Trace it in LangSmith and read the run. This is the capstone of the build-along track.

team.py · fill the TODOs
from langgraph.graph import StateGraph, START, MessagesState
from langgraph.types import Command
from typing import Literal

def supervisor(state) -> Command[Literal["researcher", "writer", "__end__"]]:
    # TODO 1: decide next worker from state; return Command(goto=...)
    ...
def researcher(state):
    # TODO 2: do a (fake) search, return Command(update={...}, goto="supervisor")
    ...
def writer(state):
    # TODO 3: write final answer, return Command(update={...}, goto="supervisor")
    ...

b = StateGraph(MessagesState)
for n in (supervisor, researcher, writer): b.add_node(n.__name__, n)
b.add_edge(START, "supervisor")
graph = b.compile()
# Run it, then set LANGSMITH_TRACING=true and read the trace.
  • The supervisor routed to a worker, got control back, and eventually ended.
  • Each worker returned a Command that updated state and went back to the supervisor.
  • You built a create_react_agent with one real tool and watched it call the tool.
  • You enabled LangSmith and read a full trace of nodes, prompts, and tool calls.
  • You can explain supervisor vs swarm as "who decides routing" in one sentence.
8

Interview check

The full bank is Module 05. = must-know cold.

Q1When should you use multiple agents instead of one?
When one agent's toolset or responsibilities grow large enough that prompt-following degrades, or when sub-tasks need different instructions, models, or independent testing. Otherwise prefer a single agent: multi-agent adds coordination overhead, latency, and token cost.
Q2Explain the supervisor pattern.
A central supervisor agent inspects state and routes to specialist workers one at a time; workers do their job and report back to the supervisor, which decides again until done. Workers do not talk to each other. In LangGraph the supervisor routes with Command(goto=...) and workers are nodes or subgraphs.
Q3Supervisor vs swarm vs agent-as-tool?
All answer "who decides routing." Supervisor: one central agent. Swarm: peers hand off directly to each other (a Command with graph=Command.PARENT). Agent-as-tool: a sub-agent is wrapped as a tool the caller invokes. Same Command routing underneath.
Q4What is create_react_agent and what is it under the hood?
A prebuilt that returns a compiled tool-calling agent running the think to call-tool to observe loop. Under the hood it is an ordinary LangGraph graph, so it supports checkpointers, streaming, and interrupts, and drops into a bigger graph as a node.
Q5How would you debug a flaky agent in production?
Read the LangSmith trace for a failing run: inspect each node, the exact prompt, every tool's input and output, plus token and latency per step, to find which node or tool call went wrong. Pair that with checkpoint history (time travel) to see the state at each step.
Q6What changes when you move a LangGraph app to production?
The in-memory checkpointer becomes a database-backed one (Postgres) for durable, shared memory; the graph runs as a service (LangGraph Platform, self-hosted server, or embedded in your API) handling concurrency and scaling; and you add tracing (LangSmith) and visual debugging (Studio).