State, Memory & Persistence
so the agent stops forgetting
Module 01's agent has amnesia: every invoke starts from nothing. Here you give it a
real memory. You will control exactly how state updates merge (reducers), where state is saved
(checkpointers), how to keep many conversations apart (threads), and how to watch a run unfold (streaming) or rewind it (time travel).
Where Sahil is stuck
Sahil's research agent works for a single question. But he wants a conversation: ask something, get an answer, then ask a
follow-up that refers back to the first. He tries it, and the agent stares blankly, because the second
invoke knows nothing about the first. The state from run one was returned to him and then thrown away.
State lives only for the duration of one invoke. Close the call and it is gone. To hold a
conversation, or to survive a server restart mid-task, the graph needs somewhere durable to write its state between steps.
That mechanism is a checkpointer, and getting there cleanly means first understanding how state updates even combine.
So this module climbs in a deliberate order: first how a node's return value merges into state (reducers), then where that state is saved (checkpointers), then how to partition saved state per user (threads), and finally the two payoffs that persistence unlocks: streaming and time travel.
Reducers: how updates merge
Back in Module 01 a node returned {"answer": text} and that simply overwrote the old answer.
Overwrite is the default reducer: last write wins. But overwrite is wrong for anything that should accumulate.
If your findings field overwrote on every research step, the agent would only ever remember its
most recent lookup. You want each step to add to the list, not replace it.
A reducer is the rule that combines the old value of a field with the update a node returns. You attach one
to a field by annotating its type. The most common reducer is "append to a list", which the standard library gives you as
operator.add (or LangGraph's add_messages for chat).
from typing import Annotated
from operator import add
from typing_extensions import TypedDict
class ResearchState(TypedDict):
question: str
findings: Annotated[list, add] # reducer = add → updates APPEND
answer: str # no reducer → updates OVERWRITE
# A node returning {"findings": ["fact A"]} now ADDS to the list.
# Two parallel nodes can each contribute; their lists are merged, not lost.
Read Annotated[list, add] as "this field is a list, and when a node returns more of it, glue the
new on with add." No annotation means "just overwrite". This is also why parallel nodes do not clobber
each other: each returns its slice, and the reducer merges all slices. Without a reducer, two parallel writers to the same
field is an error, because the runtime would not know how to combine them.
This is exactly the machinery behind MessagesState from Module 01. Its single field is declared as
Annotated[list, add_messages], and add_messages is a smarter append that also
handles updating a message by id. You were already using a reducer; now you can write your own.
Input and output schemas (optional polish)
By default the state schema is also what callers pass in and what they get back. Often you want the internal state to carry scratch fields that you do not want to expose. LangGraph lets you declare a separate input schema and output schema, so the public surface stays clean while the graph keeps private working fields.
class InputState(TypedDict): question: str
class OutputState(TypedDict): answer: str
class InternalState(TypedDict): question: str; scratch: list; answer: str
builder = StateGraph(InternalState, input_schema=InputState, output_schema=OutputState)
# callers send only {question}; they receive only {answer}; scratch stays hidden
Skip this until a graph grows enough internal scratch state that you want a tidy public contract. It is a refinement, not a requirement. Know that it exists so you recognise it in real codebases and interviews.
Checkpointers: where memory actually lives
Here is the line that solves Sahil's amnesia. A checkpointer is a backend that saves a snapshot of the
graph's state after every step. Attach one at compile time and the graph automatically persists its progress. The
simplest is InMemorySaver (state held in RAM, good for learning and tests); for production you swap in
a database-backed saver such as the SQLite or Postgres checkpointers, with no change to your graph logic.
from langgraph.checkpoint.memory import InMemorySaver
checkpointer = InMemorySaver() # swap for SqliteSaver/PostgresSaver in prod
graph = builder.compile(checkpointer=checkpointer) # ← the whole trick
That single argument changes the agent's nature. Now, between steps, the runtime writes the current state to the checkpointer. If the process dies after step three of a five-step task, restarting resumes from the saved snapshot instead of starting over. This durability is the foundation that human-in-the-loop (Module 03) is built on: pausing for a human is just "stop, the state is already saved, resume when they reply."
The moment you compile with a checkpointer, every call must include a thread id (next section). Without it the runtime does not know which saved conversation to read and write, and it will error. Checkpointer and thread id always travel together.
Threads: keeping conversations apart
One agent serves many users. Sahil's research bot should not mix Priya's conversation into Ravi's. A thread
is a named, independent timeline of state. You pick a thread by passing a thread_id in the config on
every call. Same thread id, same continuing memory. New thread id, a fresh blank conversation.
config = {"configurable": {"thread_id": "priya-001"}}
graph.invoke({"messages": [user("What is LangGraph?")]}, config)
# ... later, SAME thread id → the agent remembers the first turn ...
graph.invoke({"messages": [user("And how is it different from LangChain?")]}, config)
# a DIFFERENT thread_id would start clean, with no memory of the above
"How does a LangGraph agent remember a conversation?" Compile with a checkpointer (the storage), then pass a thread_id in config (the conversation key). After each step the runtime saves a checkpoint under that thread; the next call with the same thread id loads it and continues. Different users get different thread ids, so their histories never collide.
Short-term vs long-term memory
The checkpointer you just learned is short-term memory: it remembers this conversation, scoped to a
thread. But sometimes you want facts that outlive any single conversation: a user's name, their preferences, things learned in a
thread last week that should apply today. That is long-term memory, and LangGraph models it separately with a
Store, keyed by a namespace you choose (often the user id) rather than by thread.
Short-term (thread-scoped)
The running conversation: message history, scratch findings. Lives in the checkpointer, keyed by thread_id. Ends when the thread does.
Long-term (cross-thread)
Durable facts about a user or the world: preferences, profile, prior learnings. Lives in a Store, keyed by a namespace. Survives across conversations.
from langgraph.store.memory import InMemoryStore
store = InMemoryStore()
graph = builder.compile(checkpointer=checkpointer, store=store)
# inside a node you receive the store and read/write by (namespace, key):
def remember(state, *, store):
ns = ("user", "priya") # namespace, not thread
store.put(ns, "prefers", {"format": "bullets"})
saved = store.get(ns, "prefers") # available in ANY future thread
...
Checkpointer = "what happened in this chat". Store = "what I know about this user, always". Many real agents use both: the thread remembers the current exchange, the store remembers the person. Interviewers love this distinction.
Streaming: watch the run unfold
Calling invoke waits for the whole graph to finish, then hands you one final result. For a
multi-step agent that can take many seconds, and a blank screen feels broken. stream instead yields
output as it happens, so you can show progress or token-by-token text. You choose what granularity you want with a stream mode.
| Stream mode | What you get | Use it for |
|---|---|---|
"updates" | The state change after each node runs. | Showing "now searching... now writing..." progress. |
"values" | The full state after each step. | Debugging, or rendering the whole evolving state. |
"messages" | LLM tokens as they generate, per node. | The typewriter effect in a chat UI. |
for chunk in graph.stream(inputs, config, stream_mode="updates"):
# chunk = {node_name: {field: new_value}}, one per completed node
print(chunk)
# Same graph, same state, same memory, just a different way to consume output.
Take your looping agent from Module 01, compile it with an InMemorySaver, and run it with
stream in "updates" mode. Watch each loop iteration print as its own chunk.
That visible heartbeat is what makes long-running agents feel alive instead of frozen.
Time travel: rewind and replay
Because the checkpointer saves a snapshot after every step, a thread is not just a "current state", it is a full history of states. That history is a superpower. You can list every checkpoint, inspect what the state looked like at step three, and even resume execution from an earlier checkpoint, optionally after editing it. This is how you debug "why did the agent decide that?" and how you let a user undo a bad turn.
# current snapshot for this thread
snap = graph.get_state(config)
print(snap.values, snap.next) # state + which node would run next
# every snapshot, newest first
for s in graph.get_state_history(config):
print(s.config["configurable"]["checkpoint_id"])
# resume from an OLD checkpoint: pass its config back into the graph
graph.invoke(None, an_old_checkpoint_config) # re-runs forward from there
Persistence is not only about not losing work. It turns every run into an auditable, editable timeline. Time travel, human-in-the-loop, and "fork the conversation and try a different branch" are all the same underlying feature: saved checkpoints you can read, edit, and resume from. Hold that and Module 03 will feel inevitable rather than new.
Hands-on: give Sahil's agent a memory
Upgrade your Module 01 graph into something that holds a conversation. Do it for real, two turns on one thread, then a third turn on a new thread to prove isolation.
from langgraph.graph import StateGraph, START, MessagesState
from langgraph.checkpoint.memory import InMemorySaver
def chat(state: MessagesState):
# TODO 1: call llm on state["messages"], return {"messages": [reply]}
...
builder = StateGraph(MessagesState)
builder.add_node("chat", chat)
builder.add_edge(START, "chat")
# TODO 2: compile with an InMemorySaver()
graph = builder.compile(...)
cfg = {"configurable": {"thread_id": "t1"}}
graph.invoke({"messages": [user("My name is Sahil.")]}, cfg)
out = graph.invoke({"messages": [user("What is my name?")]}, cfg)
# TODO 3: confirm it answers "Sahil". Then repeat with thread_id "t2"
# and confirm it does NOT know the name.
- Turn two on the same thread correctly recalled a fact from turn one.
- A new
thread_idstarted blank, proving conversations are isolated. - You added a reducer (
Annotated[list, add]) to a custom field and saw it accumulate. - You called
get_state_historyand saw more than one checkpoint for a thread. - You can explain checkpointer vs Store (this-chat vs this-user) without peeking.
Interview check
More in Module 05. = must-know cold.
Q1What is a reducer and why is it needed?
Annotated[list, add] appends. It is needed so fields can accumulate (like findings or messages) and so
parallel nodes writing the same field can be combined instead of conflicting.Q2How do you make a LangGraph agent remember across calls?
Q3Short-term vs long-term memory in LangGraph?
Q4Why must you pass a thread_id once a checkpointer is attached?
Q5When would you stream instead of invoke, and what are the modes?
"updates"
yields each node's change, "values" yields the full state per step, "messages"
yields LLM tokens as they generate.Q6What is "time travel" and what makes it possible?
get_state_history and resume from one (optionally edited). It is the same persistence
machinery that powers human-in-the-loop and conversation forking.