AI
LangGraph vs CrewAI in 2026: Which for Production Agents?
October 202612 min read

LangGraph is a low-level orchestration runtime where you define an explicit graph of nodes and edges, and a checkpointer saves state after each step. CrewAI starts from role-based agents grouped into crews, and adds Flows with start, listen and router decorators when you need a fixed sequence. The practical difference is how much of the control flow you write yourself.
LangGraph is the stronger default when runs must pause for human approval, survive restarts and resume from a checkpoint, because those are its core API. CrewAI is a good fit for collaborative, open-ended agent work and gets a first version running faster. CrewAI Flows can handle production workflows too, but you must choose the persistence backend and a non-blocking feedback provider yourself.
A node calls interrupt() with a JSON-serialisable payload, and the graph saves a checkpoint and returns the payload to the caller under __interrupt__. You resume later by invoking the graph with Command(resume=value) and the same thread_id. The interrupted node re-runs from its first line, so code before interrupt() must be idempotent.
Yes. The human_feedback decorator on a Flow method asks for feedback and can emit outcomes such as approved or rejected, optionally using an LLM to classify free-text replies. By default it blocks on console input, so for a web app you pass a provider that raises HumanFeedbackPending, which persists the state, and later call resume() or resume_async().
Both can, if persistence is configured. LangGraph uses a checkpointer such as PostgresSaver keyed by thread_id, with durability modes exit, async and sync controlling when checkpoints are written. CrewAI uses the persist decorator, backed by SQLite by default, and resumes when you kick off with the saved state id.

Key Takeaway
LangGraph suits production agents that must pause for approval, survive restarts and resume from a checkpoint, because state, edges and persistence are explicit. CrewAI suits role-based agent teams and gets a first version running faster, and its Flows add persistence and human feedback, but its default feedback blocks on console input until you supply a provider.
Picture an accounts-payable agent for an Indonesian distributor. It reads a supplier invoice PDF, compares the total against the purchase order in the ERP, and posts a draft invoice. Most invoices match. The ones that are more than two percent off need a person to approve them, and that person may answer four hours later from a phone, after the server has been redeployed twice.
That last sentence is where the LangGraph vs CrewAI choice is actually made. This post builds the same three-step workflow in both frameworks, using only APIs from their current documentation: LangGraph 1.x, whose 1.0 release in October 2025 was its first stable major version, and CrewAI, whose documentation currently resolves to version 1.15.23. The comparison covers control model, durable execution, human approval, observability, managed platforms and learning curve.
Both are open-source Python frameworks and both can run this workflow. They differ in what they make explicit. LangGraph makes you draw the graph and gives you a checkpoint after every step. CrewAI makes you describe agents by role and goal, and adds Flows when you need the steps pinned down.
| Dimension | LangGraph | CrewAI |
|---|---|---|
| Control model | A StateGraph of nodes and edges. Every transition is code you wrote, including conditional routing through Command(goto=...). | Crews of role-based agents run sequentially or under a manager agent, orchestrated by event-driven Flows using start, listen and router decorators. |
| State | A typed schema, usually TypedDict, merged from each node's return value and written by a checkpointer. | A Pydantic model or a dict on the Flow, with an id field added automatically. |
| Durable execution | Checkpointers for Postgres, SQLite, MongoDB and Cosmos DB, plus durability modes exit, async and sync. | The persist decorator, backed by SQLiteFlowPersistence by default, at class or method level. |
| Resume | Invoke again with the same thread_id. The interrupted node re-runs from its first line. | Kick off with the saved state id to continue, or restore_from_state_id to fork a new run. |
| Human-in-the-loop | interrupt() inside a node returns a payload; Command(resume=...) supplies the answer. | The human_feedback decorator; console input by default, a provider for non-blocking use, then resume(). |
| History and replay | get_state_history lists every checkpoint; resuming from a checkpoint_id replays from that point. | Forking a new run from a saved state id; the Flows docs describe no per-step checkpoint history. |
| Observability and platform | LangSmith tracing; LangSmith Deployment runs graphs as a managed service. | Built-in tracing to CrewAI AMP, flow.plot() for an HTML diagram; AMP hosts crews and flows. |
| Learning curve | Steeper. You reason about state shape, edges and replay before the first useful run. | Gentler. Role, goal and backstory read like a job description, so a first crew is short. |
If the workflow has a fixed shape, money at stake and a person in the middle, LangGraph's explicitness pays for itself. If the work is open-ended research or drafting where agents should decide among themselves, CrewAI's crews are the more natural fit, and Flows let you add structure around them without switching frameworks.
LangGraph describes itself as a low-level orchestration runtime, and the word low-level is accurate. A node is a function that takes state and returns an update. An edge is a line you add. When the model should choose, you still decide what it chooses between, by returning a Command that names the next node. Nothing moves that you did not wire, which is what makes the graph auditable.
CrewAI starts from the other end. A crew is a set of agents with a role, a goal and a backstory, working through tasks either one after another or under a manager agent that delegates and checks results. Flows were added for the cases where autonomy is the wrong default: they are plain Python methods chained with start, listen and router, and a Flow method can kick off a whole crew. For an invoice pipeline you end up writing mostly Flow code with a single agent inside it, which is a sign the problem is graph-shaped.
Three steps: extract the fields with a structured-output model call, match the total against the purchase order, then either post a draft or stop for approval. The tolerance and the ERP helper functions are illustrative; the LangGraph calls are the documented ones.
from typing import Literal, TypedDict
from langchain.chat_models import init_chat_model
from langgraph.checkpoint.postgres import PostgresSaver
from langgraph.graph import END, START, StateGraph
from langgraph.types import Command, interrupt
from pydantic import BaseModel
class InvoiceFields(BaseModel):
vendor_code: str
po_number: str
total_idr: int
class InvoiceState(TypedDict, total=False):
raw_text: str
fields: InvoiceFields
variance_pct: float
erp_doc_id: str
llm = init_chat_model("openai:gpt-4.1-mini").with_structured_output(InvoiceFields)
TOLERANCE_PCT = 2.0 # example policy: above this, a person signs off
def extract(state: InvoiceState) -> InvoiceState:
return {"fields": llm.invoke(state["raw_text"])}
def match_po(state: InvoiceState) -> Command[Literal["approve", "post"]]:
po_total = erp.get_po_total(state["fields"].po_number) # read-only
variance = abs(state["fields"].total_idr - po_total) / po_total * 100
nxt = "approve" if variance > TOLERANCE_PCT else "post"
return Command(update={"variance_pct": variance}, goto=nxt)
def approve(state: InvoiceState) -> Command[Literal["post", "__end__"]]:
# Nothing with a side effect above this line: on resume the whole
# node runs again from the top, not from the interrupt() call.
decision = interrupt({
"po": state["fields"].po_number,
"variance_pct": round(state["variance_pct"], 2),
})
return Command(goto="post" if decision == "approve" else END)
def post(state: InvoiceState) -> InvoiceState:
# Idempotency key = PO number: a replayed node finds the existing draft.
doc_id = erp.create_draft_invoice(state["fields"], key=state["fields"].po_number)
return {"erp_doc_id": doc_id}
builder = StateGraph(InvoiceState)
builder.add_node(extract)
builder.add_node(match_po)
builder.add_node(approve)
builder.add_node(post)
builder.add_edge(START, "extract")
builder.add_edge("extract", "match_po")
builder.add_edge("post", END)
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
checkpointer.setup() # creates the checkpoint tables once
graph = builder.compile(checkpointer=checkpointer)
# thread_id is the checkpoint primary key: make it the business key.
config = {"configurable": {"thread_id": "inv-2026-10-0042"}}
result = graph.invoke({"raw_text": pdf_text}, config, durability="sync")
if "__interrupt__" in result:
notify_approver(result["__interrupt__"]) # hours may pass here
# Later, from the approval webhook, in any process:
graph.invoke(Command(resume="approve"), config)Two details carry most of the production weight. The thread_id is the checkpointer's primary key, so using the invoice number makes every run findable by the people who ask about it. And the resume call does not need the original process: any worker with the same checkpointer and thread_id picks up where the graph paused, which is exactly what the four-hours-later approval needs.
The Flow version maps step for step. The extractor is a single Agent with a Pydantic response format, the match is a router that returns a label, and approval is the human_feedback decorator, which emits one of the listed outcomes. The llm argument lets the reviewer type free text and has a model classify it into approved or rejected.
from crewai.agent import Agent
from crewai.flow.flow import Flow, listen, router, start
from crewai.flow.human_feedback import HumanFeedbackResult, human_feedback
from crewai.flow.persistence import persist
from pydantic import BaseModel
class InvoiceFields(BaseModel):
vendor_code: str
po_number: str
total_idr: int
class InvoiceState(BaseModel):
raw_text: str = ""
fields: InvoiceFields | None = None
variance_pct: float = 0.0
erp_doc_id: str = ""
TOLERANCE_PCT = 2.0
@persist # SQLiteFlowPersistence by default: state saved per method
class InvoiceFlow(Flow[InvoiceState]):
@start()
async def extract(self):
clerk = Agent(
role="AP clerk",
goal="Read supplier invoices into structured fields",
backstory="Indonesian supplier invoices, amounts in rupiah.",
)
result = await clerk.kickoff_async(self.state.raw_text, response_format=InvoiceFields)
self.state.fields = result.pydantic
@router(extract)
def match_po(self):
po_total = erp.get_po_total(self.state.fields.po_number)
self.state.variance_pct = abs(self.state.fields.total_idr - po_total) / po_total * 100
return "needs_review" if self.state.variance_pct > TOLERANCE_PCT else "auto_post"
@listen("needs_review")
@human_feedback(
message="Variance above tolerance. Approve posting this invoice?",
emit=["approved", "rejected"],
llm="gpt-4.1-mini", # maps free-text feedback onto an outcome
default_outcome="rejected", # silence never posts an invoice
)
def request_review(self):
return f"PO {self.state.fields.po_number}: {self.state.variance_pct:.2f}% over"
@listen("auto_post")
def post_clean(self):
self._post()
@listen("approved")
def post_approved(self, result: HumanFeedbackResult):
self._post()
@listen("rejected")
def park(self, result: HumanFeedbackResult):
erp.flag_for_ap_team(self.state.fields.po_number, result.feedback)
def _post(self):
self.state.erp_doc_id = erp.create_draft_invoice(
self.state.fields, key=self.state.fields.po_number
)
flow = InvoiceFlow()
flow.kickoff(inputs={"raw_text": pdf_text})
# Resume the same run later: InvoiceFlow().kickoff(inputs={"id": saved_state_id})
# Fork a new run from it: InvoiceFlow().kickoff(restore_from_state_id=saved_state_id)It reads more like ordinary Python than the graph does, and for a team new to agents that matters. The cost is that two things LangGraph gives you for free become your responsibility: choosing a persistence backend beyond the default SQLite file, and making sure the approval step does not block a worker, which the next sections cover.
Both frameworks persist state, but they restart at different granularities, and the difference decides where you put side effects.
For a back-office workflow the practical rule is the same in both: anything that writes to the ERP must be idempotent, keyed on a business identifier such as the PO or invoice number, so that a replayed step finds the draft it already created instead of creating a second one. Choose sync durability in LangGraph when a lost step is more expensive than a slower one.
The classic double-post: an ERP write placed before interrupt() in the same node. The first run creates the draft and pauses; the approval resumes the node from the top and creates it again. Keep the write in a separate node after approval, or make it idempotent, and never wrap interrupt() in a bare try/except, because it works by raising internally.
The approval step is where the two documentation sets differ most in their defaults. CrewAI's human_feedback blocks on console input unless you pass a provider. LangGraph's interrupt never blocks: it saves the checkpoint and returns the payload to the caller, so the pause and the answer are naturally two separate requests.
# CrewAI: the default @human_feedback calls input() and blocks a worker.
# Wrong for a web app: the request thread waits on a terminal nobody sees.
# Right: pass provider=... to @human_feedback. The provider raises
# HumanFeedbackPending, kickoff() returns it, and the flow state is
# persisted automatically. Your approval endpoint later calls:
await flow.resume_async(feedback_text) # resume() from inside a running
# event loop raises RuntimeError
# LangGraph: the interrupt payload must be JSON-serialisable, because it
# is stored by the checkpointer and handed to whoever resumes the thread.
graph.invoke(Command(resume="approve"), {"configurable": {"thread_id": tid}})In practice both end in the same architecture: the run pauses, an approval record goes to a queue or a chat message, and a webhook resumes the run with the decision. LangGraph gets there with no extra code. CrewAI gets there once you implement a provider, and in exchange the reviewer's free-text reply can be classified into an outcome by a model.
Store the pause payload in your own approvals table as well as in the framework's state. Approvers, auditors and the finance team will ask which invoices are waiting and why, and that query should hit a normal table, not a checkpoint blob.
Each framework has a commercial platform from the same company. LangGraph traces into LangSmith, and LangSmith Deployment runs graphs as a managed service with persistence and a task queue. CrewAI has built-in tracing that reports to CrewAI AMP, its agent management platform, which can also host crews and flows. Neither framework requires its platform, but each one's visual tooling assumes it.
For a team that already runs its own infrastructure, the more important question is what you can see without the platform. LangGraph's get_state and get_state_history let you inspect every checkpoint of a thread from your own database. CrewAI gives you flow.plot() for the structure and flow.usage_metrics for token counts across every LLM call in a run, which helps with cost per invoice.
Run through these in order and stop at the first one that settles it.
Neither choice is a trap. LangGraph 1.0 committed to no breaking changes until 2.0, and CrewAI publishes versioned documentation for each release, so whichever you start with, pin the version and read the changelog before upgrading.
The deciding question is not which framework has more agents or more stars. It is what happens when a run stops halfway and resumes later in another process. LangGraph answers that with explicit checkpoints and a node that replays from the top; CrewAI answers it with persisted Flow state and a feedback provider you supply. Design the side effects for that answer first, then the prompts.