LangGraph — deep dive¶
At a glance¶
- Package / repo:
langgraph— pinnedlanggraph==1.2.11,langchain-core==1.6.1,langchain-openai==1.6.0 - Licence: MIT
- Adapter:
frameworks/langgraph/adapter.py - Status: runs all seven arenas; loses one
resilienceitem (res-01). No live scorecard yet (no API key wired in)
Wiring notes¶
- LLM:
ChatOpenAI(model=..., base_url=config.base_url, api_key=config.api_key, temperature=0.0). Points cleanly at the shared gateway — no special handling needed for mock vs live. - Tools: the shared implementations wrapped with
@toolfromlangchain_core.tools. The wrapper only forwards; it never changes what a tool computes. Only the tools the arena declares are registered (arena.tools.names_for). - Iteration budget:
config={"recursion_limit": 2 * max_tool_iterations}. One tool round is two graph steps (model node + tool node), so the limit has to be doubled to line up with the other adapters' LLM-call budgets. It was briefly2N + 2, which quietly bought LangGraph one extra model call. - Metrics: summed from
msg.usage_metadataacross the returned messages; tool calls read frommsg.tool_calls. - Human-in-the-loop: implemented natively — see below.
Native interrupts¶
The request_approval tool calls langgraph.types.interrupt(...). LangGraph
checkpoints the graph at that point and invoke returns with __interrupt__
set; resume continues the same thread with Command(resume=decision), and the
interrupt(...) call inside the tool returns that decision so the graph carries
on from exactly where it stopped.
Nothing about the transcript is reconstructed by hand. That is the real
distinction from the vanilla baseline, which emulates the pause by carrying the
message list back in — an emulation that would not survive a process restart.
Two details the adapter has to get right, both of which are asserted in
tests/test_suspend_resume.py:
- A resumed
invokereturns the whole thread. The harness sums cost across legs, so the second leg must report only the messages added since the pause (seenslicing) or every token on a paused item is counted twice. request_approvalis not logged as a tool call. Asking permission is the pause, not an action taken. The baseline does not log it either, and the arena'sno_tool_before_suspendcheck compares the two adapters directly.
Gotchas¶
create_react_agentis deprecated in LangGraph 1.0 (moves tolangchain.agents.create_agentin 2.0). Deliberately not migrated yet — the replacement lives in thelangchainpackage, which this adapter does not install, on a different version track from the pinnedlangchain-core. See the deprecation register in ../dependencies.md.interruptneeds a checkpointer and athread_id, so the adapter only attachesMemorySaverfor arenas that declarerequest_approval. Attaching it everywhere would change behaviour on arenas that never pause.- Loses one
resilienceitem (res-01): it gives up on malformed tool arguments where the stdlib baseline recovers. See thecomparisonCI job.
Results¶
Mock mode only so far. Pass rates in mock mode are ~100% by construction and are not a quality signal — they prove the adapter is wired correctly. The two columns that do compare honestly are marked.
| Arena | Mode | Pass rate | Note |
|---|---|---|---|
tool_use |
mock | 15/15 | 754 prompt tok/item — ties the hand-rolled baseline byte for byte, see overhead.md (comparable) |
structured_output |
mock | 15/15 | |
rag |
mock | 15/15 | |
multi_agent |
mock | 10/10 | single-agent role-play entry |
resilience |
mock | 7/8 | (comparable) — fails res-01, malformed tool args |
human_in_the_loop |
mock | 12/12 | native interrupt + checkpointer |