Async Programming for Agents
- Async lets an agent call multiple tools or LLMs at the same time instead of waiting one-by-one β massive speedup for I/O-bound work.
async def+awaitis Python's syntax;asyncio.gather()runs multiple coroutines concurrently.- LangChain/LangGraph have async variants for everything β
ainvoke,astream,abatchβ use them insideasync deffunctions.
Agents spend most of their time waiting β for LLM responses, API calls, database lookups, web searches. Synchronous code blocks on every wait. Async code lets the agent do other work while it waits, which means parallel tool calls, concurrent sub-agent execution, and responsive streaming β all on a single thread. One submodule per idea, ending with a cheat sheet.
Why agents specifically need asyncβ
Consider a ReAct agent that needs to search three different knowledge bases:
# Synchronous β each call blocks until complete
result_1 = search_arxiv(query) # 2 seconds
result_2 = search_wikipedia(query) # 1.5 seconds
result_3 = search_web(query) # 1 second
# Total: ~4.5 seconds β sequential
# Async β all three run concurrently
result_1, result_2, result_3 = await asyncio.gather(
search_arxiv(query),
search_wikipedia(query),
search_web(query),
)
# Total: ~2 seconds β limited by the slowest call
The async version is 2x faster β and the more tools you have, the bigger the win. This is why every serious agent framework supports async natively.
async/await basicsβ
An async def function returns a coroutine β a pausable function. You await it
to get the result. Nothing runs until you await it or schedule it on the event loop.
import asyncio
async def fetch_data(source: str) -> str:
"""Simulate an API call that takes 1 second."""
print(f"Starting {source}...")
await asyncio.sleep(1) # non-blocking sleep β other coroutines can run during this
print(f"Done {source}")
return f"Data from {source}"
async def main():
# Sequential β takes 3 seconds
a = await fetch_data("API-1")
b = await fetch_data("API-2")
c = await fetch_data("API-3")
# Parallel β takes 1 second
a, b, c = await asyncio.gather(
fetch_data("API-1"),
fetch_data("API-2"),
fetch_data("API-3"),
)
print(a, b, c)
asyncio.run(main()) # entry point β starts the event loop
Key rules:
async defdefines a coroutine functionawaitpauses the current coroutine until the awaited thing finishesasyncio.gather()runs multiple coroutines concurrently and returns all resultsasyncio.run()starts the event loop β call this once at the top level
The event loopβ
The event loop is the scheduler. It runs one coroutine until it hits an await, then
switches to another coroutine that's ready. This is concurrency, not parallelism β it's
single-threaded, but it never wastes time waiting.
Event loop timeline:
βββ coroutine A βββΊ await (I/O) βββ idle βββ resume A βββΊ done
βββ coroutine B βββββββββββββββΊ await (I/O) βββ idle βββ resume B βββΊ done
βββ coroutine C βββββββββββββββββββββββββββΊ await (I/O) βββ resume C βββΊ done
β²
all three overlap their wait times
For CPU-bound work (heavy computation), async doesn't help β use concurrent.futures or
multiprocessing instead. Agents are almost always I/O-bound, so async is the right choice.
Async tool callsβ
When you define tools for an agent, make them async if they do I/O:
import httpx
from langchain.tools import tool
@tool
async def search_weather(city: str) -> str:
"""Get the current weather for a city."""
async with httpx.AsyncClient() as client:
resp = await client.get(f"https://api.weather.com/v1/{city}")
data = resp.json()
return f"{city}: {data['temp']}F, {data['condition']}"
@tool
async def search_news(topic: str) -> str:
"""Search recent news articles about a topic."""
async with httpx.AsyncClient() as client:
resp = await client.get(f"https://api.news.com/search?q={topic}")
articles = resp.json()["articles"][:3]
return "\n".join(a["title"] for a in articles)
When a ReAct agent calls both tools in the same step, the framework can gather them
automatically β both HTTP requests fly out at the same time.
Async LLM callsβ
LangChain model wrappers have async methods for every operation:
from langchain.chat_models import init_chat_model
llm = init_chat_model("openai:gpt-4o")
# Sync
response = llm.invoke("What is the capital of France?")
# Async β use inside an async function
response = await llm.ainvoke("What is the capital of France?")
# Async batch β multiple prompts concurrently
responses = await llm.abatch([
"What is the capital of France?",
"What is the capital of Japan?",
"What is the capital of Brazil?",
])
Async streamingβ
For real-time UIs, stream tokens as they arrive. The async version lets you handle other events (like user cancellation) while streaming:
async def stream_response(query: str):
"""Stream an LLM response token by token."""
async for chunk in llm.astream(query):
print(chunk.content, end="", flush=True) # print each token as it arrives
# With a LangGraph agent β stream node updates
async def stream_agent(agent, query: str):
config = {"configurable": {"thread_id": "1"}}
async for event in agent.astream({"messages": [("user", query)]}, config):
for node_name, output in event.items():
print(f"[{node_name}]", output)
Running multiple agents concurrentlyβ
The real power shows when you have multiple agents and want to run them in parallel β for example, a researcher agent and a fact-checker agent that work at the same time:
async def run_research_pipeline(question: str):
# Run researcher and fact-checker concurrently on the same question
research_task = researcher_agent.ainvoke(
{"messages": [("user", f"Research: {question}")]}
)
factcheck_task = factchecker_agent.ainvoke(
{"messages": [("user", f"Find common misconceptions about: {question}")]}
)
research_result, factcheck_result = await asyncio.gather(
research_task, factcheck_task
)
# Combine results in a synthesis step
synthesis = await synthesizer_agent.ainvoke({
"messages": [("user", f"""
Research: {research_result}
Fact-check: {factcheck_result}
Synthesize a final answer.
""")]
})
return synthesis
asyncio.gather vs asyncio.create_taskβ
Two ways to run coroutines concurrently β gather for when you want all results together,
create_task for fire-and-forget:
async def main():
# gather β waits for ALL to finish, returns results in order
results = await asyncio.gather(
fetch_data("A"),
fetch_data("B"),
fetch_data("C"),
)
# results = ["Data from A", "Data from B", "Data from C"]
# create_task β schedule and continue immediately
task = asyncio.create_task(background_sync()) # runs in background
# ... do other work ...
result = await task # get result when you need it
# gather with error handling β return_exceptions=True avoids one failure killing all
results = await asyncio.gather(
fetch_data("A"),
fetch_data("B"), # even if this raises, A and C still return
fetch_data("C"),
return_exceptions=True, # exceptions become return values instead of propagating
)
Cheat sheetβ
| Task | Code |
|---|---|
| Define async function | async def my_func(): |
| Call async function | result = await my_func() |
| Run concurrently | results = await asyncio.gather(a(), b(), c()) |
| Start event loop | asyncio.run(main()) |
| Async LLM call | await llm.ainvoke(prompt) |
| Async streaming | async for chunk in llm.astream(prompt): |
| Async batch | await llm.abatch([prompt1, prompt2]) |
| Async agent invoke | await agent.ainvoke({"messages": ...}) |
| Async agent stream | async for event in agent.astream(...): |
| Background task | task = asyncio.create_task(my_func()) |
| Non-blocking sleep | await asyncio.sleep(1) |
- Blocking calls inside async functions β calling
requests.get()ortime.sleep()inside anasync defblocks the entire event loop. Usehttpx.AsyncClientandasyncio.sleep()instead. - Forgetting to
awaitβllm.ainvoke(prompt)withoutawaitreturns a coroutine object, not the result. Your agent silently uses<coroutine object>as a string. - Calling
asyncio.run()inside an already-running loop β this crashes. In Jupyter notebooks, useawait main()directly (the notebook already has a loop). In scripts, useasyncio.run()only at the top level. - Using sync tools with async agents β a sync tool blocks the event loop while it runs, killing all the concurrency benefits. Make I/O tools async.
- Not using
return_exceptions=Trueingatherβ one failed task raises an exception and cancels the others. Usereturn_exceptions=Truewhen you want partial results.
Quick self-check
Why is async better than threading for agents?
Agents are I/O-bound (waiting for APIs/LLMs), not CPU-bound. Async handles thousands of concurrent I/O operations on a single thread without the overhead and complexity of thread synchronization.
What does asyncio.gather() do?
It runs multiple coroutines concurrently and returns all their results as a list, in the same order they were passed in.
What's wrong with time.sleep(1) inside an async def?
It blocks the entire event loop for 1 second β no other coroutine can run during that time. Use await asyncio.sleep(1) instead, which yields control back to the loop.
What's the async equivalent of llm.invoke()?
await llm.ainvoke() β the a prefix is the LangChain convention for async methods.
Related: Pydantic for Agents Β· Agents Architecture Β· Glossary
Next: LangGraph Workflows β