Pydantic for Agents
- Pydantic BaseModel gives you typed, validated data structures β agents need these for tool inputs, LLM outputs, and state.
- Structured output (
with_structured_output(MyModel)) forces the LLM to return a Pydantic model instead of raw text β no more parsing regex. - Field validators catch bad data before it reaches your agent logic β fail fast, not halfway through a tool call.
Agents pass data between components constantly β the LLM produces output, a tool consumes it, the tool returns a result, the agent updates its state. If any of that data is malformed, the agent silently breaks or hallucinates downstream. Pydantic makes every hand-off typed and validated, so bad data fails loud and early. One submodule per idea, ending with a cheat sheet.
Why agents need typed dataβ
Without Pydantic, agent I/O looks like this:
# Untyped β the agent's tool gets a raw dict from the LLM
def search_flights(args: dict):
origin = args["origin"] # KeyError if LLM forgot it
destination = args["dest"] # or was it "destination"? who knows
date = args["date"] # "tomorrow" or "2025-03-15"? no validation
max_price = args["max_price"] # string "500" or int 500?
...
Every line is a landmine. The LLM might spell a key differently, omit a field, or return a string where you need an int. Pydantic eliminates all of this.
BaseModel basicsβ
A BaseModel is a class where each attribute has a type annotation. Pydantic validates data
on construction β if the types don't match, it raises ValidationError immediately.
from pydantic import BaseModel, Field
from datetime import date
class FlightSearch(BaseModel):
"""Input schema for the flight search tool."""
origin: str = Field(description="IATA airport code, e.g. 'SFO'")
destination: str = Field(description="IATA airport code, e.g. 'JFK'")
departure_date: date = Field(description="Departure date in YYYY-MM-DD format")
max_price: float = Field(default=1000.0, ge=0, description="Maximum ticket price in USD")
# Valid β works fine
search = FlightSearch(origin="SFO", destination="JFK", departure_date="2025-06-15", max_price=500)
print(search.departure_date) # date(2025, 6, 15) β auto-coerced from string
# Invalid β raises ValidationError instantly
FlightSearch(origin="SFO", destination="JFK", departure_date="not-a-date", max_price=-50)
# ValidationError: departure_date β invalid date format, max_price β >= 0
Key things to notice: Pydantic coerces "2025-06-15" into a date object automatically,
and ge=0 on max_price rejects negative values. The Field(description=...) is critical
for agents β the LLM reads these descriptions to understand what each field expects.
Field validators and constraintsβ
Beyond basic types, you can add custom validation logic with field_validator:
from pydantic import BaseModel, Field, field_validator
class AgentAction(BaseModel):
"""Represents a single action the agent wants to take."""
tool_name: str = Field(description="Name of the tool to call")
arguments: dict = Field(default_factory=dict, description="Tool arguments")
reasoning: str = Field(description="Why the agent chose this action")
@field_validator("tool_name")
@classmethod
def tool_must_be_allowed(cls, v):
allowed = {"search", "calculator", "weather", "send_email"}
if v not in allowed:
raise ValueError(f"Unknown tool '{v}'. Allowed: {allowed}")
return v
@field_validator("reasoning")
@classmethod
def reasoning_not_empty(cls, v):
if len(v.strip()) < 10:
raise ValueError("Reasoning must be at least 10 characters β make the agent explain itself")
return v.strip()
# This passes
action = AgentAction(
tool_name="search",
arguments={"query": "latest AI news"},
reasoning="The user asked about recent developments in AI"
)
# This fails β unknown tool
AgentAction(tool_name="hack_pentagon", arguments={}, reasoning="Just curious")
# ValidationError: Unknown tool 'hack_pentagon'
Validators act as guardrails β they constrain what the agent can do before any action is executed. Think of them as the cheapest, fastest safety layer you can add.
Structured LLM outputβ
This is where Pydantic becomes essential for agents. Instead of parsing free-text LLM responses with fragile regex, you tell the model to return a specific Pydantic schema:
from pydantic import BaseModel, Field
from langchain.chat_models import init_chat_model
class MovieRecommendation(BaseModel):
"""A movie recommendation with reasoning."""
title: str = Field(description="Movie title")
year: int = Field(description="Release year")
genre: str = Field(description="Primary genre")
reason: str = Field(description="Why this movie fits the user's request")
confidence: float = Field(ge=0, le=1, description="How confident the model is (0-1)")
llm = init_chat_model("openai:gpt-4o")
structured_llm = llm.with_structured_output(MovieRecommendation)
result = structured_llm.invoke("Suggest a sci-fi movie for someone who loved Arrival")
print(type(result)) # <class 'MovieRecommendation'> β not a string!
print(result.title) # "Interstellar"
print(result.confidence) # 0.92
The LLM's response is guaranteed to match your schema β or it raises an error. No more
json.loads() inside a try/except hoping the model returned valid JSON.
Tool input validationβ
When an agent calls a tool, the tool's input should be a Pydantic model. This way, if the LLM passes garbage arguments, you catch it before executing anything:
from pydantic import BaseModel, Field
from langchain.tools import tool
class WeatherInput(BaseModel):
"""Input for the weather lookup tool."""
city: str = Field(description="City name")
units: str = Field(default="celsius", description="Temperature units: 'celsius' or 'fahrenheit'")
@field_validator("units")
@classmethod
def validate_units(cls, v):
if v not in ("celsius", "fahrenheit"):
raise ValueError(f"Units must be 'celsius' or 'fahrenheit', got '{v}'")
return v
@tool(args_schema=WeatherInput)
def get_weather(city: str, units: str = "celsius") -> str:
"""Get the current weather for a city."""
# By the time we get here, city is a valid string and units is celsius/fahrenheit
return f"Weather in {city}: 22 degrees {units}"
The args_schema=WeatherInput tells LangChain to validate the LLM's tool-call arguments
against this model before invoking the function. Bad arguments never reach your code.
Defining agent state schemasβ
In LangGraph, agent state flows between nodes. Define it as a Pydantic model so every node can trust the shape of the data:
from pydantic import BaseModel, Field
from typing import List, Optional
from langchain.schema import Document
class ResearchAgentState(BaseModel):
"""State for a research agent that gathers and synthesizes information."""
query: str = Field(description="The user's research question")
search_results: List[Document] = Field(default_factory=list)
current_summary: str = Field(default="")
sources_checked: int = Field(default=0, ge=0)
is_sufficient: bool = Field(default=False)
final_answer: Optional[str] = Field(default=None)
# Every node receives and returns this typed state
def search_node(state: ResearchAgentState) -> dict:
docs = retriever.invoke(state.query)
return {
"search_results": docs,
"sources_checked": state.sources_checked + len(docs),
}
def evaluate_node(state: ResearchAgentState) -> dict:
# The agent can check: do I have enough sources?
if state.sources_checked >= 10 or state.is_sufficient:
return {"is_sufficient": True}
return {"is_sufficient": False}
With a typed state, you get autocompletion in your editor, clear documentation of what data flows through the graph, and instant errors if a node returns malformed data.
How bad data breaks agents β a real exampleβ
Here's what happens without Pydantic in a multi-step agent:
# Step 1: LLM returns {"action": "search", "query": "AI papers"} β fine
# Step 2: Tool returns {"results": [...]} β fine
# Step 3: LLM returns {"action": "summarize", "text": None} β uh oh
# Step 4: Summarize tool does text.split() β AttributeError: NoneType has no attribute 'split'
# But the error is in step 4, caused by step 3. Good luck debugging.
With Pydantic, step 3 would have raised ValidationError: text β none is not an allowed value
immediately. The error points to the exact problem, at the exact moment it happens.
Cheat sheetβ
| Task | Code |
|---|---|
| Define a model | class MyModel(BaseModel): field: type = Field(...) |
| Add constraints | Field(ge=0, le=100), Field(min_length=1) |
| Custom validator | @field_validator("field") classmethod |
| Structured LLM output | llm.with_structured_output(MyModel) |
| Tool input schema | @tool(args_schema=MyModel) |
| Serialize to dict | model.model_dump() |
| Serialize to JSON | model.model_dump_json() |
| Parse from dict | MyModel.model_validate({"field": "value"}) |
- Forgetting
Field(description=...)β the LLM uses these descriptions to fill in the schema correctly. No description = guessing. - Using
dictinstead of a Pydantic model for tool args β you lose validation and the agent silently passes bad data. - Overcomplicating schemas β deeply nested models confuse LLMs. Keep tool input schemas flat when possible.
- Not handling
ValidationErrorβ in production, catch it and feed the error back to the agent so it can retry with corrected arguments.
Quick self-check
Why do agents need typed data more than normal programs?
Because agents pass data between non-deterministic components (LLMs, tools, state). Any hand-off can produce malformed data, and without types you won't catch it until something breaks downstream.
What does with_structured_output(MyModel) do?
It forces the LLM to return a response that matches your Pydantic model's schema β you get a validated Python object instead of raw text.
How does args_schema protect a tool?
It validates the LLM's tool-call arguments against the Pydantic model before invoking the function. Bad arguments raise a ValidationError instead of reaching your code.
What's the benefit of using Pydantic for agent state in LangGraph?
Every node can trust the shape and types of the state β you get autocompletion, clear documentation, and instant errors if a node returns malformed data.
Related: Agents Architecture Β· Agentic RAG Β· Glossary