Adaptive RAG & Self-RAG
- Self-RAG β the model decides whether to retrieve at all, generates an answer, then self-reflects: is it grounded? is it relevant? If reflection fails β retry with a better query.
- Adaptive RAG β a router inspects the query and picks the best strategy: vector store, web search, or direct answer β no one-size-fits-all pipeline.
- Both are built in LangGraph as graphs with conditional edges and reflection loops.
CRAG (previous section) checks docs after retrieval. But what if the question doesn't need retrieval at all? Or what if the generated answer is the thing that's wrong β not the docs? Self-RAG and Adaptive RAG push the intelligence further: the model reasons about when to retrieve and whether its own answer is good. One submodule per idea, ending with a cheat sheet.
Self-RAG: retrieve only when needed, then self-checkβ
The Self-RAG paper (Asai et al., 2023) trains a model to emit special reflection tokens at inference time:
- Retrieve β should I retrieve? (yes/no)
- ISREL β is the retrieved doc relevant to the query?
- ISSUP β is the generated sentence supported by the retrieved doc?
- ISUSE β is the overall response useful to the user?
In practice β rather than training a custom model β we simulate these checks with an LLM judging its own output. The pattern is:
question β should_retrieve? β (retrieve β grade docs) β generate β reflect β (retry or finish)
Adaptive RAG: route first, then executeβ
Adaptive RAG takes a different angle: instead of always doing the same pipeline, classify the query first and route it to the right strategy.
| Query type | Strategy |
|---|---|
| Simple factual ("What year was Python created?") | Direct LLM answer β no retrieval needed |
| Domain-specific ("Explain our refund policy") | Vector store retrieval β search your docs |
| Current events ("Latest OpenAI news") | Web search β your vector store is stale |
| Complex / multi-part | Multi-step retrieval β decompose and retrieve per sub-question |
The router is itself an LLM call with structured output β it reads the question and returns a strategy label.
Building Adaptive RAG + Self-RAG in LangGraphβ
We'll combine both ideas into one graph: route the query first (Adaptive), then execute the chosen strategy, then self-reflect (Self-RAG).
Step 1 β State and importsβ
from typing import List, Literal
from typing_extensions import TypedDict
from langchain.schema import Document
from langchain_openai import ChatOpenAI
from pydantic import BaseModel, Field
class AdaptiveRAGState(TypedDict):
question: str
documents: List[Document]
generation: str
route: str # "vectorstore" | "web_search" | "direct"
retry_count: int
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
Step 2 β Query routerβ
The router classifies the question and picks a strategy. Structured output keeps it clean.
class RouteDecision(BaseModel):
"""Route a user question to the best data source."""
datasource: Literal["vectorstore", "web_search", "direct"] = Field(
description="Route to 'vectorstore' for domain-specific questions, "
"'web_search' for current events, or 'direct' for simple factual questions."
)
router_llm = llm.with_structured_output(RouteDecision)
ROUTER_PROMPT = """You are a query router. Given a user question, decide the best strategy:
- "vectorstore" β the question is about domain-specific knowledge that would be in our docs.
- "web_search" β the question is about recent events or needs up-to-date information.
- "direct" β the question is simple enough to answer directly without any retrieval.
Question: {question}
"""
def route_question(state: AdaptiveRAGState) -> AdaptiveRAGState:
"""Classify the query and pick a retrieval strategy."""
question = state["question"]
result = router_llm.invoke(ROUTER_PROMPT.format(question=question))
return {**state, "route": result.datasource, "retry_count": 0}
Step 3 β The conditional edge after routingβ
def pick_strategy(state: AdaptiveRAGState) -> Literal["retrieve", "web_search", "direct_answer"]:
"""Route to the correct node based on the router's decision."""
route = state["route"]
if route == "vectorstore":
return "retrieve"
elif route == "web_search":
return "web_search"
else:
return "direct_answer"
Step 4 β Retrieval and web search nodesβ
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings
from langchain_community.tools.tavily_search import TavilySearchResults
vectorstore = Chroma(
collection_name="my_docs",
embedding_function=OpenAIEmbeddings(),
persist_directory="./chroma_db",
)
retriever = vectorstore.as_retriever(search_kwargs={"k": 4})
web_search_tool = TavilySearchResults(max_results=3)
def retrieve(state: AdaptiveRAGState) -> AdaptiveRAGState:
"""Retrieve from the vector store."""
docs = retriever.invoke(state["question"])
return {**state, "documents": docs}
def web_search(state: AdaptiveRAGState) -> AdaptiveRAGState:
"""Search the web for fresh context."""
results = web_search_tool.invoke({"query": state["question"]})
docs = [
Document(page_content=r["content"], metadata={"source": r["url"]})
for r in results
]
return {**state, "documents": docs}
def direct_answer(state: AdaptiveRAGState) -> AdaptiveRAGState:
"""Answer directly without retrieval."""
answer = llm.invoke(f"Answer concisely: {state['question']}").content
return {**state, "documents": [], "generation": answer}
Step 5 β Generate nodeβ
def generate(state: AdaptiveRAGState) -> AdaptiveRAGState:
"""Generate an answer from the retrieved context."""
context = "\n\n".join(doc.page_content for doc in state["documents"])
prompt = f"""Answer the question using ONLY the context below. If the context is
insufficient, say so.
Context:
{context}
Question: {state['question']}"""
answer = llm.invoke(prompt).content
return {**state, "generation": answer}
Step 6 β Self-reflection node (the Self-RAG part)β
After generation, the model checks its own output. Two questions:
- Is the answer grounded in the retrieved docs? (faithfulness)
- Does it actually answer the question? (relevance)
class ReflectionResult(BaseModel):
"""Self-reflection on generated answer quality."""
is_grounded: Literal["yes", "no"] = Field(
description="Is the answer supported by the provided documents?"
)
answers_question: Literal["yes", "no"] = Field(
description="Does the answer actually address the user's question?"
)
reflection_llm = llm.with_structured_output(ReflectionResult)
REFLECTION_PROMPT = """You are a quality checker. Given a question, retrieved documents,
and a generated answer, evaluate:
1. Is the answer grounded in (supported by) the documents? β is_grounded
2. Does the answer actually address the user's question? β answers_question
Question: {question}
Documents: {documents}
Answer: {answer}
"""
def reflect(state: AdaptiveRAGState) -> AdaptiveRAGState:
"""Self-reflect on the generated answer."""
docs_text = "\n".join(doc.page_content for doc in state["documents"])
prompt = REFLECTION_PROMPT.format(
question=state["question"],
documents=docs_text,
answer=state["generation"],
)
result = reflection_llm.invoke(prompt)
# if reflection fails, bump retry count
if result.is_grounded == "no" or result.answers_question == "no":
return {**state, "retry_count": state["retry_count"] + 1}
return state # passes β keep the answer
Step 7 β Decide after reflectionβ
def after_reflection(state: AdaptiveRAGState) -> Literal["rewrite", "finish"]:
"""If reflection failed and we haven't retried too many times, rewrite and retry."""
if state["retry_count"] > 0 and state["retry_count"] <= 2:
return "rewrite"
return "finish"
def rewrite_query(state: AdaptiveRAGState) -> AdaptiveRAGState:
"""Rewrite the question for a better retrieval attempt."""
rewrite_prompt = f"""Rewrite this question to be clearer and more specific for search:
Original: {state['question']}
Rewritten:"""
rewritten = llm.invoke(rewrite_prompt).content.strip()
return {**state, "question": rewritten, "documents": []}
Step 8 β Compile the full graphβ
from langgraph.graph import StateGraph, END
workflow = StateGraph(AdaptiveRAGState)
# nodes
workflow.add_node("route_question", route_question)
workflow.add_node("retrieve", retrieve)
workflow.add_node("web_search", web_search)
workflow.add_node("direct_answer", direct_answer)
workflow.add_node("generate", generate)
workflow.add_node("reflect", reflect)
workflow.add_node("rewrite", rewrite_query)
# entry
workflow.set_entry_point("route_question")
# routing after classification
workflow.add_conditional_edges(
"route_question",
pick_strategy,
{
"retrieve": "retrieve",
"web_search": "web_search",
"direct_answer": "direct_answer",
},
)
# after retrieval / web search β generate
workflow.add_edge("retrieve", "generate")
workflow.add_edge("web_search", "generate")
# direct answer skips generation and reflection
workflow.add_edge("direct_answer", END)
# after generation β reflect
workflow.add_edge("generate", "reflect")
# after reflection β finish or retry
workflow.add_conditional_edges(
"reflect",
after_reflection,
{
"rewrite": "rewrite",
"finish": END,
},
)
# rewrite loops back to retrieve
workflow.add_edge("rewrite", "retrieve")
adaptive_rag = workflow.compile()
Running itβ
# domain-specific question β routed to vectorstore
result = adaptive_rag.invoke({
"question": "How does our cancellation policy work?",
"documents": [],
"generation": "",
"route": "",
"retry_count": 0,
})
print(result["generation"])
# current events β routed to web search
result = adaptive_rag.invoke({
"question": "What did OpenAI announce this week?",
"documents": [],
"generation": "",
"route": "",
"retry_count": 0,
})
print(result["generation"])
Self-RAG vs CRAG vs Adaptive RAGβ
| CRAG | Self-RAG | Adaptive RAG | |
|---|---|---|---|
| Checks | Docs after retrieval | Docs + generated answer | Query before retrieval |
| Decision | Generate or web search | Retrieve or not, answer quality | Which strategy to use |
| Loop | No (single pass) | Yes (reflect β retry) | Depends on combination |
| Strength | Filters bad docs | Catches hallucinations | Right tool for each query |
In practice you combine them β Adaptive routing at the front, CRAG-style grading in the middle, Self-RAG reflection at the end. That's exactly what our graph above does.
Cheat sheetβ
| Task | Code |
|---|---|
| Route a query | llm.with_structured_output(RouteDecision).invoke(prompt) |
| Conditional routing | workflow.add_conditional_edges("route_question", pick_fn, {...}) |
| Self-reflection | llm.with_structured_output(ReflectionResult).invoke(prompt) |
| Retry loop | conditional edge from reflect β rewrite β retrieve (capped) |
| Direct answer (no retrieval) | skip retrieval entirely for simple questions |
- No retry cap on the reflection loop β without a max retry count, the graph can loop forever rewriting and re-retrieving. Always cap it (2β3 retries is plenty).
- Routing with a single keyword match β use an LLM for routing, not regex. Queries are ambiguous and keyword rules break on edge cases.
- Reflecting without the source docs β the reflection check needs the retrieved docs to judge groundedness. Don't just compare the answer to the question.
- Skipping reflection for direct answers β if the model answered without retrieval, there's
nothing to ground-check, so route
direct_answerstraight to END.
Quick self-check
What two things does Self-RAG's reflection check?
1) Is the answer grounded in the retrieved documents? (faithfulness) 2) Does the answer actually address the user's question? (relevance)
How does Adaptive RAG differ from CRAG?
CRAG checks docs after retrieval and falls back to web search. Adaptive RAG classifies the query before retrieval and routes to the best strategy (vectorstore, web search, or direct answer) upfront.
Why cap the reflection retry loop?
Without a cap, a query that consistently produces poor answers will loop forever β rewriting, re-retrieving, regenerating, and reflecting endlessly. A cap of 2β3 retries prevents this.
When should Adaptive RAG route to "direct" (no retrieval)?
For simple factual questions the LLM already knows well β retrieval adds latency and cost without improving the answer.
Related: Corrective RAG Β· Agentic RAG Β· Agents Architecture
Next: RAG with Persistent Memory β coming soon.