When a user asks the support assistant, a Retrieval Augmented Generation (RAG) application built with LangChain to compare two products across three dimensions, they’re effectively posing six questions simultaneously. Similarity search uses a single query vector to encapsulate all the intents. The retriever then generates the best approximation of the average of those intents. The resulting answer comes back concise. The search executes without errors. The relevance scores look reasonable. Yet the retrieved chunks, while topically relevant, only cover a fraction of what the question actually asked.
In this post, we showcase a RAG application on Amazon Bedrock Managed Knowledge Base with LangChain. We run the same multi-part question through standard and agentic retrieval, and read the trace events to see the plan the model produced. We also cover what the two retrieval paths cost and when the cheaper one is the right choice.
Agentic retrieval is available on Amazon Bedrock Managed Knowledge Base. Instead of one search, Amazon Bedrock Managed Knowledge Base plans the retrieval. It breaks the question into sub-queries, runs them, judges whether it has enough evidence, and searches again if it doesn’t. The langchain-aws package exposes both agentic and standard retrieval, so you can use either from a LangChain application.
Solution overview
Amazon Bedrock Managed Knowledge Base, the fully managed RAG capability in Amazon Bedrock, removes the self-managed vector store, embeddings, and re-ranking models from the RAG architecture. You configure a data source, and Amazon Bedrock Managed Knowledge Bases handles chunking, embedding, storage, and retrieval. This walkthrough uses Amazon Simple Storage Service (Amazon S3).
Amazon Bedrock Managed Knowledge Bases provides two APIs. We briefly discuss those differences in this post. The Retrieve API runs one hybrid search and returns scored chunks. The AgenticRetrieveStream API runs a planning loop and streams the steps back to you as trace events. In the langchain-aws package, the first is a standard LangChain retriever you can drop into a chain. The second is a function retrieval directly from a knowledge base.
The following diagram shows the solution architecture. The application queries Amazon Bedrock Knowledge Bases using either the Retrieve API (standard, single-shot) or the AgenticRetrieveStream API (multi-step planning loop). Both paths return document chunks from the knowledge base, which the application then uses to generate a grounded response.
Figure 1: Solution architecture for querying Amazon Bedrock Knowledge Bases with the Retrieve and AgenticRetrieveStream APIs
Implementation walkthrough
The following sections walk you through creating a knowledge base, querying it with both retrieval methods, and reading the trace events the agentic planner produces.
Prerequisites
To follow along you need:
- An AWS account with access to Amazon Bedrock in a Region where Amazon Bedrock Managed Knowledge Bases and agentic retrieval are available. This walkthrough uses the US East (N. Virginia) Region (
us-east-1), and the code assumes it throughout. Check the AWS documentation for other Regional availability and support. - Two AWS Identity and Access Management (IAM) identities, described in the next section: a service role the knowledge base assumes, and permissions on the identity you call the APIs from.
- Python 3.12 or later.
- An S3 bucket holding the sample documents. The corpus needs several documents that cover overlapping topics so that a comparative question has somewhere to go. A single flat document cannot demonstrate query planning.
Install the packages. The Boto3 version matters: agentic_retrieve_stream did not exist before 1.43.32.
langchain-aws>=1.6.3
langchain>=1.0
boto3>=1.43.32
Permissions
Two identities are involved and separating them is worth doing deliberately. The knowledge base assumes a service role to read your documents and call the embedding model. Your application uses an AWS Security Token Service (AWS STS) caller identity to query. Neither needs the other’s permissions.
Amazon Bedrock creates the service role for you if you let it. To supply your own, give it a trust policy that lets Amazon Bedrock assume it. Scope it with aws:SourceAccount and aws:SourceArn so that another account can’t use it as a confused deputy:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {"Service": "bedrock.amazonaws.com"},
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {"aws:SourceAccount": "111122223333"},
"ArnLike": {
"aws:SourceArn": "arn:aws:bedrock:us-east-1:111122223333:knowledge-base/*"
}
}
}]
}
The service role also needs s3:ListBucket on your bucket and s3:GetObject on its contents, both conditioned on aws:ResourceAccount. Scope the knowledge-base/* wildcard character down to specific knowledge base IDs after you have created them.
The AWS STS caller identity needs a different set. bedrock:AgenticRetrieveStream and bedrock:InvokeModelWithResponseStream can’t be scoped to a knowledge base Amazon Resource Name (ARN). bedrock:Retrieve and bedrock:GetDocumentContent can:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AgenticRetrievalAndPlannerModel",
"Effect": "Allow",
"Action": [
"bedrock:AgenticRetrieveStream",
"bedrock:InvokeModelWithResponseStream"
],
"Resource": "*"
},
{
"Sid": "RetrieveAndFullDocumentExpansion",
"Effect": "Allow",
"Action": ["bedrock:Retrieve", "bedrock:GetDocumentContent"],
"Resource": "arn:aws:bedrock:<region>:111122223333:knowledge-base/<knowledge-base-id>"
},
{
"Sid": "GenerateAnswersInTheChains",
"Effect": "Allow",
"Action": ["bedrock:InvokeModel", "bedrock:Converse", "bedrock:ConverseStream"],
"Resource": "*"
}
]
}
bedrock:GetDocumentContent is often overlooked. Agentic retrieval calls it when a FullDocumentExpansion step decides a passage lacks the context to answer. A policy with only bedrock:Retrieve works until the planner reaches for a whole document and then fails partway through a query.
To create and manage the knowledge base itself, the calling role additionally needs bedrock:CreateKnowledgeBase on *, and the GetKnowledgeBase, UpdateKnowledgeBase, DeleteKnowledgeBase, StartIngestionJob, GetIngestionJob, and ListIngestionJobs actions on knowledge-base/*. If you’re using guardrails, add bedrock:GetGuardrail and bedrock:ApplyGuardrail.
Running this walkthrough might incur costs for document storage and ingestion in the knowledge base, retrieval calls, and foundation model (FM) inference.
For more information about pricing, see the Knowledge Bases section of Amazon Bedrock pricing.
Delete the resources when you complete this experiment.
Creating and populating the knowledge base
Create the knowledge base with a managedKnowledgeBaseConfiguration. Setting embeddingModelType to MANAGED uses the service-managed embedding model.
import boto3
import os
REGION = os.environ["AWS_REGION"]
bedrock_agent = boto3.client("bedrock-agent", region_name=REGION)
response = bedrock_agent.create_knowledge_base(
name=KB_NAME,
roleArn=KB_ROLE_ARN,
knowledgeBaseConfiguration={
"type": "MANAGED",
"managedKnowledgeBaseConfiguration": {
"embeddingModelType": "MANAGED",
},
},
)
KB_ID = response["knowledgeBase"]["knowledgeBaseId"]
There’s no storageConfiguration in that request. For a self-managed knowledge base you would pass one describing your vector store. Amazon Bedrock Managed Knowledge Base does not take one, which is the clearest signal in the API that Amazon Bedrock owns the storage layer.
Attach the S3 bucket as a data source, then start an ingestion job. Ingestion is asynchronous, so poll until the job reaches a terminal state rather than sleeping for a fixed interval and hoping.
import time
SUCCESS_STATES = frozenset({"COMPLETE"})
FAILURE_STATES = frozenset({"FAILED", "STOPPED"})
def wait_for_ingestion(kb_id, ds_id, job_id, timeout_s=1800):
"""Poll an ingestion job until it reaches a terminal state."""
deadline = time.time() + timeout_s
while time.time() < deadline:
job = bedrock_agent.get_ingestion_job(
knowledgeBaseId=kb_id,
dataSourceId=ds_id,
ingestionJobId=job_id,
)["ingestionJob"]
status = job["status"]
if status in SUCCESS_STATES:
return job
if status in FAILURE_STATES:
reasons = job.get("failureReasons") or ["no reason reported"]
raise RuntimeError(f"Ingestion job {job_id} finished as {status}: " + "; ".join(reasons))
time.sleep(15)
raise TimeoutError(f"Ingestion job {job_id} did not finish in {timeout_s}s")
The full data source configuration and error handling are in the sample repository.
Querying with the LangChain retriever
AmazonKnowledgeBasesRetriever wraps the Retrieve API and behaves like any other LangChain retriever. For Amazon Bedrock Managed Knowledge Bases, pass managedSearchConfiguration. This is the part that trips people up: vectorSearchConfiguration is the previous path for knowledge bases where you run your own vector store. It is what most existing examples show.
from langchain_aws.retrievers import AmazonKnowledgeBasesRetriever
SIMPLE_QUERY = "What is the restore time objective for the checkout service?"
retriever = AmazonKnowledgeBasesRetriever(
knowledge_base_id=KB_ID,
region_name=REGION,
retrieval_config={
"managedSearchConfiguration": {
"numberOfResults": 5,
}
},
)
docs = retriever.invoke(SIMPLE_QUERY)
Each result comes back as a LangChain Document. The relevance score is in metadata["score"], and the source document’s own metadata is under metadata["source_metadata"], renamed so it does not collide. If you want to drop low-confidence results, set min_score_confidence on the retriever instead of filtering afterward.
For a question with one clear intent, this is the right tool. It is one call. The latency is the lowest of the two options, and you keep full control of how the answer gets generated. Most of the queries a production assistant sees are this shape, and reaching for a planning loop to answer them wastes money and time.
Where single-shot retrieval runs out
Now give the same retriever a question with several parts:
COMPLEX_QUERY = (
"Compare the checkout and inventory services across on-call escalation, backup and "
"restore targets, and deployment rollback procedure. Where do they differ?"
)
docs = retriever.invoke(COMPLEX_QUERY)
Five chunks come back, ranked by hybrid score against one embedding of that whole question.
That question contains six intents: two services across three dimensions. Scoring the retrieved text for evidence of each one gives a concrete measure of what a single embedding recovers.
| numberOfResults | Chunks | Share of corpus | Sub-intents covered | Missing |
| 5 | 5 | 10% | 4 of 6 | checkout on-call, inventory restore |
| 10 | 10 | 19% | 6 of 6 | none |
At five results, one embedding standing for six intents misses two of them. At ten it covers all six, with visible waste: two sub-intents are covered twice and one chunk carries none.
The retriever did its job. The limitation is structural: one vector cannot represent six intents, and there is no step in the process that asks whether the returned evidence is enough to answer the question.
Running agentic retrieval
Agentic retrieval is not a LangChain retriever, but a feature of Amazon Bedrock Managed Knowledge Bases. The langchain-aws package exposes it as a standalone function, agentic_retrieve, because the underlying API streams its results and doesn’t fit the synchronous BaseRetriever interface. There’s no flag on AmazonKnowledgeBasesRetriever that switches it on.
from langchain_aws.retrievers.bedrock import agentic_retrieve
result = agentic_retrieve(
knowledge_base_id=KB_ID,
query=COMPLEX_QUERY,
region_name=REGION,
generate_response=True,
number_of_results=10,
)
print(result["generatedResponse"]["answer"])
With generate_response=True, the service returns a grounded answer and citations alongside the retrieved chunks, so you get an answer without wiring up a separate model call. The function works only against Amazon Bedrock Managed Knowledge Base.
Internally, the service plans, retrieves, evaluates whether the evidence is sufficient, and iterates if it isn’t. The helper hides all of that and hands back the final chunks, which is convenient and means you can’t see the plan.
Reading the trace events
To watch the model decompose the question, call agentic_retrieve_stream on the bedrock-agent-runtime client directly. This is the one place in this walkthrough where we step around langchain-aws, because the helper discards trace events and doesn’t expose maxAgentIteration or a custom planner model.
runtime = boto3.client("bedrock-agent-runtime", region_name=REGION)
response = runtime.agentic_retrieve_stream(
messages=[{"role": "user", "content": {"text": COMPLEX_QUERY}}],
retrievers=[{
"configuration": {
"knowledgeBase": {
"knowledgeBaseId": KB_ID,
"retrievalOverrides": {"maxNumberOfResults": 10},
}
}
}],
agenticRetrieveConfiguration={
"foundationModelType": "MANAGED",
"rerankingModelType": "MANAGED",
# 5 is the API default. Below 4 the planner stops decomposing entirely.
"maxAgentIteration": 5,
},
generateResponse=False,
)
for event in response["stream"]:
if "traceEvent" in event:
attrs = event["traceEvent"]["attributes"]
print(f"{attrs.get('step')}: {attrs.get('status')}")
for action in attrs.get("actions", []) or []:
if "retrieve" in action:
query = action["retrieve"].get("inputQuery", {}).get("text", "")
print(f" sub-query: {query}")
elif "result" in event:
for chunk in event["result"].get("results", []):
print(chunk.get("content", {}).get("text", "")[:120])
Set generateResponse to False when you only want the retrieval behavior. The API generates a grounded answer by default, which costs an extra model call you may not need while you are inspecting the plan.
A trace event’s step tells you where the planner is. SpeculativeRetrieval runs before the first plan to cut latency and doesn’t count against your iteration budget. Planning is where the model reads the question and prior results and emits sub-queries.