Add Web Search to a LlamaIndex Agent
Write a plain Python function that calls SerpexClient.search(), wrap it with FunctionTool.from_defaults, and pass it to a FunctionAgent. The agent then runs a live web search whenever it decides it needs one. Set include_content on the same call and each top result comes back with its page as markdown, which is what you want for RAG.
This is the same pattern as our LangChain and CrewAI posts, translated to LlamaIndex's current agent API. Only the wrapper changes. The SDK call is identical.
What do you need before you start?
- Python 3.10 or newer.
llama-index-coredeclares>=3.10on PyPI. pip install serpex llama-index-core llama-index-llms-openai- A Serpex API key from app.serpex.dev. New accounts get 200 free credits, no card required.
- The key in an environment variable,
SERPEX_API_KEY. The example LLM is OpenAI's, so you'll also needOPENAI_API_KEYset, as in LlamaIndex's own agent tutorial. Any LLM that LlamaIndex supports works in its place.
How do you turn Serpex into a LlamaIndex tool?
LlamaIndex builds a tool from an ordinary function. The tool name comes from the function name, and the description comes from the signature plus the docstring, so the docstring is what the model reads when it decides whether to search.
import osfrom llama_index.core.tools import FunctionToolfrom serpex import SerpexClientclient = SerpexClient(os.environ["SERPEX_API_KEY"])def web_search(query: str) -> str:"""Search the live web and return titles, URLs, and snippets for a query."""response = client.search({"q": query})if not response.results:return "No results found."lines = [f"{r.title}\n{r.url}\n{r.snippet}" for r in response.results]return "\n\n".join(lines)search_tool = FunctionTool.from_defaults(web_search)
client.search() returns a SearchResponse. Each item in response.results is a SearchResult with title, url, snippet and position. You don't pick a search engine anywhere: the API routes each query automatically.
The SDK call is synchronous. That's fine inside an async agent, because FunctionTool runs a sync function in a thread executor rather than on the event loop.
How do you give the tool to an agent?
FunctionAgent from llama_index.core.agent.workflow is LlamaIndex's prebuilt tool-calling agent. It takes a list of tools, an LLM, and a system prompt. You could also pass web_search itself, since the agent converts plain functions into FunctionTools for you. Wrapping it yourself just makes the name and description explicit.
import asynciofrom llama_index.core.agent.workflow import FunctionAgentfrom llama_index.llms.openai import OpenAIagent = FunctionAgent(tools=[search_tool],llm=OpenAI(model="gpt-4o-mini"),system_prompt=("You are a research assistant. Use web_search for anything current ""or anything you're not sure about, and cite the URLs you used."),)async def main():response = await agent.run(user_msg="What changed in the latest LlamaIndex release?")print(response)if __name__ == "__main__":asyncio.run(main())
agent.run() is async, so plain scripts need the asyncio.run(main()) wrapper. In a notebook you can await it directly. Printing the result gives you the final answer text.
If you're building several agents that hand work to each other, LlamaIndex's AgentWorkflow takes the same tools. The search tool doesn't care which one calls it.
How do you get page content back for RAG?
Snippets are one or two lines per result. That's enough to answer "who announced this" and not enough to summarize a page or ground an answer in it. Set include_content and Serpex fetches the page for each top result in the same request and returns it as markdown in the result's content field.
def web_search_with_content(query: str) -> str:"""Search the live web and return the page content of the top resultsas markdown, for questions that need the full page, not a snippet."""response = client.search({"q": query,"include_content": True,"content_results": 5,})chunks = []for r in response.results:if r.content:chunks.append(f"# {r.title}\n{r.url}\n\n{r.content}")elif r.content_error:chunks.append(f"# {r.title}\n{r.url}\n(content unavailable: {r.content_error})")return "\n\n---\n\n".join(chunks)content_tool = FunctionTool.from_defaults(web_search_with_content)
content_results must be exactly 5 or 10. The SDK rejects anything else before sending.
Fetching is best effort. Content comes back for roughly 70% of the results you ask for. A result that couldn't be fetched carries content_error (for example timeout, or a page that robots.txt disallows) instead of content, never both, and still has its title, URL and snippet.
If you'd rather index the pages than hand them to the agent, turn each result into a LlamaIndex Document and keep the URL as metadata:
from llama_index.core import Documentresponse = client.search({"q": "llamaindex function agent", "include_content": True})docs = [Document(text=r.content, metadata={"url": r.url, "title": r.title})for r in response.resultsif r.content]
From there it's the usual LlamaIndex path, for example VectorStoreIndex.from_documents(docs).
What does it cost per call?
| Call | Credits |
|---|---|
| Plain search | 1 |
include_content, 1 to 5 pages delivered | 3 |
include_content, 6 to 10 pages delivered | 6 |
include_content, no pages delivered | 1 |
| Your org repeats the same request within about 5 minutes | 0 |
Billing follows pages delivered, not pages requested, and errors are never charged. Credits come in pay-as-you-go packs with no subscription. A plain search costs $0.80 per 1,000 on Starter, $0.65 on Standard and $0.50 on Scale, with 5, 25 and 50 concurrent requests. Prices checked 1 October 2026 on the pricing page. Full billing rules are in the Search API reference.
FAQ
Should I give the agent both tools? Yes, if your questions vary. Pass [search_tool, content_tool] and let the docstrings tell the model when full pages are worth 3 credits instead of 1.
Can I use ReActAgent instead of FunctionAgent? Yes. LlamaIndex's tools guide passes the same kind of FunctionTool list to ReActAgent(llm=llm, tools=tools). FunctionAgent is the simpler choice when your LLM supports function calling.
Do I need to handle GET versus POST? No. SerpexClient.search() builds the request and the Bearer auth header for you. If you're not in Python, the raw endpoint is documented in the Search API reference.
What if a search returns nothing? response.results is an empty list and the tool returns "No results found." An empty result that couldn't be confirmed as genuinely empty isn't charged.
Get an API key at app.serpex.dev and you'll have 200 free credits to wire this up, no card needed. The Python SDK docs cover the rest of the client.