Enrich Company Records with a Web Search API
Most company records arrive half empty: a name, maybe an email domain, maybe a city. A web search API fills the gaps that a static database can't, because it reads the web as it is today. One query per missing field is usually enough to find a company's official website, its social profile URLs, recent news and funding announcements. The part that takes care is checking each answer against what you already know before you write it to the record.
This post covers which fields search is good for, the query patterns that work, how to verify a match, and what it costs per record. The examples use Serpex, but the method works with any search API that returns URLs and snippets.
What search can fill, and how to check it
| Field | Query pattern | How to confirm the match |
|---|---|---|
| Official website | "Acme Robotics" official website | The result's domain matches the email domain on the record, or the company name appears in the page title |
| Social profiles | site:x.com "Acme Robotics", site:github.com "Acme Robotics" | The official website links to the same profile |
| Recent news | "Acme Robotics" announces | The page names the company and its city or domain, not just a similar name |
| Funding mentions | "Acme Robotics" raises or "Acme Robotics" Series A | The amount and round appear in the page text, not only in the snippet |
| Hiring | site:acme-robotics.com careers | The page is on the company's own domain |
Two habits make this work. Quote the company name, so the search doesn't match "Acme" and "Robotics" separately. And once you know the official domain, use it in later queries, because a domain is far less ambiguous than a name.
Serpex has no date filter, so for news and funding put the year in the query ("Acme Robotics" raises 2026) and check the date on the page itself.
The request
A search is one POST. This is the cURL example from our Search API reference with the query changed:
curl -X POST https://api.serpex.dev/api/search \-H "Authorization: Bearer sk_your_api_key" \-H "Content-Type: application/json" \-d '{"q": "\"Acme Robotics\" official website"}'
For funding and news, where the snippet alone isn't proof, add "include_content": true and "content_results": 5. The API then fetches the top 5 result pages and returns their text as markdown in the same response.
The response
This is the response example from the same reference page, trimmed to the fields that matter here. Its query is a generic one, but the shape is what you'll parse:
{"id": "3f6c2a9e-8b1d-4c57-9f0a-2d4e6b8c1a73","query": "best JavaScript frameworks 2025","results": [{"title": "Top JavaScript Frameworks in 2025","url": "https://example.com/js-frameworks","snippet": "React, Vue, and Svelte remain the top choices for JavaScript developers...","position": 1,"content": "# Top JavaScript Frameworks in 2025\n\nReact, Vue, and Svelte remain the top choices for JavaScript developers heading into 2025..."},{"title": "2025 Frontend Framework Comparison","url": "https://example.com/frontend-comparison","snippet": "A deep dive into performance and DX across the major frameworks...","position": 2,"content_error": "blocked"}],"metadata": {"number_of_results": 10,"credits_used": 3,"from_cache": false,"status": "success","content_requested": 5,"content_delivered": 4}}
For enrichment, four fields matter:
urlis the candidate value for website and social fields, and the source you store for everything else.snippetis often enough to reject a wrong company quickly.contentis the page as markdown, present only when you asked for it and the fetch worked. That's where you confirm a funding amount or a date.content_errorreplacescontentwhen a page couldn't be fetched, for exampleblocked,timeoutor a page that robots.txt disallows. A result carries one or the other, never both.
A small enrichment step in Python
This finds the official website for a record and only accepts a result whose domain matches the email domain you already have. It uses the Python SDK.
import osfrom urllib.parse import urlparsefrom serpex import SerpexClientclient = SerpexClient(os.environ["SERPEX_API_KEY"])def host(url: str) -> str:h = urlparse(url).netloc.lower()return h[4:] if h.startswith("www.") else hdef find_website(company: str, email_domain: str):response = client.search({"q": f'"{company}" official website'})for r in response.results:h = host(r.url)if h == email_domain or h.endswith("." + email_domain):return {"website": f"https://{h}", "source": r.url}return Nonerecord = {"name": "Acme Robotics", "email_domain": "acme-robotics.com"}match = find_website(record["name"], record["email_domain"])
If nothing matches, the function returns None and the field stays empty. That's the right outcome. A wrong website is worse than a missing one, because everything you look up next is keyed on it.
For funding, run the search with include_content, then look for the company name and an amount in r.content rather than in the snippet. Keep the URL of the page you read as the source.
Rules that keep enriched data clean
- Store a source and a date with every field. The URL you read it from and the day you checked. When someone asks where a number came from, you can answer, and you know when it's due a recheck.
- Don't overwrite what a person typed. Write search results to a separate candidate field, and promote them only when they pass your checks or a person approves them.
- Match on two things, not one. Name plus domain, or name plus city. Plenty of companies share a name.
- Treat directories and news pages as leads, not answers. A directory page that mentions the company isn't its website. Follow it to the company's own domain.
- Leave blocked pages alone. Content fetching honours robots.txt, and a disallowed page comes back as
content_errorinstead of text. Don't route around it.
What it costs per record
Say you run four searches per record: website, social profiles, and news as plain searches, and funding with page content for the top 5 results. On Serpex that's 1 + 1 + 1 + 3 = 6 credits per record.
| Records | Credits | At $0.80 per 1,000 (Starter) | At $0.65 (Standard) | At $0.50 (Scale) |
|---|---|---|---|---|
| 1,000 | 6,000 | $4.80 | $3.90 | $3.00 |
| 10,000 | 60,000 | $48.00 | $39.00 | $30.00 |
| 100,000 | 600,000 | $480.00 | $390.00 | $300.00 |
A few billing rules bring the real number down. The content search is billed on pages actually delivered, and costs 1 credit if none come back. Errors are never charged. A search that returns nothing is free unless it's confirmed as genuinely empty. And repeating the same request from your organization within 5 minutes costs 0, so a retry inside that window doesn't double the bill. Prices checked on our pricing page on 3 October 2026, and new accounts get 200 free credits with no card.
When search is the wrong tool
Search is good at the long tail and at fresh facts: the startup no database has caught yet, last week's funding round, a new careers page. It's a poor way to fill the same fifty fields for a million companies. At that scale a company dataset is usually the better buy, and search is better used to check and update the records that matter most.
If what you need is the exact Google results page for a company name, with ads and rankings, that's a SERP API's job. SerpApi vs Serper vs DataForSEO compares the three main ones, and Search API Pricing in 2026 has the wider price table.
FAQ
How many searches does one record need?
Usually one per missing field. Start with the website, because once you have the domain every later query gets more precise.
Can I use site: in a query?
Yes. site:example.com restricts results to one domain, which is how you find a careers page or check a social profile handle.
Should I trust the snippet?
For rejecting a wrong match, yes. For a number such as a funding amount, read the page text with include_content and store the URL of the page you read.
Does this work from no-code tools?
Yes. Any tool that can send an HTTP POST with a header can call the API. There's a walkthrough for web search in n8n, and the same request works from other automation tools that can make HTTP calls.