Skip to main content

Web Search

Search the live web and get clean, deduplicated, LLM-ready results โ€” titles, URLs and snippets, ranked by how many independent search engines agree on each source.

Powered by our own self-hosted metasearch cluster. No third-party search API sits in the middle, so results are fast, cheap, and not subject to someone else's pricing page.

Endpointsโ€‹

GET/search?q=next.js+16+release
Web search โ€” up to 30 ranked results.
GET/search?q=what+is+mcp&deep=1
Research mode โ€” also reads the top pages and returns their extracted text.
GET/search?q=india+ai+news&news=1
Fresh results only โ€” last 24 hours.
GET/v1/search?q=...
Versioned alias โ€” same behaviour.

Query Parametersโ€‹

ParameterTypeDefaultDescription
qstringโ€”โœ… Required. The query (2โ€“400 characters)
limitnumber10How many results to return (1โ€“30)
deep0 | 10Research mode: also read the top pages and return their text
pagesnumber4How many pages to read in deep mode (1โ€“8)
news0 | 10Only results from the last 24 hours

Responseโ€‹

{
"success": true,
"query": "next.js 16 release",
"count": 10,
"took_ms": 1240,
"deep": false,
"sources_read": 0,
"cached": false,
"results": [
{
"title": "Next.js 16",
"url": "https://nextjs.org/blog/next-16",
"snippet": "Next.js 16 is here with Turbopack by default...",
"engines": ["duckduckgo", "bing"],
"agreement": 2
}
]
}

In deep mode each result also carries a content field with the readable body of the page (up to ~5,000 characters).

agreement is a trust signal

agreement counts how many independent engines returned that URL. Most results come from a single engine (agreement: 1); roughly 1 in 5 is surfaced by two or more, which is a decent relevance signal. Treat it as a ranking hint, not a guarantee โ€” always rank by your own relevance judgement too.

Feed it straight to an LLMโ€‹

This is the pattern the service is built for: search, read, then answer with citations.

import requests

r = requests.get(
"https://api.apimitra.in/search",
params={"q": "latest ruby release", "deep": 1, "pages": 3, "limit": 6},
headers={"x-api-key": "YOUR_KEY"},
)

data = r.json()
context = "\n\n".join(
f"[{i+1}] {x['title']}\nURL: {x['url']}\n{x.get('content', x['snippet'])}"
for i, x in enumerate(data["results"])
)

# Send `context` as system context to your model and it can cite [1], [2], ...

Examplesโ€‹

# Basic search
curl "https://api.apimitra.in/search?q=best+ai+chat+apps&limit=5" \
-H "x-api-key: YOUR_KEY"

# Research mode โ€” returns page text too
curl "https://api.apimitra.in/search?q=what+is+mcp&deep=1&pages=3" \
-H "x-api-key: YOUR_KEY"

# Fresh news only
curl "https://api.apimitra.in/search?q=india+ai+funding&news=1" \
-H "x-api-key: YOUR_KEY"

Notesโ€‹

  • Caching โ€” identical queries are cached for 15 minutes; repeats return in under 1 ms and do not touch the upstream engines.
  • Deduplication โ€” the same URL from multiple engines is merged into one result, with engines[] listing where it came from.
  • Noise filtering โ€” known low-value domains are dropped before ranking.
  • Safety โ€” q is validated and capped at 400 characters; results are returned as data only, never rendered as HTML.