跳转到内容

Exa agent search API

Primary sources: Homepage · Introducing Exa Agent (2026-06-16) · Agent guide · Search guide (coding agents) · Create a run · Search OpenAPI · Pricing · MCP · OpenAI SDK compat · llms.txt

Related vault: 2026-09-26 Cloud agent orchestrators in the wild · deep-research

Pin checked 2026-09-28 (Shanghai / UTC+8).


Positioning from docs/homepage: “A powerful web search tool designed for agents” / “web search built for AI agents” — index + retrieval + contents optimized for LLM context, not for human SERPs.

SurfaceEndpoint(s)Role
Search APIPOST https://api.exa.ai/searchNatural-language search → ranked results + optional highlights/text/summary; optional outputSchema synthesis
Contents APIPOST /contentsSame content options when you already have URLs (no search)
Deep Search/search with type ∈ deep-lite / deep / deep-reasoningMulti-step retrieve + synthesize on the Search path (seconds, not minutes)
Answer/answer (also /chat/completions model exa)LLM answer grounded on Exa search
Agent APIPOST /agent/runs (+ get/list/cancel/stop/events)Async context agent: fans out searches, reads, subagents, verification, Connect partners → one grounded structured result
Exa ConnectdataSources on Agent runsPremium partner DBs (Fiber, Similarweb, Baselayer, …) inside the Agent loop
Monitors / Websets / Batchseparate APIsRecurring search; verified datasets; async request batches — out of scope for this digest beyond naming

Auth (all products): Authorization: Bearer $EXA_API_KEY or x-api-key: $EXA_API_KEY. Keys: dashboard.exa.ai/api-keys. OpenAPI source of truth: exa-spec.yaml.

SDKs: pip install exa-py · npm install exa-js (exa-labs/exa-py, exa-labs/exa-js).

Enterprise claims on homepage (sourced): SOC 2 Type II, GDPR, CCPA, HIPAA (BAA available), Zero Data Retention (“queries and results are never stored or trained on” — ZDR is Enterprise, per-team; product matrix differs — see Gotchas).


Terminal window
POST https://api.exa.ai/search
Authorization: Bearer $EXA_API_KEY
Content-Type: application/json

Required: query (natural language). Default returns up to 10 results; numResults 1–100 (no pagination). No search without a query.

From Search guide + OpenAPI + Pricing:

typeTypical latencyBase price (≤10 results)Use when
auto (default)~1 s$7 / 1kDefault balance
fast~450 ms$7 / 1kLatency-sensitive UI
instant~250 ms$7 / 1kReal-time (chat/voice/autocomplete)
deep-lite~4 s$12 / 1kLightweight research + synthesis
deep4–15 s$12 / 1kMulti-step + structured outputs
deep-reasoning12–40 s$15 / 1kMax reasoning on Search path

Docs explicitly say: for long-running research / list building / multi-hop enrichment, prefer Exa Agent over deep-reasoning.

Additional results above 10: 1/1kresults∗∗.AIpagesummaries:∗∗1 / 1k results**. AI page summaries: **1 / 1k pages.

ParamNotes
numResults1–100; default 10
categorycompany, publication, news, personal site, financial report, people (+ string hints). company / people do not support startPublishedDate, endPublishedDate, excludeDomains (400 if used)
includeDomains / excludeDomainsHost, path prefix (anthropic.com/news), or *.substack.com. Prefer filters over site: in query. Max 1200 entries each
startPublishedDate / endPublishedDateISO 8601 publication window (hard filter)
userLocationISO country code (e.g. US)
contentsNested: highlights, text, summary, maxAgeHours, subpages, snapshotAsOf, extras…
outputSchemaRoot type: "text" or "object" → response output.content + output.grounding. Object schemas: ≤2 nesting levels, ≤10 properties. Adds ~2 s synthesis latency. Works with every type
systemPromptBehavior / source prefs for synthesis (not the response shape)
streamSSE only when outputSchema is set; else normal JSON
additionalQueriesDeep-search variants only (1–10); expands research queries

Recommended default: contents: { "highlights": true } — query-relevant excerpts, token-efficient.

  • highlights — excerpts sized to relevance (recommended for agents).
  • text — clean page body; often { "maxCharacters": N }.
  • summary — LLM summary per result (extra $ per page).

Pick one content view unless you need both — each view is billed separately. /search nests options under contents; /contents puts the same fields at top level next to urls.

Freshness (contents.maxAgeHours): omit = cache-with-fallback; positive = max cache age then refetch; 0 = always live fetch; -1 = cache only. This is content freshness, not publication-date filtering. Max 720 hours. Prefer this over deprecated livecrawl.

Terminal window
curl -s -X POST "https://api.exa.ai/search" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $EXA_API_KEY" \
-d '{
"query": "recent techniques for improving retrieval in RAG systems",
"type": "auto",
"contents": { "highlights": true }
}'

Introduced 2026-06-16: single API for frontier web research at lower cost via model fusion + Exa highlights (blog claims up to 94% token reduction from highlights). Strengths: deep research, list-building, entity enrichment; can split large jobs into parallel subagents.

POST /agent/runs → create (returns immediately unless SSE)
GET /agent/runs/{id} → poll
GET /agent/runs → list
POST …/cancel | …/stop → cancel / graceful stop
GET /agent/runs/{id}/events → replay (non-ZDR); SSE with Accept

Statuses: queued → running → completed | failed | cancelled.

Create with Accept: text/event-stream (or SDK stream: true) for SSE: agent_run.created / .started / .completed / .failed / .cancelled. Treat live agent_run.source.added as preview; authoritative citations are terminal output.grounding.

Completed output:

  • output.text — prose
  • output.structured — JSON when outputSchema set
  • output.grounding — field-level citations (+ confidence)
  • usage / costDollars — ACUs, searches, emails, phones, Connect
FieldRole
queryRequired. Task spec: entity, count, criteria, evidence bar
systemPromptExtra behavior / judging rules
effortminimal | low | medium | high | xhigh | auto (default) | ultra
outputSchemaJSON Schema → output.structured (draft-07 / 2019-09 / 2020-12 via $schema)
input.dataRows/entities to enrich
input.exclusionRecords/entities to skip
previousRunIdContinue from a completed prior run (new ID each time)
dataSourcesUp to 5 Connect providers (fiber, similarweb, baselayer, affiliate, particle, jinko, polymarket, macrobond, financial_datasets)
budget.maxCostDollarsCap for auto/ultra only (1–1–100; defaults 5/5 / 20)
budget.maxDurationSecondsSoft wall-clock for ultra only (300–10800); stopReason: time_limit_reached
metadataCaller key-value strings

stopReason examples: schema_satisfied, budget_reached, time_limit_reached, stopped, error, cancelled.

From blog + Agent guide + Pricing — numbers match as of pin date:

EffortPriceBest for
minimal$0.012 / requestLowest-cost lookups, short answers
low$0.025 / requestSimple lookups, narrow facts
medium$0.10 / requestDefault starting point for standard research
high$0.50 / requestHarder research, more citations
xhigh$1.00 / requestCompleteness ≫ cost/latency (fixed)
autoMetered; default cap $5Variable scope / list building
ultraMetered; default cap $20Large lists, exhaustive research

Metered usage rates (also apply under auto/ultra caps):

ComponentPrice
Agent Compute Unit (ACU)$0.10 / ACU
Search tool call$0.005 / search
Email enrichment$0.02 / email
Phone enrichment$0.07 / phone

Connect is additive (examples from Connect overview: Fiber 0.02/credit,Similarweb0.02/credit, Similarweb 0.30/credit, Baselayer 0.10–0.10–4.00/order, etc.).

Free tier (Search/etc.): **10∗∗creditsonsignup,resetsto10** credits on signup, resets to 10 monthly; + one-time $10 onboarding bonus (Pricing).

Still accurate per OpenAI SDK Compatibility:

base_url = https://api.exa.ai
model = "exa-agent" # → Agent API
"exa" # → /answer via /chat/completions
ModeHowBehavior
SyncdefaultBlocks until complete
Streamstream: trueOpenAI Responses SSE → response.completed
Backgroundbackground: trueImmediate in_progress; poll GET /responses/{id}

Map effort via reasoning.effort (same enum). high / xhigh / ultra cannot be synchronous (400) — use stream or background. No budget field on /responses; ultra uses default cap. Continue with previous_response_id. Cancel: POST /responses/{id}/cancel.

Terminal window
curl -s -X POST "https://api.exa.ai/agent/runs" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $EXA_API_KEY" \
-d '{
"query": "Find engineering leaders at AI infrastructure companies that raised a Series A or B in the last 6 months.",
"effort": "auto",
"outputSchema": {
"type": "object",
"required": ["people"],
"properties": {
"people": {
"type": "array",
"maxItems": 10,
"items": {
"type": "object",
"required": ["name", "job_title", "linkedin_url"],
"properties": {
"name": { "type": "string" },
"job_title": { "type": "string" },
"linkedin_url": { "type": "string", "format": "uri" }
}
}
}
}
}
}'

URL: https://mcp.exa.ai/mcp (docs, exa-labs/exa-mcp-server).

ToolDefaultPurpose
web_search_exaonOrdinary web search + ready content
web_fetch_exaonContents for known URLs
web_search_advanced_exaopt-inFilters, dates, highlights, freshness, subpages
agent_runon with OAuth/API keyMulti-step Agent (not on free keyless limits)

Auth modes: keyless (rate-limited), OAuth (?login), API key (x-api-key). Cursor example: { "mcpServers": { "exa": { "url": "https://mcp.exa.ai/mcp" } } }. Local: npx -y exa-mcp-server with EXA_API_KEY.

agent_run long jobs: tool may return status: "running" + id; client re-calls with runId. Follow-ups on completed work use previousRunId (not the same as waiting on runId).

Terminal window
npx skills add exa-labs/agent-skills

Repo: exa-labs/agent-skills — build-with-exa, exa-search, exa-contents, company-research, lead-generation. Also installable as Exa plugin in ChatGPT/Codex, Claude (claude plugin install exa@claude-plugins-official), Grok Build marketplace.

Attach partners on Agent runs via dataSources: [{ "provider": "…" }]. Agent picks partner tools when outputSchema field descriptions name the source (e.g. “from Similarweb”). Index search remains available on every run; Connect is premium overlay. Unavailable under ZDR (400 if requested).


NeedStart with
Ranked pages + highlights for your LLM / tool loopSearch auto / fast / instant
One-shot synthesis / small structured JSON in secondsDeep Search (deep-lite / deep) + outputSchema
Known URLs onlyContents
Async list building, multi-hop, enrichment over rows, verificationAgent
Premium B2B / traffic / KYB / markets data fused with webAgent + Connect
Coding agent in Cursor/Claude/Codex without custom HTTPMCP (web_search_exa / agent_run)

Rule of thumb from Agent best practices: Search when you need pages quickly and your app does the reasoning; Agent when the product would otherwise be a loop of search → read → verify → enrich.


Search

  • Put content options under contents on /search (not top-level text/highlights).
  • Do not copy legacy params: useAutoprompt, includeUrls/excludeUrls, livecrawl, crawl-date filters, includeText/excludeText, context — see Search best practices → Common mistakes.
  • maxAgeHours ≠ publication date; use date filters for “newer articles,” freshness for “live page content.”
  • company/people categories: unsupported filters → 400.
  • Combining highlights+text+summary bills each view.
  • Python SDK: snake_case everywhere (max_age_hours, output_schema, …).
  • Search outputSchema object limits: 2 levels / 10 properties; citations live in output.grounding, not your schema.
  • Prefer Agent over stacking deep-reasoning for list-building.

Agent

  • Create response ≠ final answer — persist id, poll or stream.
  • Schema adherence validates shape, not truth; fields may be null despite required; stopReason: schema_satisfied allows nulls.
  • Bound arrays with maxItems for cost predictability (esp. contact formats email / phone).
  • Keep rows in input.data, exclusions in input.exclusion, shape in outputSchema — don’t paste tables into query.
  • Verification schemas should include uncertainty enums (cannot_verify vs absent).
  • Fixed efforts = predictable ;‘auto‘/‘ultra‘without‘budget.maxCostDollars‘canrunto; `auto`/`ultra` without `budget.maxCostDollars` can run to 5/$20 defaults.
  • Event replay / previousRunId / Connect blocked under ZDR.

ZDR / compliance (ZDR docs)

ProductZDR
Search, Contents, AgentAvailable (Enterprise, per team)
Answer, WebsetsNot currently supported

Agent ZDR: stream or poll within ~10 minutes after terminal status; then irretrievable. Homepage “never stored” marketing applies in the ZDR/Enterprise framing — enable via sales (sales@exa.ai), don’t assume default.

MCP

  • Explicit ?tools= replaces defaults — include every tool you want.
  • agent_run needs OAuth or API key (not keyless).

  • Did not call the live API (no key exercised); pricing copied from official Pricing + Agent guide + Jun 2026 blog — all three agreed on fixed effort $.
  • Rate limits / concurrency numbers deferred to Billing / Agent limits (not fully pulled).
  • Full Connect per-operation price tables only sampled from overview; Fiber/Similarweb credit math is deeper in partner pages.
  • Absolute Search guide URL …/search-api-guide also exists; coding-agents guide was the requested primary and is current.
  • No screenshots embedded (API/reference tables suffice; Attachments unused).