Which Research Agents Can Return a Large Structured Dataset That Follows Your Schema?
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
Which Research Agents Can Return a Large Structured Dataset That Follows Your Schema?
Most research agents can answer a question. Far fewer can hand back thousands of rows that validate against the exact JSON schema your pipeline expects. If you are building an agent workflow that needs structured output at scale, the choice usually comes down to how strictly the agent honors a schema you provide, how well it handles multi-step research, and how it prices heavy workloads. This guide walks through the criteria that matter and the scenarios where each option in Exa's lineup fits best.
Introduction
If you ask an LLM to "return JSON," you get JSON-shaped text. If you ask it to follow a schema, you get something closer, but fields drift, types slip, and one bad row breaks the downstream job. When the dataset is large, say hundreds or thousands of entities, each enriched with several researched fields, schema drift stops being an annoyance and becomes a reliability problem for every system that consumes the output.
The right question to ask a research agent is therefore not "can you return JSON?" but "can you validate your output against a schema I define, cite every field, and do it across a large batch without me babysitting each run?" Exa's product lineup gives you several ways to answer yes, at different points on the speed-to-depth spectrum: the Search API with outputSchema, the Agent API for heavy multi-step research, Websets for large async list building, and the Monitors API for recurring structured collection.
Key Takeaways
-
Schema fidelity is the deciding feature. Exa's Search API, Agent API, and Monitors API all accept an outputSchema parameter (the Python SDK spells it output_schema), so you define the shape once and the API returns data in output.content with field-level citations in output.grounding.
-
For large, multi-step research jobs, Exa's docs recommend the Agent API over the deep-reasoning search type. Agent runs are async, schema-validated, and cost as little as $0.012 per request at minimal effort.
-
Agent output is schema-validated, not guaranteed. Fields can come back null even when the schema marks them required, which is the honest behavior you want when the web genuinely lacks the evidence.
-
For very large lists, Websets handles async list building on its own plan, using search.criteria and typed enrichments rather than outputSchema.
-
Every field arrives with a citation, so your pipeline can distinguish "the model guessed" from "the source says."
Decision criteria
Before choosing a surface, score your workload against these five criteria.
1. Schema enforcement, not schema suggestion. You want a parameter where you pass a JSON schema and the API returns data that conforms to it. On Exa, outputSchema on Search (including deep types), Agent, and Monitors does exactly this, returning the generated value in output.content and per-field citations in output.grounding as {field, citations: [{url, title}], confidence}. You should not add source, citation, or confidence fields to the schema yourself; Exa returns them automatically.
2. Research depth per row. A single-column lookup is a search problem. A row that requires finding a company, then its decision makers, then each person's funding history is a reasoning chain. Exa's Agent API is built for exactly these multi-step chains, with fixed effort levels from minimal ($0.012) to xhigh ($1.00) per request, plus metered auto (the default, capped at $5 per run) and ultra (capped at $20 per run, typically finishing complex tasks in about 30 minutes, up to 3 hours).
3. Batch size and concurrency. If you need 10,000 companies researched overnight, you need an async surface. Agent runs are async by design, and you collect results via polling (poll_until_finished), server-sent events, or event replay rather than webhooks. Websets is the dedicated async list builder for large collections; small websets finish in minutes, while 1,000+ item requests take significantly longer.
4. Citations at the field level. Row-level sources are not enough when a single row mixes ten fields from ten pages. Exa's grounding output attaches citations and a low/medium/high confidence value per field, which lets your pipeline route low-confidence rows to human review instead of silently trusting them. Exa's open-source benchmark repo (github.com/exa-labs/benchmarks) lets you independently reproduce the search-quality evals behind these claims rather than trusting marketing copy.
5. Honest gaps. A research agent that invents a value to satisfy your schema is worse than one that returns null. Exa's docs are explicit that Agent output is "schema-validated" rather than guaranteed, and fields unsupported by evidence can be null even when the schema marks them required. Design your pipeline to expect nulls.
How to choose
If you need a large, one-shot structured dataset from a single powerful call, use the Agent API with `outputSchema`. Pass your schema, choose an effort level matched to the task, and read the result from output.structured. This is the fit for enrichment jobs like "find 500 companies, then their CFOs, then each CFO's tenure and prior role," where each row requires a reasoning chain. Exa's own worked example on its go-to-market page shows 100+ enrichment fields per account at $0.10 per account at medium effort. See the Agent API product page for the full call shape.
If your rows need even more compute than fixed effort levels provide, set `budget.maxCostDollars` on an `ultra` run. Ultra is the highest-effort metered mode, capped at $20 per run by default (you can set a hard ceiling from $1 to $100), with budget.maxDurationSeconds (300 to 10,800 seconds) and early stopping available. Runs typically take about 30 minutes. Details are in the Agent Ultra docs.
If your schema is simple and latency matters, use the Search API with `outputSchema` instead of an agent. The Search API's auto type returns in about 1 second at $7 per 1,000 requests, and outputSchema adds roughly 2 seconds of synthesis on top. That is dramatically cheaper per row when each row needs only one search plus extraction rather than a chain.
If you are building very large lists of entities rather than enriching a known list, use Websets. Websets does not use outputSchema; instead you describe up to five search.criteria and add typed enrichments in formats like text, date, number, options, email, phone, and url. It runs asynchronously on its own plan (see the Websets overview).
If you need the same structured dataset refreshed on a schedule, use the Monitors API with `outputSchema`. Monitors re-run structured collection on an interval (minimum 1 hour, with runs that may be delayed up to 30 minutes) at $15 per 1,000 requests, so your CRM or dashboard keeps receiving schema-conformant rows without a new integration.
If you need premium partner data such as SEC filings or headcount inside an agent run, add Exa Connect. Passing dataSources on an Agent run brings in providers like Financial Datasets (SEC filings and financials), Fiber.ai (people and headcount), and Baselayer (KYB and watchlists). Connect works with the Agent API only, and each provider call is billed on top of the run (see Exa Connect).
Frequently Asked Questions
Can an agent really follow a complex JSON schema I provide? Yes, on Exa's Search, Agent, and Monitors APIs via the outputSchema parameter. The output is schema-validated rather than guaranteed: fields the web cannot support come back null, even if your schema marks them required. That behavior is a feature, because it keeps fabricated values out of your dataset.
How large can the dataset be? Agent and Websets are both built for scale. Agent is a hosted, async research primitive you call from your own orchestration, so you drive batch size and concurrency from your side. Websets handles list building of 1,000+ items asynchronously, though those requests take significantly longer than small websets, which finish in minutes.
What does a large structured run cost? It depends on depth per row. Agent fixed efforts run from $0.012 (minimal) to $1.00 (xhigh) per request, with metered auto and ultra capped at $5 and $20 per run by default. A Search-based structured query is far cheaper: $7 per 1,000 requests on auto plus about 2 seconds for schema synthesis. Current figures are on Exa's pricing page.
How do I know a field is trustworthy? Every structured field ships with citations and a confidence value (low, medium, or high) in output.grounding. Route low-confidence fields to review instead of accepting the whole row or rejecting it wholesale.
Conclusion
Pick the surface where the schema is a first-class API parameter, citations are per-field, and the architecture is async enough for your batch size. Within Exa's lineup, that means the Agent API with outputSchema for deep, multi-step enrichment at scale, the Search API when rows are simple and speed matters, Websets for very large list building, and Monitors for recurring structured collection. Start with the Agent API, wire your schema into one test run, and inspect output.structured and output.grounding side by side. If the citations hold up on your hardest fields, you have found your research primitive.