I Ran Three Research Projects Through Tavily and Exa. Exa Won All Three.

I ran the same deep-research plan through Tavily and Exa three times, blind-judged and single-variable. Exa won all three — cheaper, faster, more complete.

8 min readBy Matthew Stublefield
A wooden sign pointing in opposite directions in a wooded area

Angle: Source-grounded proof-of-craft, grounded in Matthew's own head-to-head receipts. Spine: for the deep, multi-thread, entity-heavy research Fieldway actually does, Exa wins across the board — 3-for-3 in blind judging, 1.8-2.6x cheaper and faster every run, more complete first pass, denser entity extraction. Method exposed in full (single variable = provider; neutral curl fact-checker so no provider grades its own homework; blind position-swapped judging; reconciled against dashboard cost exports). Later beat: while the competition was running, Tavily published 'we rank #1 on SealQA/SimpleQA' — and they're not wrong, but their method dropped cost and speed entirely and tested single-fact snippet retrieval, not deep research. Both scoreboards are honest; only one measured my job. The Fieldway point: a vendor's headline benchmark tells you what THEY optimized for, not what YOUR work needs — so measure the job you actually do. Primary keyword: tavily vs exa SEO hint: RARE proven-demand keyword: 'tavily vs exa' / 'exa vs tavily' = 90/mo each at KD 0 (commercial/informational intent) — winnable head-on, unlike the null-volume buyer questions. Timely (react to Tavily's 2026-08 post while current). Counts as the 2nd (final) AI-commentary slot but is primary-receipt proof-of-craft, not pure reaction. Cross-audience (Camille research-rigor + Marcus research-that-de-risks-the-bet); tagged Camille. Service: managed-intelligence

Exa did the same research 1.8 to 2.6 times cheaper than Tavily. Every run. Never close.

I know that because I didn't read it off anyone's benchmark page. I ran it – three separate research projects, each one a real deliverable I needed for actual work, each one executed twice under conditions I held identical except for a single swapped part: the retrieval provider feeding the research agent underneath. Same plan, same agent prompts, same models doing the reading and the judging. One variable. Exa won all three.

I'm not a benchmarks lab, and I don't have a stake in a vendor race. I'm one consultant who needed to know which tool to trust with a client's budget and, more to the point, a client's answer. So I built the fairest test I could and let it tell me something I might not have wanted to hear.

The job I was actually testing

Most published comparisons of research tools measure a specific thing: you ask a question that has one correct answer, the tool fetches some snippets, and you check whether the answer is in there. That's a real job. It's just not my job.

The research I do for clients is deep, multi-threaded, and entity-heavy, with no answer key at the back of the book. Map a competitive landscape. Pressure-test whether a market is real before someone builds for it. Work out who the players are, who runs them, what they charge, and where the quiet gaps sit. That work runs a dozen threads at once, chases each one until it either pays off or dead-ends, and has to come back sourced well enough that I'd put my name on it. And it has to do all of that without spending a client's whole research budget in an afternoon.

So that's what I tested. Not "can it find a fact," but "can it do a month's worth of real research well, and what does that cost?"

How I ran it, so you can trust the result

A comparison is only worth as much as its fairness, so here's the whole method, including the parts that were a pain to build.

One variable. Each project ran through the exact same research plan – the same threads, the same agent prompts, the same Claude models in each role. The only thing that changed between the two runs was which provider's search sat underneath the agent. Anything else that drifted would have quietly invalidated the comparison, so nothing else was allowed to.

No provider grades its own homework. Both runs verified their citations through plain curl – a neutral fetcher that just pulls the page – rather than through either vendor's own extract or read API. If I'd let each provider check its own sources, I'd be measuring marketing, not retrieval.

Blind judging, positions swapped. Every qualitative call – which report had denser evidence, which pulled cleaner quotes, which left fewer gaps – was made on sanitized reports labeled only "System A" and "System B." The assignment was randomized per project and sealed before judging, and each head-to-head was run twice with the positions flipped, so a judge that leaned toward whatever it read first couldn't tip the result.

Nothing counted until it was on disk. Every run had to reconcile against its own files before I'd call it done, and every cost figure was checked against the provider's dashboard export – I pulled Tavily's actual billing CSV and corrected two of my own early estimates downward when they ran high, rather than let a flattering number stand. Exa's costs are exact from its fixed-price ledger.

None of that is exotic. It's just the discipline that separates "here's what I found" from "here's what I wanted to find."

What happened

Three projects. Same fairness rules each time. Exa took every one.

ProjectJudged tally (Exa / even / Tavily)Tavily costExa costCost gapMedian runFollow-ups (Tavily / Exa)
Apprenticeship pathways3 / 3 / 0$30.22$17.11Exa ~1.8× cheaper254s / 188s10 / 3
A B2B ICA profile4 / 2 / 0$31.45$15.15Exa ~2.1× cheaper260s / 160s8 / 5
A competitive landscape5 / 0 / 0$44.80$17.50Exa ~2.6× cheaper216s / 183s7 / 4

Read down the columns and the same pattern repeats. Exa was cheaper on every run, and never by a little. It was faster on every run. And it needed far fewer follow-up passes to fill the gaps in its first sweep – three where Tavily needed ten, on the first project – which means its opening pass simply left less undone.

The texture of the evidence was better, too. Exa pulled more named entities, more verbatim quotes and prices, and it attached people to the firms they run – pinning a named principal to the boutique consultancy she founded, the kind of connection a bare list of ten snippets can't make. When the judges read Exa's longer reports, they generally found the extra length earned its place. When they read Tavily's extra volume, they found it more of a mixed bag.

I want to be precise about where Tavily won, because it did win, consistently, on a few axes. It found more unique domains every time – broader sourcing, less concentration in the top few sites. Its cited links resolved more reliably in the run where it mattered most: on the competitive landscape, only 1.4% of Tavily's links had rotted versus 8.6% of Exa's. (On another project the genuine dead-link rate flipped the other way, so treat that one as task-dependent rather than a law.) And it brought more raw sources per finding – better corroboration density, even when that density converted into fewer distinct, named results.

That's a real edge, and it's the edge you'd want if your job were verification – checking a claim against as many independent sources as possible. It's just not the edge that wins deep synthesis on a budget. For my job, Exa was cheaper and better, and Tavily's strengths were the runner-up's strengths.

And then Tavily told me they win

Here's the part that makes this really interesting. While I was running these competitions, Tavily published a post titled "Tavily ranks #1 on SealQA and SimpleQA" – benchmarks where they beat Exa and several others, with strong numbers: 97.6% on SimpleQA Verified, 55.7% on the hard split of SealQA.

They're not wrong. And I still disagree with them. Both of those can be true, because we measured different things.

Their test queried each provider's raw Search API for ten results, passed the snippets to a reader model, and graded the answer against a known correct one. They kept it deliberately short and simple – their words – to measure pure retrieval, not a research harness wrapped around it. That's a legitimate benchmark. It's an honest measurement of the snippet-retrieval layer, and Tavily may genuinely be the best in the field at it.

But look at what that test leaves out. It measures accuracy on single-fact questions that have a gold answer. My work has neither – it's open-ended synthesis with no answer key. It scores the raw search API; I was scoring a full research agent wrapped around that API, which is where Exa's lead compounds through fewer follow-ups and richer extraction. And the two columns that decided my competitions – cost and speed – don't appear anywhere in their post. A benchmark that reports zero dollars and zero latency has quietly removed the two axes where Exa beat them 1.8 to 2.6 times, every run.

None of that is a gotcha. Their method is disclosed, not hidden. And the axes where they won my runs – source diversity, links that resolve – are exactly what a single-fact accuracy benchmark rewards, so it's coherent that they'd top it. Their own post even closes with the right advice: treat any benchmark as one signal, and run your own evaluations on queries from your real work. That's precisely what I did. Three times.

Both scoreboards are honest. Only one measured my job.

A vendor's headline benchmark tells you what that vendor chose to optimize for. It does not tell you what your work needs. Those are different questions, and the gap between them is where a confident, well-produced number can walk you straight into the wrong tool.

Tavily benchmarked the retrieval floor and may well own it. I benchmarked the deep-research floor – multi-thread synthesis, sourced, on a budget – and Exa won it three times out of three. Neither of us is lying. We're measuring different floors of the same building, and the only floor that matters to you is the one you actually work on.

That's the part I'd hand to anyone choosing research tooling for real stakes – a market you're about to commit engineering to, a competitor you're about to reposition against. The answer to "which is better" is another question: better at what, measured how, and priced how. When the cost of being wrong is a client's money or a company's direction, you don't get to borrow someone else's benchmark. You run the test that looks like your own work, and you keep the receipts.

Exa won my three. Whether it wins yours is a question only your own work can settle – no benchmark, mine included, gets to answer it for you. When the cost of being wrong lands on someone who trusted you, a number borrowed from a vendor's homepage isn't an answer you can stand behind. A test you ran yourself is.

Want help running a sharper practice?

The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.

See how this works for advisors