Parallel vs Exa: We Put Their AI Search Products Through a Real Recruiting Brief. Only One Passed.
Type a brief in plain English, get a named list back. We tested four AI search tools from Parallel and Exa on a real recruiting brief and graded all 849 results. Here is what actually held up.

Deep Singh
Principal Talent Engineer & Co-Founder, Effi Flo
By Deep Singh, Principal Talent Engineer and Co-Founder, Effi Flo
In short: Type a brief in plain English, get a named list of people or companies back. That's the pitch from Parallel and Exa. What nobody tells you is how much of that list is actually right. So we ran four of their tools through a real recruiting brief and graded all 849 results. Exa won. And the accurate option turned out to be the cheap one.
Key Takeaways
- Exa took both top spots. Websets at 91.5% usable, Exa Agent at 85.7%. Parallel trailed, at 72.6% and 54.0%.
- Fast search breaks at requirement three. Perfect for simple asks. Add a filter it can't verify, and it starts guessing. We call it the depth cliff.
- A verified 25-name list costs about $1.50. Roughly 3.6 times cheaper than Parallel's checked mode, and more accurate. You pay in minutes, not dollars.
- Pedigree still needs a human skim. A founder's school was the hardest thing to verify in the whole study. Check school and past-employer claims by hand.
The pitch, and the catch
Type "machine learning engineers at fintech companies in New York," the way you'd say it to a colleague. Get back real people, real companies, profiles attached, in seconds. No Boolean strings. No fifteen open tabs.
It's called entity search. Google hands you web pages. This hands you entities: the actual people and companies that fit.
Parallel and Exa lead the space, and both are genuinely good. But the marketing never answers the one question that matters. When the list lands, how much of it holds up? Send a client a "targeted list" where a third of it misses the brief, and you didn't save time. You spent trust.
Nobody publishes that number. So we did.
Four tools, four temperaments
Two tools from Parallel, two from Exa. Same job, four ways of doing it. One sprints, one audits, one researches, one reasons.
Parallel Entity Search
The sprinter
Instant ranked search. Type a description, get a ranked list back in 1-5 seconds. The catch: results are ranked, not checked. The product's best guess, unverified.
- Speed
- seconds
- Price
- ~$0.005 per search
- Checks its work?
- No
Parallel FindAll
The auditor
Condition-checked matching. You list requirements explicitly; it evaluates every candidate against each one, with citations. Runs take minutes and are billed per matched result.
- Speed
- 2-9 minutes
- Price
- $2 + $0.15 per match
- Checks its work?
- Yes
Exa Websets
The researcher
Agentic verified list-building. It finds candidates, then dispatches automated research to verify each against each criterion. You are only charged for results that pass every check.
- Speed
- 1-11 minutes
- Price
- ~6 cents per verified result
- Checks its work?
- Yes
Exa Agent
The analyst
Autonomous research. A research agent plans its own searches, reasons across steps, and returns a structured, citation-backed answer. Built for multi-hop asks.
- Speed
- 1.5-3.5 minutes
- Price
- ~$0.87 per search
- Checks its work?
- Researches, but no per-item verification
How we stress-tested them
Real briefs aren't one line. They stack. "Space tech startups" is easy. "US space tech startups that raised $1M or more in the last two years, founders from a top school" is a real client brief. So we built two ladders, companies and candidates, and added one requirement per rung.
Company ladder
- 1Startups in space tech or rocket engineering
- +2...based in the US
- +3...that raised at least $1M in funding
- +4...within the last 2 years
- +5...with founders from a top US/Canadian school
Candidate ladder
- 1Machine learning engineers
- +2...in New York
- +3...currently at fintech companies
- +4...with 5+ years of experience
- +5...who previously worked at a big-tech company
Ten searches. Up to 25 results each. Then every result, all 849 of them, graded independently against every requirement, same strict standard for all four tools. No tool got to mark its own homework.
One rule for "usable": would it survive your desk check? Nothing invented, nothing clearly wrong. Out of what it handed you, how much could you actually send to a client?
Quick glossary. Entity search returns named people or companies, not web pages. Depth is how many filters you stack: "ML engineers" is depth 1; add location, employer, experience, and background and you're at depth 5. Usable rate is the share of results that pass an independent check of every filter. That's the whole ballgame.
The depth cliff
Read those two lines and you've basically read the study.
Fast search is brilliant, right up until the third requirement. One or two filters, Parallel Entity Search scored 88 to 100%, instantly, for half a cent. Add "at fintech companies" and it dropped to 8%. One more filter, and zero. It's not a bug. A public profile doesn't list an employer's funding stage, so an unchecked tool can't know it. So it guesses.
Verification is what holds the line. Exa Websets stayed at 88 to 91% at depth four, right where the fast tool hit zero. It gets there by researching each result before it shows you anything. Time is the price of the accuracy.
And at the hardest depth, honesty looks like a shorter list. On the toughest candidate search, Websets returned 13 names instead of 25. All 13 checked out. A tool that pads to 25 with confident wrong answers costs you more than one that admits the pool is small.
Exa Agent is the strong runner-up. Perfect through depth three on companies, 92% on the depth-four people search. Its one weakness shows at the end: at depth five it filled the list to 25, and about half didn't survive grading. Deep research, but no final per-item check before results reach you.
Every requirement you add is a test. Does your tool check, or guess? Guessers break at three.
Four very different animals
Accuracy, depth-resistance, speed, cost, recall, and self-checking. Score each tool across those six traits, and the shapes tell the story on their own.
Spikes on speed and price. Collapses on depth. No self-checking.
Decent everywhere, best nowhere, at the highest per-list price.
The balanced shape: accuracy, depth-resistance, honest self-checks.
Near-Websets accuracy with steadier speed, but no per-item verification.
Axes (10 = best): accuracy = average usable rate · depth 4-5 = accuracy on the hardest searches · speed and cost = faster and cheaper score higher · recall = how much of the requested list it fills · self-check = when the product says a result matches, how often the independent grader agrees.
You buy accuracy with minutes
Accuracy isn't the only thing that moves with difficulty. So does the wait.
The fast tool is flat, 3 to 6 seconds no matter what you ask, because it does the same work every time: rank, don't check. The verified tools earn their scores with time, and you can watch it happen. Websets runs under a minute on easy asks, then climbs toward several as each filter means more digging per name. Exa Agent holds steady around two to three minutes. The cliff never disappears. You just pay your way over it in minutes.
What it actually costs
| Product | Typical speed | Price / 25-name list | Usable rate | Cost / good result |
|---|---|---|---|---|
| Parallel Entity Search | ~4 seconds | ~$0.005 | 54% | <$0.001 |
| Parallel FindAll | ~3 minutes | ≈$5.50 | 72.6% | ≈$0.32 |
| Exa Websets | ~1 min (up to 11 on hard searches) | ≈$1.53 | 91.5% | ≈$0.07 |
| Exa Agent | ~2 minutes (steady) | ≈$0.87 | 85.7% | ≈$0.04 |
Here's the part that surprises people. The most accurate tool is also one of the cheapest. Websets bills only for names that pass every check, about 6 cents each, and nothing for the ones that fail. A full 25-name list runs about $1.53. Parallel's checked mode, FindAll, ran about $5.50 for a worse list. More accurate and cheaper, at the same time.
There's a quieter win in that model. When Websets can't verify enough people, it hands you a shorter list and charges less, instead of padding to 25 and billing you for guesses. The pricing is on your side.
So what do you actually use?
One or two filters, a quick market read? Fast search. Instant, pennies, plenty good for "who's even out there" before a kickoff call.
Three or more filters, a list going to a client? A tool that checks its work. For us, that was Exa Websets, clearly. Budget minutes, not seconds.
Anything about a school or a past employer? Skim it yourself. Pedigree is where results wobble: on the founder-school search, the hardest of the ten, the best any tool managed was 60%. Ten minutes of checking protects the whole list.
Either way, search is one slice of a bigger stack. We keep a tested list of the recruiting tools we'd actually put in front of a client, and we've written up how we'd build an AI recruiting stack around them.
Where we'd stay careful
Three honest notes before you take these numbers as gospel. "Unverified" isn't "wrong", it's unchecked, and on easy searches the fast tool was flawless. These are young tools that change monthly, so treat this as a July 2026 snapshot; we'll run it again. And we tested one kind of search per side, deep rather than broad. Your niche may behave differently. Which is exactly why we test before we build.
See every number behind the charts
Company searches - % of results that held up, by requirement depth
| Tool | 1 req | 2 reqs | 3 reqs | 4 reqs | 5 reqs |
|---|---|---|---|---|---|
| Parallel Entity Search | 100% | 88% | 92% | 56% | 0% |
| Parallel FindAll | 88% | 84% | 76% | 64% | 32% |
| Exa Websets | 100% | 100% | 96% | 88% | 60% |
| Exa Agent | 100% | 100% | 100% | 76% | 56% |
Candidate searches - % of results that held up, by requirement depth
| Tool | 1 req | 2 reqs | 3 reqs | 4 reqs | 5 reqs |
|---|---|---|---|---|---|
| Parallel Entity Search | 100% | 96% | 8% | 0% | 0% |
| Parallel FindAll | 92% | 100% | 68% | 72% | 50% |
| Exa Websets | 100% | 100% | 80% | 91% | 100% |
| Exa Agent | 100% | 100% | 76% | 92% | 57% |
Average time per search, seconds, by requirement depth
| Tool | 1 req | 2 reqs | 3 reqs | 4 reqs | 5 reqs |
|---|---|---|---|---|---|
| Parallel Entity Search | 3.5s | 4.5s | 5.7s | 3.2s | 3.6s |
| Parallel FindAll | 144s | 269s | 155s | 329s | 361s |
| Exa Websets | 57s | 44s | 149s | 168s | 520s |
| Exa Agent | 84s | 113s | 103s | 113s | 164s |
Product profile scores, 0 to 10 (10 = best)
| Tool | accuracy | depth 4-5 | speed | cost | recall | self-check |
|---|---|---|---|---|---|---|
| Parallel Entity Search | 5.4 | 1.4 | 10 | 10 | 9.3 | 0 |
| Parallel FindAll | 7.3 | 5.5 | 3.5 | 1.5 | 9.4 | 7.3 |
| Exa Websets | 9.2 | 8.5 | 5 | 5 | 9.4 | 9.2 |
| Exa Agent | 8.6 | 7 | 4.5 | 6 | 9.6 | 0 |
We log every run, and cache the expensive ones
None of this lived in a spreadsheet. Every search and every grade writes straight to Supabase from the tools themselves, over MCP, into structured tables: one row per result, per requirement, per verdict. That buys two things.
It keeps the study honest. Nothing gets overwritten, and any number here traces back to a stored row we can re-open.
And it lets us cache. Before a run fires an expensive verified search, it checks for a recent result for the same query and reuses it instead of paying to run it again. On a bake-off where each list costs real money, that's the difference between testing once and testing whenever we want. It's also how this becomes a living benchmark, not a one-off: re-run a slice, the new rows land next to the old ones, and the numbers update themselves.
How we ran it
Ten searches, five company and five candidate, at depths one through five, up to 25 results each. Every result graded against every requirement by the same independent automated evaluator, on a strict standard where no supporting evidence counts as not met. Pricing as published by the vendors, July 2026: Parallel Entity Search at $5 per 1,000 requests; FindAll at $2 plus $0.15 per match; Exa Websets at 10 credits per verified result on the $49 plan, roughly 6 cents each and about $1.53 for 25 names; Exa Agent usage-based on compute and search calls. Times are wall-clock, averaged across both ladders per depth.
Ready to automate your recruiting?
Book a 30-minute strategy session with Deep Singh to see how Effi Flo can transform your pipeline.
Book a callFrequently Asked Questions
Related Articles
Claude Code for Recruitment: How Staffing Agencies Build AI-Powered Workflows Without a Dev Team
How staffing agencies use Claude Code in production recruiting stacks. 5 workflows Effi Flo runs across 110+ agencies, where each breaks, and how Claude Code fits alongside Clay, n8n, Supabase, and your ATS.
Why Your ATS Is Lying to You: Stale Candidate Data, Fragmented Records, and the Hidden Problem Killing Agency Productivity
An ATS tracks applicants; it was never built to keep candidate data fresh or hold your relationships. Here's why that gap costs placements, and how modern teams fix it.
The 5-Layer AI Recruiting Stack for Staffing Agencies (2026)
After building automation systems for 110+ agencies, we've mapped the exact tools and architecture behind the staffing firms that are scaling without growing headcount. Here's the 5-layer stack that actually works.
