October 1, 2026

Why We Built GTMBench: A Benchmark for Prospecting

Buyers evaluate prospecting tools on database size and filter counts. Nothing in that measures what reaches the sales floor. GTMBench runs 26 natural-language prompts against four vendors to show.
Research
  • Button with overlapping square icons and text 'Copy link'.
Table of Contents

Major Takeaways

Why did GTM work need its own benchmark?
Software engineering has SWE-bench. Language models have MMLU. Go-to-market work has had nothing of the kind, so buyers evaluate list-building tools on database size, filter counts and coverage claims. None of those numbers predict what reaches the sales floor.
How does the GTMBench methodology work?
26 natural-language prompts were fixed before any vendor ran. Each system turned the prompt into a list of up to 1,000 rows. Every row was re-scraped from LinkedIn and judged against the same criteria by the same model. The run produces four measures: precision, qualified recall share, pool size and time to list.
What is the single most useful number the benchmark produces?
First-100 accept rate. That is the precision on the first 100 rows, the page a rep works in week one. Landbase scored 76.1% on first-100, against 40.6% for Clay, 29.2% for Apollo and 18.3% for ZoomInfo. A top-of-list that is already noisy stays noisy however long the file runs.

Software engineering has SWE-bench, developed at Princeton to measure how coding agents handle real GitHub issues. Language models have MMLU and a dozen successors. Go-to-market work has had nothing of the kind. Buyers evaluate list-building tools on database size, filter counts and coverage claims, and those numbers predict almost nothing about what reaches the sales floor.

The 2026 GTMBench was built on the SWE-bench model. 26 real list-building prompts, fixed before any vendor ran. One automated attempt per prompt. Every returned row re-scraped from LinkedIn and judged row by row against the same criteria by the same model. Salesforce State of Sales research reports that reps spend only 28% of a working week actually selling, with most of the remainder absorbed by list hygiene and re-verification. Harvard Business Review coverage of B2B selling documents the same pattern. GTMBench measures the metric that governs how much selling actually happens: how clean is the list that a vendor hands to a rep. The full report lives on the GTMBench landing page.

The problem with database size as a buyer signal

The headline number in most prospecting evaluations is catalog scale. Apollo publishes 240 million contacts. ZoomInfo aggregates business intelligence at similar scale. Clay reaches more than 200 data providers through its marketplace. Buyers see those numbers and infer coverage.

Coverage is not precision. GTMBench measured that gap. ZoomInfo's filters matched a geometric mean of 60,314 records per prompt, 46x Landbase's 1,304. ZoomInfo's precision came in at 23.2%. Landbase's at 76.7%. The pool is much larger. The yield after verification is much smaller. Expected true matches came in at 4,868 for ZoomInfo against 846 for Landbase, closer than the pool sizes suggest.

Catalog scale and verified yield are different measures. A benchmark that only reports catalog scale answers a different question than the one buyers care about.

What a prompt-level benchmark reveals

The benchmark used 26 prompts drawn from a corpus of 50,000 natural-language queries that Landbase operators had run on the platform. The prompts cover four families: firmographic (facts a vendor already stores, like industry and headcount), derived (facts a vendor would have to compute from raw data), lookalike (a seed set plus find more like these) and concept (loose labels like sticky brands or creator economy startups).

Four systems were in the set: Landbase, Clay, Apollo and ZoomInfo. The four cover the three ways list-building work gets done today. Apollo and ZoomInfo are catalogs with fixed filter dropdowns. Clay is a workflow builder that reaches further through multi-step pipelines. Landbase is an agent that takes the prompt in plain language and reasons over the data.

Each competitor was driven by a fresh coding agent reading only that vendor's own API documentation. That measured what a capable new user gets on the first attempt, not what a seasoned power user coaxes out of the platform after months of tuning.

Four measures per vendor

Precision is the share of returned rows that met every criterion. Landbase scored 76.7% precision across all 26 prompts. Clay 47.9%. Apollo 36.0%. ZoomInfo 23.2%. On the first 100 rows, Landbase scored 76.1%, Clay 40.6%, Apollo 29.2%, ZoomInfo 18.3%. The first-100 number matters more because that is the page a rep works in week one.

Qualified recall share measures how qualified rows split across the four vendors within a prompt. The share sums to 100% within a prompt, so above 25% is more than an even share. Landbase scored 47% across all 26. Clay 24%, Apollo 19%, ZoomInfo 10%.

Pool size captures reach: how many records the vendor's filters matched before the row cap. ZoomInfo's 60,314 and Apollo's 13,519 dominate on reach. Landbase's 1,304 trails. Expected true matches multiplies precision by pool, which captures the qualifying companies each vendor reaches. ZoomInfo 4,868, Apollo 3,764, Landbase 846, Clay 190-plus.

Time to list measures seconds from prompt to finished list. Clay 17, Apollo 20, ZoomInfo 20, Landbase 134. The agent takes seven times longer than any catalog. The extra time is where the precision gap comes from. Gartner research on sales technology adoption has argued that buyers trade latency for accuracy when the downstream cost of a bad list is high.

How rows get judged

Judging precision cleanly is where most industry comparisons fail. GTMBench built the judging layer on the SWE-bench pattern. Each vendor's list and the Landbase list were split into the rows they share and the rows only one of them found. The same model judged every segment: 100 rows per segment, scored against the original prompt criteria, re-scraped from LinkedIn at judgment time.

None of the judging time counted toward any vendor's time to list. Judging consistency across segments comes from using the same model with the same criteria rubric, instead of relying on human raters who vary across sessions.

The row-level judgment is where catalog claims come apart. A vendor can hold a company record that nominally meets the criterion. Re-scraping from LinkedIn at judgment time tests whether the record is current. A stale catalog loses rows at this step.

Who ran it

The Landbase AI Lab designed and ran the benchmark. The report is published for buyers and operators who want a measure of what each system returns under the same conditions. McKinsey research on sales productivity has argued that benchmarks created by platform operators carry obvious orientation, which is why GTMBench publishes the methodology, the prompts and the per-prompt scores.

Every row the four systems returned is listed in the full report. The judging rubric and the LinkedIn re-scrape pattern are documented. Future editions will add systems and expand the prompt set as new patterns surface in operator data.

What buyers can take from the benchmark

Three measures separate systems on dimensions that catalog claims do not predict. First-100 precision describes the top of the file. Precision-across-26 describes consistency. Expected true matches describes qualified reach after verification.

A buyer who optimizes for raw reach reads expected true matches and tolerates lower precision in exchange for pool size. A buyer who optimizes for SDR productivity reads first-100 precision and accepts a smaller pool for a cleaner top. The two choices lead to different stacks.

GTMBench does not resolve the choice. It makes the trade-off visible. For the vendor-by-vendor view, read our Landbase vs Clay, Landbase vs Apollo and Landbase vs ZoomInfo comparisons.

What Landbase delivers

Landbase is an agent that reads a plain-language prompt, reasons across more than 1,500 enrichment fields per company, verifies each row before returning it, and dial-tests the file before handoff. The platform runs inside Claude Code and Codex, and connects to Salesforce, HubSpot and CSV export. HubSpot sales statistics have documented how much of an SDR week depends on list quality upstream of the dial.

Across GTM teams from agencies to enterprise revenue organizations, customers have reported a 2 to 4x uplift in connect and meeting-booked rates. Send Landbase a list already pulled from Clay, ZoomInfo or Apollo and the platform will qualify, expand and score it in one pass at no cost. Start a qualification pass here.

Frequently asked questions

How was each vendor driven to produce the lists?

Each competitor was driven by a fresh coding agent reading only that vendor's own API documentation. That measures what a capable new user gets on the first attempt. Human operator tuning across many sessions would likely raise all four vendors' numbers; the benchmark isolates the out-of-the-box precision.

Why 26 prompts and not 100 or 1,000?

Each prompt required judging every returned row, up to 1,000 per vendor per prompt, by re-scraping from LinkedIn. 26 prompts across four vendors produced over 100,000 row-level judgments. The 26 were chosen to represent the four prompt families proportionally. The number is a cost-of-judgment constraint, not a measurement limit.

Why re-scrape from LinkedIn instead of trusting the vendor's own record?

Catalog freshness is one of the things the benchmark measures. A stored attribute that no longer matches the LinkedIn profile fails the criterion the prompt actually asks for. Judging from current LinkedIn state puts all four systems on the same reference data.

Will there be a 2027 edition?

Yes. The benchmark pattern is annual. Future editions will add systems (new entrants that emerge in the market) and expand the prompt set as new patterns surface in operator data.

Where can I see the full prompt-by-prompt results?

The full report is at landbase.com/gtmbench. The methodology, the 26 prompts, the per-prompt scores and the four summary measures are all documented.

Build a GTM-ready audience

Qualify your list in one pass

  • Button with overlapping square icons and text 'Copy link'.

Turn this list into a GTM-ready audience

Match this list to your ICP, prioritize accounts, and identify who to contact using live growth signals.

Run your list through Landbase

Send a list you have already pulled from Clay, ZoomInfo or Apollo. Landbase will qualify, expand and score it in one pass at no cost.

Stop managing tools. 
Start driving results.

See Agentic GTM in action.
Get started
Our blog

Lastest blog posts

Tool and strategies modern teams need to help their companies grow.

Research

Buyers evaluate prospecting tools on database size and filter counts. Nothing in that measures what reaches the sales floor. GTMBench runs 26 natural-language prompts against four vendors to show.

Daniel Saks
Chief Executive Officer
Research

Two prominent GTM platforms returned nothing on three prompts in the 2026 GTMBench. The common thread was derived criteria no catalog stores as a field.

Daniel Saks
Chief Executive Officer
Insight

Database size predicts reach. Precision predicts how many of those rows a rep can work. The 2026 GTMBench measured both.

Daniel Saks
Chief Executive Officer

How GTM teams turn this list into pipeline

See how GTM teams use fastest-growing lists to define TAM, prioritize accounts, and launch campaigns.