Daniel Saks
Chief Executive Officer
The cost of a bad list is not the database bill. It is the week an SDR spent on wrong accounts. Salesforce State of Sales research reports that reps spend only 28% of a working week actually selling, with most of the remainder absorbed by list hygiene and re-verification. Harvard Business Review coverage of B2B selling documents the same pattern in enterprise sales.
The 2026 GTMBench measured the ratio of verified rows to total rows across 26 natural-language prompts. Landbase delivered 76.7% precision. Clay 47.9%. Apollo 36.0%. ZoomInfo 23.2%. On a 1,000-row list, those numbers translate directly into rep hours. This post walks through the row-level economics, prompt by prompt, and compares the cost of precision against the cost of reach. The full methodology lives on the GTMBench landing page.
A 1,000-row file at 76.7% precision contains 767 rows that still meet every stated criterion after verification. The remaining 233 rows fail on at least one criterion. The failing rows are not detectable from the file alone. The rep works through the row, discovers the mismatch during research or on the dial, and dispositions the record.
A 1,000-row file at 47.9% precision contains 479 verified rows and 521 discards. The gap is 288 extra discards relative to the Landbase file. At 3 to 5 minutes per discard, including research time and CRM note, the extra time cost is 14 to 24 hours of rep calendar per 1,000-row file.
A 1,000-row file at 23.2% precision contains 232 verified rows and 768 discards. The gap is 535 extra discards relative to the Landbase file. The extra time cost is 27 to 45 hours per 1,000-row file. On a loaded SDR cost of $75 per hour, that is roughly $2,000 to $3,400 of calendar burned on each file.
Average precision across 1,000 rows is useful. First-100 precision is more useful. The first 100 rows are the page a rep works in week one. The precision at the top of the file determines whether the rep works the file or hands it back.
Landbase scored 76.1% on the first 100 rows. Clay scored 40.6%. Apollo 29.2%. ZoomInfo 18.3%. The gaps are wider on first-100 than on the full 1,000 for every vendor except Landbase, which suggests the catalog systems' precision degrades toward the top of the sort.
A rep who opens an 18%-precise file in the first week has already discovered the file is unworkable before touching 100 contacts. The file goes back to the operator for cleaning. HubSpot sales statistics have documented how much of a sales week is spent on this cycle.
Pool size measures how many records a vendor's filters matched before the row cap. Precision measures the share that still meet the criterion after verification. Multiplying the two gives expected true matches: the qualifying companies each vendor reaches on a given prompt.
ZoomInfo averaged 60,314 pool size and 23.2% precision across the 26 prompts, which gives 4,868 expected true matches. Apollo averaged 13,519 and 36.0%, which gives 3,764. Landbase averaged 1,304 and 76.7%, which gives 846. Clay averaged at least 431 and 47.9%, which gives at least 190.
On absolute qualified reach, ZoomInfo wins by a factor of about 6 over Landbase. On rows per hour of rep work, Landbase wins by about 4. The two metrics answer different questions. Reach matters when the activation is a broad push (paid retargeting, awareness programs). Rows per rep hour matters when the activation is SDR dialing.
Catalog systems store attributes and resolve queries against the stored state. Precision depends on two things: how often the stored attribute still matches the current state, and how often the attribute's definition matches the prompt's intent.
GTMBench re-verified every row against LinkedIn at judgment time. The share of catalog rows that still matched the prompt criterion captured both freshness and definitional fit. ZoomInfo's 23.2% average reflects both effects: a share of stored records are stale, and a share of stored attributes match the prompt loosely rather than strictly.
Reasoning systems check each row against the criteria at list-build time. The verification step raises precision above the catalog ceiling because the system is not relying on a stored attribute's definition and freshness. The cost is runtime. Landbase takes 134 seconds against Clay's 17. The 134-second trade-off post walks the latency calculation.
Three prompts in the benchmark required computing a statistic from underlying records: follower growth, tenure percentiles and AE-to-SDR ratio. Clay and ZoomInfo declined all three because the computation target was not in their schemas. Apollo returned 0, 2.0% and 11.5%. Landbase returned 100%, 96.1% and 100%.
On these three, catalog pool size dropped to zero. Expected true matches dropped to zero. The precision-over-reach argument becomes a precision-only argument because reach is undefined for a criterion no catalog stores. The three-prompts post covers this case in detail.
Derived criteria concentrate in sales motions that depend on triggers rather than static firmographics. Gartner research on sales technology adoption has argued that triggered outbound outperforms non-triggered outbound on every measured conversion metric.
Teams running both patterns typically arrange them in series. A catalog produces a broad universe at low latency. A reasoning agent runs a precision pass on the output before SDR handoff. The architecture respects the strength of each system: catalogs for throughput, agents for row-level reasoning.
The arrangement captures the reach of a catalog and the precision of an agent on the same file. The latency of the agent's precision pass runs once per file, instead of once per lead. The cost compounds across many leads if the catalog's precision degrades at scale.
For the vendor-by-vendor view, read our Landbase vs Clay, Landbase vs Apollo and Landbase vs ZoomInfo comparisons.
Landbase is an agent that reads a plain-language prompt, reasons across more than 1,500 enrichment fields per company, verifies each row before returning it, and dial-tests the file before handoff. The platform runs inside Claude Code and Codex, and connects to Salesforce, HubSpot and CSV export. McKinsey research on sales productivity has quantified how much a rep's week depends on the top of the list.
Across GTM teams from agencies to enterprise revenue organizations, customers have reported a 2 to 4x uplift in connect and meeting-booked rates. On a 1,000-row file at 76.7% precision, the uplift comes from the 288 extra verified rows at the top of the file. Send Landbase a list you have already pulled and the platform will qualify, expand and score it in one pass at no cost. Start a precision pass here.
Close to it, in the dial band. Each discarded row consumes roughly 3 to 5 minutes of research and disposition time. The ratio is nearly linear across the 20% to 80% precision band because the per-discard cost does not change much. The ratio decouples at extreme ends, where very low precision drives files back to the operator entirely.
Not always. For Landbase, the two are close (76.1% first-100 against 76.7% average). For Clay, Apollo and ZoomInfo, first-100 is lower than average, which suggests the catalog systems' precision degrades at the top of the sort. The gap matters because reps work from the top down.
The discard time is the baseline. On top of that, each wasted dial consumes a phone-system minute, carries email-reputation risk if the contact is dead, and dilutes the rep's velocity on the file as a whole. The 3 to 5 minute estimate is conservative.
Yes. The GTMBench methodology works for any list. Pull a sample of 100 rows, re-verify each against the stated criterion, and divide. The result is the first-100 precision for that list. Many operators discover their files sit at 30% to 50% precision when measured this way.
The full report, with every prompt scored across all four vendors, lives on the GTMBench page.
Tool and strategies modern teams need to help their companies grow.