The 2026 GTMBench

GTMBench is the prospecting benchmark evaluating the efficacy of natural language prompts. Landbase, Clay, Apollo and ZoomInfo were given the same 26 list-building prompts. Every row the four systems returned was scored by one judge against the same criteria.

Scoreboard

The full run: Landbase leads on quality, ZoomInfo on reach

Every measure from the benchmark sits in this table, six metrics across all 26 prompts and all four systems, with the median under each mean. Landbase leads precision at 76.7%, first-100 accept at 76.1% and qualified recall share at 47%, and ZoomInfo matches 60,314 companies for 4,868 expected matches by accepting a far lower precision.

Judged precision and the reach and speed each system traded for it. Click any column to sort.
System Precision % First-100 % Recall share % Pool size Expected matches Time to list Prompts won
Table 01

Agents lead three of the four prompt families

A filter can only match what the vendor already stores, so Clay leads firmographic prompts, where the prompt names those stored attributes outright. An agent reads the underlying records and works the criterion out for itself, so Landbase leads derived prompts at 75.6% against 18.2%.

Mean precision by family, %. Each of the 26 prompts sits in one family, grouped by what the request asks the system to find.
Chart 02

Clay and ZoomInfo cannot express three of the 26 prompts

Neither product has an attribute for follower growth, tenure percentiles or team ratios, so both builders returned an empty filter set and scored 0. Apollo answered all three between 0.0% and 11.5%, and Landbase scored 100%, 96.1% and 100%. A catalog answers only what its schema already stores, so reading the underlying records is the only route to a criterion with no field behind it.

Precision on the three prompts Clay and ZoomInfo had no attribute for.
Chart 01

Landbase returns the most precise list on 18 of 26 prompts

Landbase clears 95% on eleven of the 26 prompts, and Clay wins four. The eleven carry a criterion no filter can express, and Clay's four are narrow, well-bounded populations where a dropdown is enough.

Judged precision for all four systems on each prompt. Declined marks a prompt a system could not express at all, which scores 0.
# Prompt LandbaseClayApolloZoomInfo Winner
Method

How a prompt becomes a judged list

The 26 prompts are drawn from over 50,000 natural-language prompts run on the Landbase platform, grouped into four families and chosen to cover the range of requests a GTM team makes.

Same inputs

The 26 prompts were fixed before any system ran. Each list was capped at 1,000 rows, and at 200 rows for ZoomInfo person prompts. Every returned row was re-scraped from LinkedIn, so each judgment runs against fresh reference data rather than the vendor's own record.

Same criteria, same judge

Qualification criteria were generated once per prompt and applied verbatim to every system. One judge model scored up to 100 rows per segment, one criterion at a time, and a row counts only when every criterion holds. Rows two lists shared were judged once and credited to both.

How each system was driven

Each competitor's filters were written by a fresh Claude Code session given only that vendor's public API documentation and field reference, with no corrective turns. That measures what a capable new user gets from each product on the first attempt. Landbase ran its own agent.

How the AI judge works

The benchmark uses an evidence-based AI qualification workflow, powered by OpenAI models and web search tools, to assess how well audience records satisfy the requested targeting criteria. For each criterion, the judge follows three steps.

Step 1 / Plan the evidence

The judge identifies potential sources and signals that could support an assessment, such as company websites, public professional profiles and relevant databases. It also estimates whether sufficient evidence is likely to be available.

Step 2 / Validate the plan

The workflow checks the proposed sources for feasibility and consistency, flagging sources whose existence is uncertain. This helps identify criteria that may be difficult to assess reliably from available information.

Step 3 / Evaluate each record

A tool-enabled agent retrieves evidence according to the plan, combines it with available offline context, and assesses each audience record against each criterion.

Bring your own listand see where it lands

Send a list you pulled from your current tool. Landbase scores it for precision against the criteria you actually asked for, shows what it can add, and returns the result in one pass.