October 1, 2026

The 3 Prompts Clay and ZoomInfo Could Not Express

Two prominent GTM platforms returned nothing on three prompts in the 2026 GTMBench. The common thread was derived criteria no catalog stores as a field.
Research
  • Button with overlapping square icons and text 'Copy link'.
Table of Contents

Major Takeaways

Which three prompts broke Clay and ZoomInfo in the benchmark?
The three prompts were LinkedIn followers grew more than 20% last quarter, 90th-percentile engineering tenure under 18 months, and AE-to-SDR ratio between 1.5 and 3. Each requires computing a statistic from underlying records that neither Clay nor ZoomInfo stores as a field.
What did Apollo and Landbase score on the same three prompts?
Apollo scored 0, 2.0% and 11.5%. Landbase scored 100%, 96.1% and 100%. Apollo's AI research layer attempted the computation. Landbase's agent read the raw LinkedIn signals and returned only the companies where the computed value cleared the threshold.
What makes a criterion derived in the GTMBench sense?
A derived criterion requires computing a statistic from underlying records, instead of matching a stored attribute directly. Examples include trends over time (follower growth), distributional measures (tenure percentiles), and ratios across people within a company (AE-to-SDR). Catalogs can match stored attributes. Only reasoning systems can compute derived ones.

The 2026 GTMBench ran 26 natural-language list-building prompts against Landbase, Clay, Apollo and ZoomInfo. Three of the 26 broke two of the four systems entirely. Clay and ZoomInfo returned empty filter sets on those three. Apollo returned near-zero. Landbase cleared 96% on all three.

The common thread was a criterion that no catalog stores as a field. Forrester research on B2B data readiness has argued that enterprise buyers increasingly ask questions their vendor catalogs cannot answer, especially when the question involves a time trend or a ratio across people inside a company. Harvard Business Review coverage of B2B selling has documented how precision at the top of a list determines whether reps work the file or hand it back. This post walks through the three prompts and what they reveal about catalog limits. The full GTMBench report lives on the GTMBench page.

Prompt one: LinkedIn followers grew more than 20% last quarter

The prompt asked for companies whose LinkedIn page gained more than 20% in follower count during the previous quarter. The criterion is a trend over time, computed by comparing two counts at two moments. It is not a stored attribute in any of the vendor schemas.

Clay declined. ZoomInfo declined. Apollo returned 0% precision. Landbase scored 100%.

Landbase computed the follower count at the end of each quarter from the raw LinkedIn record, differenced the two values, and returned only companies where the ratio cleared 20%. The derivation happens per company at list-build time. A catalog that stores a single current follower count has no way to answer a prompt about change.

Trend criteria appear routinely in sales motions that depend on momentum. Rising follower counts correlate with hiring, funding and product-launch activity. LinkedIn research on sales trends has catalogued how buyer organizations grow attention and headcount in parallel. A prospecting system that can derive the trend is answering a different question than a system that can only read the current state.

Prompt two: 90th-percentile engineering tenure under 18 months

The prompt asked for companies where the 90th-percentile tenure of engineering employees is under 18 months. The criterion is a distributional measure, computed by ranking tenure values across the engineering function at a company and reading the value at the 90th percentile.

Clay declined. ZoomInfo declined. Apollo returned 2.0%. Landbase scored 96.1%.

Landbase enumerated the engineering roster per company from LinkedIn, computed the tenure distribution, read the 90th-percentile value, and returned only companies where the value cleared the threshold. The computation is per company, across the people within. No catalog stores a company-level percentile measure of a people-level attribute. The prompt is structurally outside the schema.

Distributional measures surface in sales motions that depend on organizational health. A rapidly turning engineering team is a different buyer than a stable one. McKinsey research on organizational effectiveness has quantified how tenure distributions shift at inflection points. A GTM team that targets inflection points needs the distribution, not the average.

Prompt three: AE-to-SDR ratio between 1.5 and 3

The prompt asked for companies where the ratio of Account Executives to Sales Development Representatives falls between 1.5 and 3. The criterion is a ratio of two counts within one company, computed by enumerating both roles and dividing.

Clay declined. ZoomInfo declined. Apollo returned 11.5%. Landbase scored 100%.

Landbase enumerated the AE and SDR populations per company, computed the ratio, and returned only companies where the ratio sat in the band. The ratio is a company-level statistic computed from a people-level roster. Catalogs that store company-level attributes do not have the field. Catalogs that store people-level profiles do not aggregate across the company at query time.

Team-ratio criteria matter when the sales motion is targeting a specific GTM maturity. A 2-to-1 AE-to-SDR ratio describes a company whose outbound function is scaling but still SDR-driven. A 10-to-1 ratio describes a company that has moved past that stage. Different motions, different message.

What the three prompts have in common

Each prompt asks for a value that is computed, not stored. Follower growth is a difference. Tenure percentile is a distributional statistic. AE-to-SDR ratio is a within-company ratio. None of the three corresponds to a column in a vendor catalog.

A workflow builder like Clay can compose multi-step pipelines against stored attributes. If the attribute is not in any provider's schema, the composition has nothing to compose from. The platform returns an empty filter set and the prompt scores 0. A catalog like ZoomInfo faces the same structural limit. Apollo's AI research layer attempted computation, which raised precision above zero on one of the three, and the result is still an order of magnitude below the agent approach.

Landbase scored 100%, 96.1% and 100%. The agent reads underlying records at list-build time, computes the derived measure per company, and verifies each row against the computed value before returning it. The derivation lives at the prompt layer rather than the schema layer.

Why derived criteria matter to GTM teams

Derived criteria concentrate buying signals that catalog criteria dilute. A list of companies in a given industry at a given size captures a population. A list of companies in that industry, at that size, whose engineering roster turned over in the last 12 months captures a subset where a specific trigger has fired.

The subset is the file that reps want to work. Gartner research on sales technology adoption has argued that triggered outbound outperforms non-triggered outbound on every measured conversion metric. Triggers are derived. Catalogs store state, not change.

GTMBench selected the three prompts to probe this boundary. The 23 other prompts in the benchmark required the systems to answer within their schema. The three derived prompts required computation. The results separated the two architectures cleanly.

What this means for prospecting infrastructure

Buyers evaluating prospecting tools often compare database size and filter count. The three prompts show that both measures miss a structural capability. A tool with 10x the records and 10x the filters still returns 0 on a prompt that requires computing a value the schema does not store.

The gap matters most for teams whose best-performing lists use derived criteria. Hiring signals, team-mix ratios, growth-stage transitions, product-launch cadence, funding timing and executive-tenure patterns all require derivation. Teams whose outbound depends on those signals need a system that can derive at list-build time.

For teams whose motion is firmographic (industry, size, geography), catalog systems compete head to head. For teams whose motion depends on derived signals, the three prompts in GTMBench suggest the structural gap is not closable through filter composition alone.

What Landbase delivers

Landbase is an agent that reads a plain-language prompt, reasons across more than 1,500 enrichment fields per company, computes derived measures at list-build time, verifies each row before returning it, and dial-tests the file before handoff. The platform runs inside Claude Code and Codex, and connects to Salesforce, HubSpot and CSV export. HubSpot sales statistics have documented how much of an SDR week depends on list quality upstream of the dial.

Across GTM teams from agencies to enterprise revenue organizations, customers have reported a 2 to 4x uplift in connect and meeting-booked rates. If your best-performing lists use derived criteria, send Landbase the prompt and the platform will build and score it in one pass at no cost. Start a derived-criteria pass here. For related reading, see the Landbase vs Clay comparison and Landbase vs ZoomInfo comparison.

Frequently asked questions

Why can't Clay reach these criteria through workflow composition?

Clay composes against stored attributes across 200-plus data providers. If no provider stores the attribute (and none in the marketplace stores follower growth, tenure percentiles or team ratios as queryable fields), the composition returns nothing. The three prompts returned empty filter sets, not partial matches.

Could ZoomInfo add these fields to its catalog?

Each of the three criteria is a computed statistic across underlying records. Adding them as stored fields would require pre-computing the measure for every company in the catalog, refreshed at a cadence that captures change. The computation volume scales with the number of possible measures a buyer might ask for, which is why catalog platforms generally do not pre-compute arbitrary derived fields.

Does Apollo's AI research layer handle derived criteria?

Partially. On the three prompts, Apollo returned 0, 2.0% and 11.5% precision. The AI research layer did more than catalog filtering alone, but trailed the agent approach by 85 to 100 percentage points. The gap suggests the architecture reaches derived criteria in principle but not at the precision an SDR workflow requires.

Which other kinds of criteria are derived in this sense?

Any criterion that requires a computation across underlying records qualifies. Hiring velocity (rate of new hires per quarter), funding-stage transitions, product-launch cadence, revenue-per-employee ratios, geographic concentration measures, and any percentile or quartile across a people-level attribute all fit the pattern. The three in the benchmark are a sample of the category.

Build a GTM-ready audience

Qualify your list in one pass

  • Button with overlapping square icons and text 'Copy link'.

Turn this list into a GTM-ready audience

Match this list to your ICP, prioritize accounts, and identify who to contact using live growth signals.

Run your list through Landbase

Send a list you have already pulled from Clay, ZoomInfo or Apollo. Landbase will qualify, expand and score it in one pass at no cost.

Stop managing tools. 
Start driving results.

See Agentic GTM in action.
Get started
Our blog

Lastest blog posts

Tool and strategies modern teams need to help their companies grow.

Research

Buyers evaluate prospecting tools on database size and filter counts. Nothing in that measures what reaches the sales floor. GTMBench runs 26 natural-language prompts against four vendors to show.

Daniel Saks
Chief Executive Officer
Research

Two prominent GTM platforms returned nothing on three prompts in the 2026 GTMBench. The common thread was derived criteria no catalog stores as a field.

Daniel Saks
Chief Executive Officer
Insight

Database size predicts reach. Precision predicts how many of those rows a rep can work. The 2026 GTMBench measured both.

Daniel Saks
Chief Executive Officer

How GTM teams turn this list into pipeline

See how GTM teams use fastest-growing lists to define TAM, prioritize accounts, and launch campaigns.