Daniel Saks
Chief Executive Officer
Buyers write prompts in the language they use at planning meetings. Sticky brands. Creator economy startups. PLG motion. Developer-first companies. These labels describe real buyer patterns, but they do not map cleanly to any column in a vendor catalog.
The 2026 GTMBench included six concept prompts across the 26-prompt set to measure how four systems interpret fuzzy criteria. Landbase averaged 65.6% precision across the family. Apollo averaged 49.1%. Clay averaged 40.0%. ZoomInfo averaged 20.7%. Harvard Business Review coverage of B2B selling has documented how buying signals increasingly cluster around behavioral and category-level patterns that resist precise definition. Gartner research on sales technology adoption notes that concept-driven prompts dominate planning conversations in modern GTM teams. The full methodology lives on the GTMBench landing page.
A concept prompt uses a label whose meaning is understood among operators but not stored as a vendor attribute. Three characteristics separate concept prompts from firmographic or derived ones.
First, the label is interpretive. Sticky brands has no single definition. It could mean high repeat-purchase rate, high brand recall in surveys, or persistent category share over time. Different operators mean slightly different things by the label.
Second, the pattern that defines the label spans several signals. A creator economy startup might combine a specific business model, a specific customer segment, and a specific product shape. No single attribute captures the pattern.
Third, the decision at the row level is a classification rather than a lookup. Is this company a sticky brand returns a judgment, not a value. The judgment requires evaluating the pattern against the company's record.
The concept family in GTMBench included Microsoft-ecosystem consulting partner firms, creator economy startups, sticky subscription consumer-goods brands, US-HQ B2B SaaS 10 to 100 employees product-led, Series A to C companies where engineering is 40% or more of headcount, and B2B software companies with an SDR to AE structure. Each prompt required the system to decide what the label meant before it could apply the criterion.
Landbase scored above 80% on four of the six. Apollo scored above 50% on three. Clay scored above 50% on two. ZoomInfo scored above 50% on none. The pattern mirrors the derived family: agent reasoning handles interpretive criteria more consistently than filter composition.
Filter composition in a workflow builder like Clay reaches concept criteria by chaining keyword matches, descriptive-text searches and taxonomy filters. Sticky brands might be approximated as consumer-goods companies in a stored vertical with high repeat-customer ratios in a stored CRM-metric field. The approximation works when the vendor's stored attributes carry the right proxies.
The approximation breaks when the proxies are weak. If no stored attribute captures retention rate at the company level, the sticky brands filter collapses into consumer-goods companies with no retention signal. The returned rows nominally match the vertical but miss the sticky part of the criterion.
Clay scored 40.0% average on the concept family. Apollo scored 49.1%. The precision came from prompts where the filter approximation sat close to the true criterion. The gap came from prompts where no stored attribute approximated the label well. Forrester research on B2B data readiness has documented how approximation gaps appear most sharply on criteria that span business-model dimensions.
At list-build time, a reasoning agent parses the concept label and decomposes it into underlying patterns. For sticky brands, the agent considers repeat-purchase cadence, subscription model shape, customer-tenure signals and category persistence. For creator economy startups, the agent considers business model (platform versus publisher), customer segment (individual creators versus brand clients) and revenue model.
Each decomposition is checked against each company's record before the row returns. The verification pass tests whether the company's signals match the pattern, not whether a stored attribute matches a value. Landbase averaged 65.6% precision across the family because the decomposition recovers more of the operator's intent than a filter approximation does.
The ceiling on concept precision is lower than on firmographic prompts because the labels themselves carry definitional ambiguity. 65.6% average reflects the share of returned rows where the agent's interpretation of the label matched the judge's interpretation at verification. Perfect concept precision would require that every operator and every judge share the same interpretation, which the labels resist.
Concept prompts dominate real planning conversations. GTM teams target creator economy startups, PLG companies, developer-first infrastructure businesses, enterprise AI-native platforms, and dozens of similar labels. The labels describe real patterns that correlate with buying behavior. The patterns resist clean filter definition.
Teams whose best-performing lists use concept criteria face a vendor-selection question that catalog claims do not predict. A tool with the deepest firmographic schema does not necessarily interpret concept labels well. The 25-point family gap between Landbase and the next-best vendor on concept prompts suggests the interpretation layer is the differentiator, not the schema layer.
For teams whose motion depends on concept-level targeting, the vendor evaluation shifts from how big is the catalog to how well does the system interpret labels. See the derived prompts post for the related case where no stored attribute exists at all.
At list-build time, the Landbase agent interprets the concept label by decomposing it into the underlying signals the operator plausibly intends. Each candidate decomposition is weighted against the prompt context: adjacent constraints in the prompt, the industry, the apparent sales motion.
Each company is then checked against the decomposed pattern. The check uses both stored attributes (industry, size, geography) and reasoning over underlying records (business model shape, customer-base signals, product cues). Only companies that match the pattern at a reasonable confidence return as rows.
The verification pass runs on the pattern match, not on a stored attribute. The precision cost is runtime: Landbase takes 134 seconds against Clay's 17. For concept prompts, the per-company reasoning is a larger share of the runtime than it is for firmographic prompts.
Landbase is an agent that reads a plain-language prompt, reasons across more than 1,500 enrichment fields per company, interprets concept labels at list-build time, verifies each row before returning it, and dial-tests the file before handoff. The platform runs inside Claude Code and Codex, and connects to Salesforce, HubSpot and CSV export. Salesforce State of Sales research has documented how much of the buying-pattern recognition in modern GTM depends on behavioral and category-level signals that catalog systems do not store.
Across GTM teams from agencies to enterprise revenue organizations, customers have reported a 2 to 4x uplift in connect and meeting-booked rates. If your best-performing lists use concept criteria, send Landbase the prompt and the platform will build and score it in one pass at no cost. Start a concept-criteria pass here. For the vendor-by-vendor view, read our Landbase vs Clay comparison.
Concept labels concentrate buying patterns that firmographic criteria dilute. A list of consumer-goods companies at a given size captures a population. A list of sticky subscription consumer-goods brands captures a subset where a specific buying pattern holds. The subset converts at a higher rate on the dial because the pattern filters out companies without the trigger.
Yes. Prompts can specify both the label and the decomposition. For example, sticky brands with repeat-purchase rate above 60% or subscription retention above 90% narrows the agent's interpretation to the operator's definition. The precision rises when the decomposition is specified.
Clay's research layer can enrich individual rows after the catalog composition returns them. On concept prompts, the composition often returns the broader population (consumer-goods companies for sticky brands), and the research layer attempts to classify each row against the label. The approach raises precision above pure filter composition but trails per-row agent verification at build time.
Yes. When the label decomposes to a small set of stored attributes with strong proxies, filter composition can approximate well. Microsoft-ecosystem consulting partner firms is close to the boundary: there is a stored partner directory that captures most of the population. Clay led this family on one of the six concept prompts in the benchmark.
The full report, with every prompt scored across all four vendors and the family-level averages, lives on the GTMBench page.
Tool and strategies modern teams need to help their companies grow.