Daniel Saks
Chief Executive Officer
Data engineering covers the systems and processes used to collect, move, transform, store, govern, and prepare data for downstream applications such as analytics and machine learning.
The category has expanded as AI applications require larger volumes of structured, reliable, and continuously available data. Modern data engineering platforms now span data integration, transformation, orchestration, databases, real-time processing, observability, and AI-oriented infrastructure.
The companies below stand out based on recent revenue or ARR growth, customer adoption, financing, acquisitions, or other publicly documented indicators. Because private and public companies disclose different metrics, the list is a growth watchlist rather than a strict ranking.
Data engineering sits between raw information and the applications that rely on it.
Modern data infrastructure commonly includes:
Workflow orchestration is one important layer. The Apache Airflow architecture, for example, organizes workflows as directed graphs of dependent tasks that can fetch data, run analysis, or trigger other systems.
Real-time processing introduces another model. Event streaming allows data to be captured, processed, stored, and routed continuously rather than relying entirely on scheduled batch movement.
These architectures increasingly support analytics and AI systems alongside traditional data workloads.
Founded: 2013
CEO: Ali Ghodsi
Headquarters: San Francisco, California
Databricks surpassed a $7 billion annualized revenue run-rate in 2026, with the company reporting growth of more than 80% year over year during its second quarter.
In August 2026, Databricks also raised $5 billion at a $190 billion valuation, substantially above its valuation earlier in the year.
Databricks provides a unified data and AI platform covering:
The company is closely associated with the lakehouse architecture, which combines characteristics of data lakes and warehouses.
Databricks has expanded from large-scale data processing into a broader infrastructure platform supporting data engineering, analytics, databases, machine learning, and AI applications.
Its current growth illustrates increased demand for platforms that connect data preparation and AI development within the same environment.
Fivetran Founded: 2012
dbt Labs Founded: 2016
CEO: George Fraser
President: Tristan Handy
Fivetran and dbt Labs completed their merger in June 2026, so they should no longer be presented as two independent growth companies.
At the time of the combination, the businesses represented approximately $600 million in combined ARR and more than 10,000 customers, while their broader products and communities supported more than 100,000 data teams.
The two platforms cover complementary parts of the data engineering lifecycle.
Fivetran focuses primarily on:
dbt focuses on:
The merger reflects consolidation within the modern data stack.
Data movement and transformation have historically been handled through separate platforms. Combining them creates a broader infrastructure layer spanning ingestion, transformation, metadata, and downstream data preparation.
This development also reflects increased demand for structured and governed data that can support both analytics and AI applications.
Company Founded: 2021
CEO: Aaron Katz
Headquarters: San Francisco Bay Area
ClickHouse surpassed $250 million in annual run-rate revenue in May 2026, more than tripling its ARR year over year.
The company also reported more than 4,000 cloud customers, including more than 1,000 added during the first part of 2026.
ClickHouse raised a $400 million Series D in January 2026 at a reported $15 billion valuation.
ClickHouse develops a column-oriented analytical database designed for high-volume analytical workloads.
Its platform supports:
The ClickHouse database originated as an internal project before being released as open source in 2016. ClickHouse Inc. was formed in 2021 around the technology.
The company's growth reflects demand for analytical databases capable of processing rapidly changing and high-volume datasets.
This is particularly relevant as analytics, observability, and AI applications increasingly need results with lower latency than traditional batch-oriented data architectures provide.
Founded: 2018
CEO: Clint Sharp
Location: Remote-first, with an office in San Francisco
Cribl surpassed $300 million in ARR in 2025, up from $200 million at the end of 2024.
The company also reported that its cloud ARR exceeded $130 million with more than 75% year-over-year growth, while multi-product adoption grew more than 90%.
Its platform currently serves more than 1,400 customers.
Cribl focuses primarily on telemetry data used by IT and security teams.
Its products support:
Cribl sits adjacent to traditional enterprise data engineering but addresses a similar infrastructure problem: how to move, shape, retain, and route high-volume data without tying every source to a single destination.
Growth in telemetry and AI-generated operational data is increasing the relevance of these architectures.
CEO: Pete DeJoy
Primary Product: Astro
Category: Data orchestration
Astronomer reported 122% year-over-year ARR growth in its EMEA business during its latest fiscal period, while its EMEA customer count increased 80%.
The company also reports that more than 900 organizations use Astronomer.
These are regional rather than company-wide growth figures, so they should not be interpreted as total corporate ARR growth.
Astronomer develops Astro, a data orchestration platform based on Apache Airflow.
Its capabilities include:
Orchestration determines when data jobs run, how dependent processes interact, and how failures are handled.
As data pipelines increasingly support production AI and machine-learning workloads, orchestration is becoming relevant beyond traditional ETL workflows.
Astronomer's current expansion reflects this broader role for workflow management within enterprise data infrastructure.
Founded: 2020
CEO: Michel Tricot
Category: Data integration
Airbyte currently reports approximately 7,000 companies using its platform, 1.5 million data pipelines synchronized daily, and more than 600 data-replication connectors.
The company has raised approximately $181 million since its founding.
Its 2026 product development also extends beyond traditional data replication into tools designed to provide business context to AI agents.
Airbyte began as an open-source data integration platform.
Its current capabilities include:
Data integration remains one of the basic building blocks of data engineering because downstream systems need reliable connections to operational sources.
Airbyte's open-source model also demonstrates how developer communities can accelerate connector coverage and adoption before an infrastructure platform expands into additional enterprise products.
Founded: 2022
CEO: Jordan Tigani
Headquarters: Seattle, Washington
MotherDuck reported approximately 850 paying customers after 18 months of commercial operation in mid-2026.
The company has raised approximately $100 million and continued expanding its data-engineering capabilities in August 2026 through the acquisition of Tower, a company focused on runtime infrastructure for data-engineering tasks.
MotherDuck develops a serverless analytical data platform built in collaboration with the DuckDB ecosystem.
Its platform covers:
MotherDuck takes a different approach from platforms designed primarily for extremely large distributed datasets.
Its architecture reflects increased interest in simpler analytical systems for workloads where developers and AI agents need fast access to manageable datasets without operating a large distributed infrastructure stack.
The company's recent acquisition also points toward broader support for agent-driven data engineering tasks.
Founded: 2014
CEO: Raj Dutt
Headquarters: New York, New York
Grafana Labs surpassed $600 million in ARR in August 2026, up from approximately $400 million less than a year earlier.
The company also crossed 10,000 customers worldwide, while monthly active Grafana Cloud users grew from approximately 127,000 to more than 251,000 over two years.
Grafana Labs operates primarily in observability rather than traditional data integration.
Its platform includes:
Observability overlaps with data engineering because data teams need visibility into whether pipelines, infrastructure, and downstream systems are behaving correctly.
As AI applications introduce more services, agents, models, and data flows, monitoring those systems becomes another important layer in production data infrastructure.
Grafana's current growth demonstrates increasing demand for this visibility.
Several themes emerge across these companies.
Moving and transforming data remain important, but the category increasingly includes orchestration, databases, observability, metadata, AI context, and workflow infrastructure.
Fivetran + dbt Labs covers movement and transformation, Astronomer focuses on orchestration, ClickHouse and MotherDuck address analytical storage and processing, while Cribl and Grafana operate closer to telemetry and observability.
Several companies on the list grew around widely adopted open-source technologies.
Examples include:
Open-source adoption can create a large technical community before commercial products reach equivalent scale.
Batch processing remains useful for many workloads, but real-time and near-real-time architectures are becoming more relevant to operational analytics, personalization, fraud detection, monitoring, and AI agents.
Technologies such as Apache Kafka illustrate how event-streaming systems continuously publish, process, and route information between applications.
AI applications do not remove the need for data engineering.
They increase the number of systems consuming organizational data and can make problems in quality, freshness, permissions, or definitions more consequential.
Organizations building AI-driven systems therefore still need controls around data quality, lineage, security, access, and governance. The NIST AI Risk Management Framework provides broader guidance for managing risk when AI capabilities move into production environments.
General-purpose data engineering platforms manage information across many departments and use cases.
GTM teams face a narrower version of several of the same problems:
For RevOps teams, these problems increasingly overlap with technical data operations rather than remaining purely administrative CRM tasks.
Landbase focuses specifically on B2B and GTM data rather than general enterprise data infrastructure.
Its connected web platform and CLI allow technical GTM teams to work with audience creation, matching, enrichment, signals, and datasets through structured workflows.
One direct connection to data engineering is record reconciliation.
Landbase supports matching and enrichment workflows that begin with existing company or person data, match records against Landbase entities, and add available attributes.
This is conceptually similar to broader data-engineering processes where raw records are resolved, standardized, enriched, and prepared before use by downstream applications.
The Landbase CLI runs inside AI-assisted environments such as Claude Code and Codex while using the same underlying Landbase data and agent available through the web platform.
This makes it possible to incorporate GTM data operations into workflows involving scripts, connected tools, CRM records, and other technical systems.
Landbase therefore does not replace general-purpose data engineering platforms. It applies several related principles specifically to the creation and operationalization of B2B and GTM datasets.
Relevant indicators include revenue or ARR growth, customer adoption, funding, valuation increases, product usage, acquisitions, and expansion into new data workloads. Because private companies disclose different metrics, growth should be evaluated using several indicators rather than a single ranking formula.
Data engineering focuses on building systems that collect, move, transform, store, and prepare information. Data science focuses more heavily on using prepared data for statistical analysis, experimentation, prediction, and machine-learning models. Reliable data engineering provides the infrastructure on which many data-science workloads depend.
The category can include data integration, transformation, orchestration, warehousing, lakehouses, databases, streaming, data quality, observability, metadata, lineage, and governance. Modern AI infrastructure increasingly overlaps with several of these areas.
Data workflows often involve multiple tasks that must execute in a particular order. Orchestration systems manage dependencies, schedules, retries, execution environments, and monitoring so that pipelines can run reliably without every step being coordinated manually.
Landbase applies related concepts specifically to B2B data. Existing records can be uploaded, matched, enriched, organized into datasets, refined through audience logic, and transferred into downstream workflows. Its CLI allows these operations to be incorporated into technical environments such as Claude Code and Codex rather than being limited to manual web-interface workflows.
Tool and strategies modern teams need to help their companies grow.