Topic · topic
Lakehouse Architecture
Lakehouse Architecture is a data architecture paradigm that combines the best features of data lakes and data warehouses, providing ACID transactions, schema enforcement, and governance on low-cost storage with support for both business intelligence and machine learning workloads.
Resources
-
Lakehouse Architecture
Resources and reference implementations for the Lakehouse Architecture data platform paradigm.
Links
Providers working in Lakehouse Architecture
Providers whose own tags share at least two of this topic's tags, most shared first — the top 30 of 66.
| Provider | About | Rating | APIs |
|---|---|---|---|
| Amazon Redshift | Amazon Redshift is a fast, fully managed cloud data warehouse that makes it simple and cost-effective to analyze all your data using standard SQL and your existing Business Intelligence (BI) tools. | strong | 1 |
| Treasure Data | Treasure Data — rebranded Treasure AI in April 2026 — is an enterprise customer data platform that unifies first-party customer data and activates it across marketing, service and AI agent workloads. It publishes eight OpenAPI descriptions… | exemplar | 7 |
| Azure Synapse Analytics | Azure Synapse Analytics is an enterprise analytics service that accelerates time to insight across data warehouses and big data systems. It brings together the best of SQL technologies used in enterprise data warehousing, Spark technologie… | strong | 30 |
| AWS Redshift | Amazon Redshift is a fast, fully managed, petabyte-scale data warehouse service that makes it simple and cost-effective to analyze all your data using standard SQL and existing Business Intelligence tools. | developing | 2 |
| Ocient | Ocient is a Chicago-based data platform company founded in 2016 that builds OcientAIQ, a unified data platform for petabyte-scale analytics and production AI. Its Compute-Adjacent Storage Architecture (CASA) colocates NVMe storage with com… | developing | 1 |
| Google BigQuery | Google BigQuery is a fully managed, serverless data warehouse that enables scalable analysis over petabytes of data using SQL. | developing | 1 |
| Qubole | Qubole is a cloud-native data lake platform (now part of Idera) that lets teams run multiple open-source data-processing engines - Apache Spark, Presto, Hive, Hadoop, and Airflow - together in a single, cost-optimized, self-managing enviro… | thin | 1 |
| Azure Data Lake Storage | Azure Data Lake Storage Gen2 REST API provides a file system interface for big data analytics workloads on Azure Blob Storage. It supports creating file systems, managing directories and files with hierarchical namespace, setting ACLs, and… | thin | 1 |
| SQream Technologies | SQream Technologies is an Israeli data and analytics company, founded in 2010 in Tel Aviv, that builds SQreamDB — a GPU-accelerated, SQL-compliant analytics database for petabyte-scale workloads on NVIDIA hardware — alongside AISQream for… | thin | 0 |
| Etleap | Etleap is a managed ETL and data-integration platform that streamlines data ingestion, transformation, and observability so data teams can build cloud data warehouses and lakes with minimal engineering effort. Originally built as an "autop… | thin | 1 |
| Netezza | IBM Netezza Performance Server is a cloud-native data warehouse and analytics appliance for running large-scale SQL analytics and in-database machine learning on structured data. Originally a standalone data-warehouse appliance vendor acqu… | emerging | 0 |
| Cazena | Cazena was a Waltham, Massachusetts-based enterprise software company, founded in 2014, that delivered a fully-managed SaaS data platform — a "Big Data as a Service" cloud data lake that let enterprises run analytics, machine learning, and… | minimal | 0 |
| Amiato | Amiato was a Palo Alto, California big-data startup founded in 2011 by Nathan Binkert and Mehul A. Shah, originally incubated at Y Combinator in 2012 under the name Nou Data. Amiato built a real-time analytics pipeline around its Schema-li… | 0 | |
| BitYota | BitYota was a Data-Warehouse-as-a-Service (DWaaS) startup founded around 2011 in Santa Clara, California, offering a SaaS analytics data warehouse that ran on top of Amazon Web Services and Rackspace infrastructure and natively queried sem… | 0 | |
| Snowflake | Snowflake is a cloud-based data platform delivering data warehousing, data lakes, data engineering, data science and data application development as a single managed service across AWS, Azure and Google Cloud. Its developer surface is a 47… | exemplar | 84 |
| Sigma Computing | Sigma Computing is a warehouse-first analytics and business-application platform. Instead of extracting data into its own store, Sigma queries the customer's cloud data warehouse live — Snowflake, Databricks, BigQuery, Redshift, Amazon Ath… | exemplar | 3 |
| TextQL | TextQL is an enterprise AI data platform built around Ana, an AI data scientist that connects to a company's warehouses, databases, BI tools and SaaS APIs and answers questions in plain language. Ana writes SQL, runs Python in a managed gV… | exemplar | 15 |
| Cloudera | Cloudera is a hybrid data platform company offering the Cloudera Data Platform (CDP) for data engineering, data warehousing, machine learning, streaming, and operational data. The platform exposes multiple REST APIs including the CDP Publi… | exemplar | 38 |
| Monte Carlo | Monte Carlo is a data and AI observability platform that monitors data warehouses, lakes, and pipelines for freshness, volume, schema, and quality anomalies, helping data teams detect, resolve, and prevent data downtime across Snowflake, D… | strong | 1 |
| Azure Databricks | Azure Databricks is an Apache Spark-based analytics platform optimized for Microsoft Azure. It provides a collaborative workspace for data engineers, data scientists, and analysts to work together on big data and machine learning workloads. | strong | 1 |
| Amazon Lake Formation | AWS Lake Formation is a service that makes it easy to set up a secure data lake in days, providing centralized governance and security for data stored in Amazon S3 and other AWS data stores with fine-grained access control. | strong | 1 |
| Beaconstac | Beaconstac (now Uniqode) is a B2B SaaS platform for creating, customizing, and tracking dynamic QR Codes and Digital Business Cards at scale, connecting physical touchpoints to measurable digital experiences for 50,000+ brands. Its REST AP… | strong | 1 |
| Databricks | Collection of Databricks REST APIs for managing workspaces, clusters, jobs, and data operations. | strong | 1 |
| Funnel | Funnel (funnel.io) is a marketing intelligence and marketing data hub that helps agencies and brands become more data-driven. It connects to hundreds of advertising, analytics, CRM, and social data platforms, then automatically collects, n… | strong | 3 |
| Amazon EMR | Amazon EMR is a cloud big data platform for running large-scale distributed data processing jobs, interactive SQL queries, and machine learning applications using open-source analytics frameworks such as Apache Spark, Apache Hive, Apache H… | strong | 1 |
| Supermetrics | Supermetrics is a marketing intelligence platform that automates the pipeline of marketing and advertising data from 100+ sources (Google Ads, Facebook/Meta Ads, TikTok, Google Analytics, LinkedIn Ads, and more) into spreadsheets, BI tools… | developing | 1 |
| Definite | Definite is an all-in-one, AI-native data platform that consolidates data integration, warehouse storage, a semantic layer, BI dashboards, and AI agents into a single product. It ships 500+ managed data connectors, a DuckDB / DuckLake lake… | developing | 1 |
| AWS Kinesis | Amazon Kinesis is a family of fully managed AWS services for collecting, processing, and analyzing real-time streaming data. The family includes Kinesis Data Streams for scalable record ingestion, Amazon Data Firehose (formerly Kinesis Dat… | developing | 4 |
| Apache Iceberg | Apache Iceberg is an open table format for large analytic datasets that provides ACID transactions, schema evolution, hidden partitioning, and time travel. It works with Spark, Flink, Hive, Presto, Trino, DuckDB, ClickHouse, and many more… | developing | 1 |
| Ryft | Ryft is the intelligent Apache Iceberg management platform - a data lakehouse optimization layer that continuously monitors, manages, and optimizes Iceberg tables across query engines and clouds. It provides automated table compaction, sna… | developing | 1 |
Tags
AnalyticsBig DataData ArchitectureData LakeData Warehouse