Topic · topic
Cloud Storage and Data Acquisition
Cloud Storage and Data Acquisition is a topic profile in the API Evangelist Network covering APIs and tooling for ingesting, moving, and persisting bulk and streaming data into cloud-resident storage. It groups object storage services, data-lake foundations, managed ingestion pipelines, change-data-capture connectors, transfer appliances, and data-broker APIs that provide source material for cloud-storage workloads. The topic is intended as an entry point for developers and architects evaluating how data lands in cloud storage from on-premises systems, SaaS APIs, IoT devices, public data sources, and partner exchanges.
Resources
-
Object Storage Surface
The object-storage surface includes Amazon S3, Google Cloud Storage, and Azure Blob Storage REST APIs. These APIs provide the canonical landing zone for cloud data acquisition and are the most common targets for ingestion pipelines.
-
Streaming Ingest Surface
The streaming-ingest surface covers Amazon Kinesis Data Streams, Google Cloud Pub/Sub, and Azure Event Hubs. These services ingest high-volume event streams and make them durable for downstream landing into object storage and analytics sys…
-
Managed Ingestion Pipelines
Managed ingestion pipelines such as AWS Glue, Google Cloud Dataflow, and Azure Data Factory expose REST APIs for orchestrating extract-transform-load and extract-load-transform jobs that land source data in cloud storage.
-
Change Data Capture Connectors
Change Data Capture connectors (Debezium, AWS DMS, Fivetran, Striim) replicate row-level changes from operational databases to cloud storage and warehouses, providing a low-latency feed for analytics and data-lake hydration.
-
Bulk Transfer Services
Bulk transfer services (AWS DataSync and Snow family, Google Storage Transfer Service, Azure Data Box) move large datasets from on-premises and edge locations into cloud storage over network or via offline appliances.
-
Data Marketplaces
Data marketplaces and open-data registries expose third-party datasets through REST APIs, subscription delivery, and shared buckets. They are an increasingly common acquisition channel for cloud-storage data lakes.
Links
Providers working in Cloud Storage and Data Acquisition
Providers whose own tags share at least two of this topic's tags, most shared first — the top 30 of 48.
| Provider | About | Rating | APIs |
|---|---|---|---|
| Nexla | Nexla is an enterprise data integration and AI-data platform, founded in 2016 and headquartered in San Mateo, California. Its core abstraction is the Nexset — a logical, schema-aware, bi-directionally usable data product that Nexla generat… | strong | 4 |
| Artie | Artie is a real-time data replication platform that streams database changes to cloud data warehouses and lakehouses with sub-minute latency and exactly-once delivery. It captures change data (CDC) from sources such as PostgreSQL, MySQL, M… | developing | 1 |
| Popsink | Popsink is a real-time data replication and change data capture (CDC) platform that continuously moves data out of mission-critical and legacy systems into cloud data platforms with low latency and minimal production impact. It offers a br… | developing | 2 |
| Estuary Flow | Estuary Flow is a real-time data movement and transformation platform combining streaming infrastructure, a runtime, and an open-source ecosystem of connectors. It supports change data capture (CDC), SaaS integration, database replication,… | thin | 4 |
| Streamkap | Streamkap is a real-time streaming ETL and change data capture (CDC) platform built on Apache Kafka and Apache Flink. It streams data from operational databases (PostgreSQL, MySQL, MongoDB, SQL Server, Oracle) to cloud warehouses, lakes, a… | thin | 1 |
| Amazon S3 | Amazon Simple Storage Service (S3) is an object storage service offering industry-leading scalability, data availability, security, and performance. | exemplar | 2 |
| Azure Data Factory | Azure Data Factory is Microsoft's cloud-based data integration service, orchestrating and automating the movement and transformation of data across ETL and ELT workloads that span cloud and on-premises stores. Its public interface is the M… | exemplar | 1 |
| Google Cloud Storage | Object storage service offering high durability, availability, and scalability for storing and accessing data on Google Cloud Platform. | strong | 1 |
| Azure Blob Storage | Microsoft Azure Blob Storage is a service for storing large amounts of unstructured object data, such as text or binary data, that can be accessed from anywhere in the world via HTTP or HTTPS. | strong | 1 |
| Feldera | Feldera is an incremental compute engine for running complex SQL data pipelines in real time. Rather than reprocessing entire datasets, it updates materialized views proportionally to the changes in the input data, delivering low-latency,… | strong | 1 |
| LucidLink | LucidLink Corp. is a cloud file-streaming company whose product is a "filespace" — a shared, cloud-native filesystem that mounts on macOS, Windows, Linux, iOS and Android and streams only the bytes an application actually asks for, so dist… | strong | 1 |
| Amazon Redshift | Amazon Redshift is a fast, fully managed cloud data warehouse that makes it simple and cost-effective to analyze all your data using standard SQL and your existing Business Intelligence (BI) tools. | strong | 1 |
| Archil | Archil is the cloud filesystem for AI. It turns an existing object-storage bucket (Amazon S3, Google Cloud Storage, Cloudflare R2, Azure Blob, MinIO, Wasabi, Backblaze B2, DigitalOcean Spaces) into an unlimited, POSIX-compatible local disk… | strong | 1 |
| Pure Storage | Pure Storage is an American publicly traded technology company specializing in all-flash data storage hardware and software products. The company provides enterprise data storage platforms including FlashArray, FlashBlade, and Pure1 fleet… | strong | 3 |
| Delta Lake | Delta Lake is a graduated project of the Linux Foundation AI & Data Foundation providing an open source storage framework for building Lakehouse architectures. Originally contributed by Databricks, it adds reliability, quality, and perform… | developing | 1 |
| WarpStream | WarpStream is a diskless, Apache Kafka-compatible data streaming platform built directly on top of cloud object storage such as S3, GCP, and Azure. It eliminates the need for local disks, brokers to rebalance, and ZooKeeper, delivering Kaf… | developing | 1 |
| Google Cloud Datastream | Google Cloud Datastream is a serverless change data capture (CDC) and replication service that allows you to synchronize data across heterogeneous databases, storage systems, and applications reliably and with minimal latency. | developing | 1 |
| Backblaze | Backblaze is a cloud storage and data backup provider offering B2 Cloud Storage - a low-cost, S3-compatible object storage service. Backblaze provides both a native B2 API and an S3-compatible API, enabling developers to build applications… | developing | 7 |
| Striim | Unified data integration and streaming platform offering change data capture (CDC), real-time streaming analytics, and data validation. Exposes a token-authenticated REST API (WActionStore queries, system health, Application Management) co… | developing | 6 |
| Flatfile | Flatfile is a data exchange platform that helps teams import, transform, validate, and collaborate on file-based data. The Flatfile API provides programmatic access to spaces, workbooks, sheets, records, files, documents, jobs, events, age… | developing | 1 |
| Weld | Weld is a programmable data-infrastructure platform for moving and transforming data. It runs near real-time ELT/ETL pipelines from 300+ prebuilt connectors, log-based Change Data Capture (CDC), SQL-based data transformations with lineage… | developing | 1 |
| Cloudflare R2 | Cloudflare R2 is S3-compatible object storage with zero egress fees. It provides both an S3-compatible REST API and a Cloudflare API for managing buckets, objects, and storage policies at global scale. R2 offers a generous free tier includ… | developing | 2 |
| Unstructured | Unstructured is a document parsing and pre-processing platform that provides a REST API for ingesting PDFs, HTML, DOCX, images, and more than 50 other file formats, transforming them into clean structured JSON chunks ready for RAG pipeline… | developing | 2 |
| Eon | Eon is a next-generation cloud backup and data-protection platform that turns cloud backups into live, searchable, strategic assets across AWS, Google Cloud, and Azure. It provides agentless backup, cloud backup posture management, ransomw… | developing | 1 |
| Talend | Talend (now part of Qlik) provides data integration, quality, and API management capabilities through cloud-native APIs for ETL, data pipelines, and application integration. The Qlik Talend Cloud platform exposes REST APIs for orchestratin… | developing | 2 |
| Filebase | Filebase is an S3-compatible object storage and IPFS pinning platform that combines familiar cloud storage APIs with decentralized, blockchain-backed infrastructure. Developers can store, manage, and pin files to IPFS using standard S3 too… | developing | 4 |
| Apache NiFi | Apache NiFi is a dataflow management system designed to automate the flow of data between systems. It provides a web-based user interface for designing, controlling, and monitoring data flows with real-time operational control, data proven… | developing | 1 |
| Sequin | Sequin is an open-source Postgres change data capture (CDC) engine that streams Postgres rows and changes to streams, queues, and search indexes - Kafka, SQS, SNS, Kinesis, Redis, NATS, RabbitMQ, Elasticsearch, Typesense, GCP Pub/Sub, Azur… | thin | 1 |
| ByteArk | ByteArk is a Thailand-based video streaming and content delivery platform founded in 2012 and headquartered in Bangkok. It provides video-on-demand (ByteArk Stream), live streaming (Fleet / Teatro), an S3-compatible object storage service,… | thin | 1 |
| Infoworks | Infoworks is an Enterprise Data Operations and Orchestration (EDO2) platform that automates data onboarding, preparation and operationalization onto Databricks, Snowflake, BigQuery, Synapse and Apache Spark. Unlike a multi-tenant SaaS, Inf… | thin | 1 |
Tags
Bulk TransferChange Data CaptureCloud StorageData AcquisitionData IngestionData LakeETLObject StoragePipelinesStreaming