Avocado Datalake
Avocado Datalake simplifies data lake management for your organization.
Avocado Datalake Architecture
A high-level architecture of our proposed solution for managing all sources of data into a unified Data Lake.

For structured data sources, we support relational databases such as MySQL, PostgreSQL, SQL Server, Amazon Aurora, and GCP Cloud SQL. For semi-structured sources, we support MongoDB, Amazon DynamoDB, and Google Cloud Bigtable. For unstructured and file-based sources, we handle Apache Kafka streams as well as CSV, JSON, Parquet, Avro, and ORC files stored on GCS, S3, or Azure Blob.
By leveraging open table formats — Apache Hudi, Delta Lake, and Apache Iceberg — we ensure Change Data Capture (CDC), ACID-like guarantees, schema evolution, and optimized read/write operations across all cloud platforms.
Our Data Engineering Products
Production-ready, highly configurable ETL/ELT pipelines delivered as codebase — set up in your own cloud account in days, not months.
RDBMS to Data Lake
Batch, incremental, and CDC pipelines from MySQL, PostgreSQL, SQL Server, Oracle, AWS Aurora, and GCP Cloud SQL into Iceberg, Hudi, or Delta Lake.
NEWGCP Cloud SQL → Iceberg Lakehouse
HOCON-driven Spark JDBC-to-Iceberg ELT pipeline on GCP Dataproc Serverless — incremental checkpoints, upserts, and Airflow/Livy orchestration. Up to 98% cheaper than Cloud Datastream.
NEWGCS Files → Iceberg Lakehouse
Checkpoint-driven Spark pipeline that ingests Parquet, CSV, JSON, and Avro files from Google Cloud Storage directly into Apache Iceberg tables — with idempotency guards and schema inference.
Semi-Structured Sources to Data Lake
Ingest MongoDB, DynamoDB, Firestore, Cassandra, and JSON/XML feeds into your data lake on AWS, GCP, or Azure with automated Spark pipelines and full data cataloguing.
Unstructured Data to Data Lake
File-based ingestion pipelines for CSV, JSON, Parquet, Avro, ORC, and web API responses — with Spark ETL, data cataloguing, and open table format support.
Data Lake → Data Warehouse
Sync curated data lake tables into BigQuery, Redshift, Snowflake, or Azure Synapse Analytics — DIY YAML/HOCON configuration, no bespoke code required.
Featured Technical Blogs
Deep-dive guides, cost benchmarks, and hands-on tutorials covering Data Lake architecture, Apache Iceberg, Apache Hudi, Delta Lake, and GCP data engineering.
GCP Cloud SQL to Iceberg Lakehouse vs. Google Cloud Datastream
Detailed cost comparison showing 98% savings with our Spark JDBC-to-Iceberg pipeline vs. Datastream.
Read more →GCS Files to Iceberg vs. GCP Managed Iceberg Tables
Architecture and cost breakdown comparing our checkpoint-driven pipeline against BigLake managed tables.
Read more →GCS Files to Iceberg vs. Google Cloud Dataplex Ingestion
Governance, scheduling, and cost tradeoffs between our Spark pipeline and Dataplex ingestion tasks.
Read more →Getting Started with GCP Iceberg Lakehouse — Scala / SBT
Step-by-step guide to running Spark + Iceberg on Dataproc Serverless using the BigLake REST catalog.
Read more →What Is a Data Lake?
Core concepts, architecture patterns, and implementation strategies for AWS S3, GCP Cloud Storage, and Azure.
Read more →Data Lake vs Data Warehouse
Schema-on-read vs schema-on-write, raw storage vs curated analytics — when to use each.
Read more →View all blogs → · See why Avocado Datalake is the right choice for your organisation.


