Skip to main content

Open Source Ingestion & Maintenance Tool for Apache Iceberg

Achieve zero-effort Apache Iceberg table maintenance and the fastest ingestion into Iceberg and Parquet on S3 with OLake

Trusted by Engineers at

Benchmarks

4.01 billion rows processed in less than 2 hours. For $6

Time to move 4.01 Billion rows, Postgres to Apache Iceberg, full load

MetricsOLake GoOpen SourceAirbyteFivetranDebeziumEstuary
Rows synced4.01 Billion12.7 Million4.01 Billion1.28 Billion0.34 Billion
Elapsed time1.91 hours7.5 hours (failed sync)24 hours24 hours24 hours
Speed (Rows/Sec)5,80,113 RPS457 RPS46,395 RPS14,839 RPS3,982 RPS
Comparison1270x slower12.5x slower39.1x slower146x slower
CostOLake is OSS and self-hosted — only pay for your infrastructure.< $ 6$ 5,560$ 0 (free full load)$ 75$ 1,668

1 Parameter, 1.8 billion rows compacted in 27 minutes.

Faster & half the cost

View OLake Fusion benchmarks
OLake FusionSpark Compaction
Total Compaction Time27m 02s2.06x faster55m 47s
Compaction Cost/ Job$1.0652% less cost$2.19
Config Parameters110x simpler10+

Two Engines

Move data fast,
keep it fast

OLake Go

Replicate your databases fast, into Apache Iceberg & Parquet on S3.

12.5X

Faster than traditional tools

90%

Cost Savings with OSS

OLake Fusion

Keep your Apache Iceberg tables consistently performant and scalable.

2X

Faster than Apache Spark

50%

Cheaper Per Compaction Cycle

Features

Why teams pick OLake.

Running Ingestion

# run it
$ OLake sync
orders      1,204,553 rows   00:41  ✓
customers     318,090 rows   00:12  ✓
Run Progress57%

Fast

Replicate databases at scale, instantly

Sync MySQL, Postgres, MongoDB, Oracle, and more to Apache Iceberg with parallelised chunking, incremental sync, and change data capture, so your lakehouse tables stay fresh with minimal overhead.

s3://lake/warehouse

orders/
└─ metadata/
└─ v3.metadata.json4.3 KB
└─ snap-8814402.avro6.7 KB
└─ manifest-list.avro2.1 KB
└─ data/
└─ 00000-3-a9f2b.parquet118 MB
└─ 00001-3-c4e81.parquet121 MB
└─ 00002-3-1d7ac.parquet116 MB

Open

Built on open standards, no lock-in

Write directly to Apache Iceberg or Parquet with support for multiple catalogs including AWS Glue, Apache Hive Metastore, and REST catalogs like Nessie, Polaris, and Unity Catalog — your data stays queryable by any engine you choose.

YOUR-VPCAPC SOUTH 1
Postgress

res

OLake

container

s3://lakeiceberg

Controlled

Self-hosted, on your infrastructure

Deploy OLake entirely within your own cloud or on-prem environment, keeping full control over where your data lives — critical for regulated industries like financial services that can't compromise on data residency.

/orderlast 90 days

P95 Query Time

0.9s

Table Size

12.4 TB

Maintained

Keep tables fast as data keeps growing

Automate Iceberg table maintenance through compaction, clearing delete files, and trimming metadata so query performance and storage costs don't degrade as your lakehouse scales.

Running Ingestion

# run it
$ OLake sync
orders      1,204,553 rows   00:41  ✓
customers     318,090 rows   00:12  ✓
Run Progress57%

Fast

Replicate databases at scale, instantly

Sync MySQL, Postgres, MongoDB, Oracle, and more to Apache Iceberg with parallelised chunking, incremental sync, and change data capture, so your lakehouse tables stay fresh with minimal overhead.

s3://lake/warehouse

orders/
└─ metadata/
└─ v3.metadata.json4.3 KB
└─ snap-8814402.avro6.7 KB
└─ manifest-list.avro2.1 KB
└─ data/
└─ 00000-3-a9f2b.parquet118 MB
└─ 00001-3-c4e81.parquet121 MB
└─ 00002-3-1d7ac.parquet116 MB

Open

Built on open standards, no lock-in

Write directly to Apache Iceberg or Parquet with support for multiple catalogs including AWS Glue, Apache Hive Metastore, and REST catalogs like Nessie, Polaris, and Unity Catalog — your data stays queryable by any engine you choose.

YOUR-VPCAPC SOUTH 1
Postgress

res

OLake

container

s3://lakeiceberg

Controlled

Self-hosted, on your infrastructure

Deploy OLake entirely within your own cloud or on-prem environment, keeping full control over where your data lives — critical for regulated industries like financial services that can't compromise on data residency.

/orderlast 90 days

P95 Query Time

0.9s

Table Size

12.4 TB

Maintained

Keep tables fast as data keeps growing

Automate Iceberg table maintenance through compaction, clearing delete files, and trimming metadata so query performance and storage costs don't degrade as your lakehouse scales.

Running Ingestion

# run it
$ OLake sync
orders      1,204,553 rows   00:41  ✓
customers     318,090 rows   00:12  ✓
Run Progress57%

Fast

Replicate databases at scale, instantly

Sync MySQL, Postgres, MongoDB, Oracle, and more to Apache Iceberg with parallelised chunking, incremental sync, and change data capture, so your lakehouse tables stay fresh with minimal overhead.

s3://lake/warehouse

orders/
└─ metadata/
└─ v3.metadata.json4.3 KB
└─ snap-8814402.avro6.7 KB
└─ manifest-list.avro2.1 KB
└─ data/
└─ 00000-3-a9f2b.parquet118 MB
└─ 00001-3-c4e81.parquet121 MB
└─ 00002-3-1d7ac.parquet116 MB

Open

Built on open standards, no lock-in

Write directly to Apache Iceberg or Parquet with support for multiple catalogs including AWS Glue, Apache Hive Metastore, and REST catalogs like Nessie, Polaris, and Unity Catalog — your data stays queryable by any engine you choose.

YOUR-VPCAPC SOUTH 1
Postgress

res

OLake

container

s3://lakeiceberg

Controlled

Self-hosted, on your infrastructure

Deploy OLake entirely within your own cloud or on-prem environment, keeping full control over where your data lives — critical for regulated industries like financial services that can't compromise on data residency.

/orderlast 90 days

P95 Query Time

0.9s

Table Size

12.4 TB

Maintained

Keep tables fast as data keeps growing

Automate Iceberg table maintenance through compaction, clearing delete files, and trimming metadata so query performance and storage costs don't degrade as your lakehouse scales.

Running Ingestion

# run it
$ OLake sync
orders      1,204,553 rows   00:41  ✓
customers     318,090 rows   00:12  ✓
Run Progress57%

Fast

Replicate databases at scale, instantly

Sync MySQL, Postgres, MongoDB, Oracle, and more to Apache Iceberg with parallelised chunking, incremental sync, and change data capture, so your lakehouse tables stay fresh with minimal overhead.

s3://lake/warehouse

orders/
└─ metadata/
└─ v3.metadata.json4.3 KB
└─ snap-8814402.avro6.7 KB
└─ manifest-list.avro2.1 KB
└─ data/
└─ 00000-3-a9f2b.parquet118 MB
└─ 00001-3-c4e81.parquet121 MB
└─ 00002-3-1d7ac.parquet116 MB

Open

Built on open standards, no lock-in

Write directly to Apache Iceberg or Parquet with support for multiple catalogs including AWS Glue, Apache Hive Metastore, and REST catalogs like Nessie, Polaris, and Unity Catalog — your data stays queryable by any engine you choose.

YOUR-VPCAPC SOUTH 1
Postgress

res

OLake

container

s3://lakeiceberg

Controlled

Self-hosted, on your infrastructure

Deploy OLake entirely within your own cloud or on-prem environment, keeping full control over where your data lives — critical for regulated industries like financial services that can't compromise on data residency.

/orderlast 90 days

P95 Query Time

0.9s

Table Size

12.4 TB

Maintained

Keep tables fast as data keeps growing

Automate Iceberg table maintenance through compaction, clearing delete files, and trimming metadata so query performance and storage costs don't degrade as your lakehouse scales.

Architecture

Works with what
you run

OLake Go supports ingestion from 8 different sources into Iceberg and Parquet. OLake Fusion keeps your Iceberg tables fast, through scheduled compaction and maintenance.

Get in touch

Sources

  • Postgres
  • MySQL
  • MongoDB
  • Oracle
  • Kafka
  • S3
  • DB2 LUW
  • MSSQL

Catalogs

  • AWS Glue
  • Hive
  • JDBC
  • REST
  • Lakekeeper
  • Nessie
  • S3 Tables
  • Unity
  • Apache Polaris
  • BigLake

Destinations

  • Apache Iceberg
  • Parquet on S3

Query Engines that read results

  • Amazon Athena
  • Apache Spark
  • Apache Flink
  • Trino
  • Starburst
  • Presto
  • Apache Hive
  • Apache Impala
  • Dremio
  • DuckDB
  • ClickHouse
  • StarRocks
  • Apache Doris
  • BigQuery
  • Snowflake
  • Databricks
  • AWS Redshift
  • Azure Synapse
Eight sources flow into OLake Go, which writes Iceberg tables and Parquet on S3; OLake Fusion optimises those tables.
Latest ReleaseWe just released our OLake v0.10.1Read More

Latest News

OLake Bulletin

For Enterprises

Want to bring OLake to your enterprise?

While we are open source, we offer custom support for enterprises based on your GRC requirements. Leave your contact information and our team will get in touch.

Request a callback