5 posts tagged with "Apache Spark"
Blogs on Apache Spark
View All TagsThe Iceberg Interoperability Myth: Why Open Table Format Doesn't Mean Write Once, Read Anywhere
See how row-level delete support can make an Iceberg table readable in Spark but incompatible with Snowflake, even when both engines use the same catalog and data.
Top 5 Debezium Alternatives of 2026
Discover the best Apache Spark alternatives in 2026. Compare Databricks, Flink, Ray, Dask, DuckDB, Polars, and Trino on performance, cost, and use cases.
Apache Spark Alternatives in 2026
Discover the best Apache Spark alternatives in 2026. Compare Databricks, Flink, Ray, Dask, DuckDB, Polars, and Trino on performance, cost, and use cases.
Apache Iceberg Row Lineage: Tracking Data Lineage at the Row-Level
How Apache Iceberg v3 row lineage tracks row-level changes for CDC, with a tested look at _row_id preservation across Spark 3.5 and Iceberg 1.9 vs 1.10.
50% Cheaper (2x Faster) Iceberg Compaction: OLake Fusion (Open Source) Beats Spark
We benchmark Spark rewrite_data_files against OLake Fusion compaction on Apache Iceberg by running a full TPCH lineitem load from Postgres to GCP, applying 200k-record CDC batches every 2 minutes, and tracking TPC-H Query 6 performance, runtime, resource usage, and infrastructure cost.







