4 posts tagged with "Apache Spark"
Blogs on Apache Spark
View All TagsTop 5 Debezium Alternatives of 2026
Discover the best Apache Spark alternatives in 2026. Compare Databricks, Flink, Ray, Dask, DuckDB, Polars, and Trino on performance, cost, and use cases.
Apache Spark Alternatives in 2026
Discover the best Apache Spark alternatives in 2026. Compare Databricks, Flink, Ray, Dask, DuckDB, Polars, and Trino on performance, cost, and use cases.
Apache Iceberg Row Lineage: Tracking Data Lineage at the Row-Level
How Apache Iceberg v3 row lineage tracks row-level changes for CDC, with a tested look at _row_id preservation across Spark 3.5 and Iceberg 1.9 vs 1.10.
50% Cheaper (2x Faster) Iceberg Compaction: OLake Fusion (Open Source) Beats Spark
We benchmark Spark rewrite_data_files against OLake Fusion compaction on Apache Iceberg by running a full TPCH lineitem load from Postgres to GCP, applying 200k-record CDC batches every 2 minutes, and tracking TPC-H Query 6 performance, runtime, resource usage, and infrastructure cost.





