Tag:Apache Iceberg
Articles tagged "Apache Iceberg", page 3.
- 30 MIN READ•Sep 2, 2026
The Iceberg Table Properties That Actually Matter
The Iceberg table properties that decide file count, pruning, write amplification, retention, and metadata growth, by workload.
Apache IcebergTable PropertiesTuning - 33 MIN READ•Sep 2, 2026
Inside the Puffin File Format
The Puffin file format inside out, byte by byte, covering Theta sketches for distinct values and deletion vectors.
Apache IcebergPuffinFile Format - 30 MIN READ•Sep 2, 2026
The Lakehouse Ingestion Tool Landscape: Fivetran, Airbyte, dlt, and CDC vs Batch
How Fivetran, Airbyte, dlt, and CDC and streaming tools land well-behaved Apache Iceberg tables, and how to choose and maintain them.
IngestionCDCFivetran - 30 MIN READ•Sep 2, 2026
Local Iceberg Development Environments: Docker, MinIO, and In-Memory Catalogs for CI
Local Iceberg development environments: in-process catalogs, a Docker Compose stack with MinIO, and CI configurations that run either.
Apache IcebergLocal DevelopmentMinIO - 31 MIN READ•Sep 2, 2026
Moving Iceberg Tables Between Catalogs Without Rewriting Data
Why moving Iceberg tables between catalogs is a pointer copy, and the protocol that makes a cutover safe for one table or thousands.
Apache IcebergCatalogsMigration - 30 MIN READ•Sep 2, 2026
Logs, Traces, and Metrics as Tables: Building an OpenTelemetry Data Lake on Iceberg
Building an OpenTelemetry data lake on Iceberg: schemas for spans, logs, and metrics, ingestion, query patterns, and retention.
OpenTelemetryApache IcebergLogs - 30 MIN READ•Sep 2, 2026
Schema Registries and Event Schemas: Avro, Protobuf, and JSON Schema on the Way Into the Lakehouse
How Avro, Protobuf, and JSON Schema evolve through a registry, and how that maps to the schema evolution rules of Iceberg.
Schema RegistryAvroProtobuf - 30 MIN READ•Sep 2, 2026
Storage-Partitioned Joins and the Bucket Transform
How the spec-defined bucket transform lets engines skip the shuffle in joins, and how to set it up and keep it engaged in Spark.
Apache IcebergBucketingQuery Planning - 31 MIN READ•Aug 25, 2026
Agent-Driven Storage Tiering for Apache Iceberg: Moving Cold Data Without Breaking Queries
A background agent can move cold Iceberg partitions to cheaper tiers without breaking live queries. Heatmaps, path-safe moves, and restore paths.
Apache Icebergstorage tieringcost optimization - 31 MIN READ•Aug 25, 2026
DataFusion Comet 1.0 and What Native Rust Scans Change for Spark on Iceberg
DataFusion Comet 1.0 replaces Spark Iceberg scans with native Rust. What speeds up, what still falls back to the JVM, and how to deploy it.
Apache IcebergApache SparkDataFusion Comet - 31 MIN READ•Aug 25, 2026
High-Throughput Branch Merging: Automating Concurrency and Conflict Resolution in Multi-Branch Iceberg Pipelines
High-throughput Iceberg branch merges need conflict detection and automation. How to reconcile concurrent writes without stalling pipelines.
Apache Icebergbranchesconcurrency - 31 MIN READ•Aug 25, 2026
Multi-Cloud REST Catalog Topologies: Running Apache Polaris Across AWS, Azure, and GCP
Polaris can catalog Iceberg tables across AWS, Azure, and GCP. Four topologies, credential vending, and the tradeoffs of each design.
Apache PolarisREST catalogmulti-cloud