Configure metadata cleanup, expire snapshots, compact files, and remove orphans from
Spark to limit metadata growth and storage use on Iceberg tables.
Invoke stored procedures with CALL against
<catalog>.system.<procedure>, for example
spark_catalog.system.expire_snapshots.
Each write to an Iceberg table creates a snapshot. Snapshots support time travel
and rollbacks, but they accumulate over time. Without maintenance, snapshot
metadata, manifest files, and data files can increase storage cost and planning
time.
Maintenance includes one-time or per-table configuration (metadata cleanup
properties) and periodic steps such as compaction, snapshot expiration, and orphan
file removal.
Cloudera recommends that you run maintenance on a schedule. The
expire_snapshots procedure removes snapshots you no longer need
and deletes files that only those snapshots reference. Compaction procedures such as
rewrite_data_files help make old data files eligible for
deletion when you expire snapshots.
Consult the Apache Iceberg documentation for the version shipped with your CDP
release. For example, CDP Private Cloud Base 7.1.9 includes Iceberg 1.3.
-
Inspect snapshot history and note snapshot timestamps.
spark.sql("SELECT * FROM default.my_table.history").show(false)
-
Configure table properties to remove old metadata JSON files after maintenance
commits.
This step is largely independent of compaction and snapshot expiration. Set
these properties before your first expire_snapshots run so
that subsequent commits, including that maintenance pass, can prune obsolete
metadata.json files that Iceberg still tracks.
Expiring snapshots does not remove old metadata.json
files by default. Set
write.metadata.delete-after-commit.enabled=true and
write.metadata.previous-versions-max as described in
Iceberg table properties. To prevent expiration of
recent snapshots, set
history.expire.min-snapshots-to-keep.
Property-based cleanup does not delete untracked metadata files left after
failed writes; use orphan file removal for those files.
-
As needed, compact data files, position delete files, or manifests before you
expire snapshots.
Run one of the following procedures when you need to reduce small files,
reorganize manifests, or make obsolete data files unreachable:
rewrite_data_files — compacts data files and can
incorporate delete files
rewrite_position_delete_files — compacts position
delete files
rewrite_manifests — rewrites manifest files for better
query planning
Compaction can temporarily increase file counts because Iceberg writes new
files before you expire old snapshots.
spark.sql("CALL spark_catalog.system.rewrite_data_files(
table => 'default.my_table',
options => map('rewrite-all', 'true'))").show()
-
Expire snapshots older than a timestamp by using the
older_than argument.
Do not use snapshot_ids to expire a range of snapshots.
snapshot_ids removes only the specific snapshot IDs you
pass and must not include the current snapshot.
spark.sql("CALL spark_catalog.system.expire_snapshots(
table => 'default.my_table',
older_than => timestamp'2024-10-31 02:49:30.000')")
-
Review procedure output and snapshot history to confirm the maintenance
run.
The expire_snapshots procedure returns counts such as
deleted_data_files_count,
deleted_manifest_files_count, and
deleted_manifest_lists_count.
-
As needed, remove orphan files that are not referenced by any table
metadata.
Orphan files can remain after failed writes or other operations. The Spark
SQL procedure remove_orphan_files corresponds to the
Iceberg Spark action deleteOrphanFiles. Run it on a
schedule or after incidents, separate from snapshot expiration.
By default, the procedure considers files older than about three days as
orphan candidates. Use dry_run => true first to list
candidates without deleting them. For arguments and behavior, see delete orphan files in the Apache
Iceberg maintenance guide and remove_orphan_files in the Spark
procedures reference.
spark.sql("CALL spark_catalog.system.remove_orphan_files(
table => 'default.my_table',
dry_run => true)")
Iceberg removes only files that are no longer required by any non-expired snapshot.
The effect depends on what changed before you expire snapshots:
- After inserts only — Expiring snapshots with
older_than
typically removes manifest list files for the expired snapshots. Data and
manifest files often remain because newer snapshots still reference
them.
- After compaction and later writes — Expiring snapshots older than a compaction
can delete data files, manifest files, and manifest lists that are no longer
referenced.
If storage file counts do not decrease immediately after
expire_snapshots, run rewrite_data_files first.
Then expire snapshots older than the compaction timestamp.
For syntax details, optimization guidance, and Impala equivalents, see the
following resources: