What's New in Spark

Learn about the new features of Spark in Cloudera Runtime 7.3.2, its service packs and cumulative hotfixes.

Cloudera Runtime 7.3.2.10000 SP1

Rebase on Apache Spark 3.5.8

Spark shipped with Cloudera Runtime 7.3.2.10000 SP1 is based on Apache Spark 3.5.8 (previously 3.5.4 in Cloudera Runtime 7.3.2). This maintenance rebase includes bug fixes and security updates from upstream Apache Spark. For more information, see the Apache Spark documentation.

FIPS 140-3 support for Spark 3

Spark 3 supports FIPS 140-3 on FIPS-enabled clusters in this service pack. Spark cryptographic modules and Cloudera SASL are updated for FIPS 140-3 validation.

For configuration details, see Spark and Livy on FIPS 140-3 clusters.

Encryption in transit for Spark

Spark 3 supports stronger encryption in transit when you enable Encrypt all ports in Cloudera Manager, including Spark network encryption, HTTPS-only History Server and UI listeners, and updated IO encryption defaults. Spark History Server and UI TLS can use TLS 1.3 when your Cloudera Manager TLS policy allows it.

For configuration and behavioral changes, see Encryption in transit for Spark and Livy in Cloudera Runtime 7.3.2.10000 SP1, Configuring TLS/SSL encryption manually for Spark, and Behavioral Changes in Spark.

Upgrade Spark Iceberg library to Apache Iceberg 1.10.0
In Cloudera Runtime 7.3.2 SP1, the Apache Iceberg library used by Spark is upgraded to Apache Iceberg 1.10.0. Existing Iceberg V1 and V2 table operations are unchanged; you need not run new SQL syntax for current workflows.

This upgrade aligns Spark with the platform-wide Iceberg 1.10.0 library and provides the foundation for Iceberg V3 capabilities in Spark. For supported Iceberg features, see the Apache Iceberg Feature Support Matrix.

For Spark-specific limitations, see Prerequisites and limitations for using Iceberg in Spark.

Spark Iceberg V3 deletion vector support [Technical Preview]

Spark supports deletion vectors on Iceberg V3 tables. Deletion vectors are the Iceberg V3 encoding for position deletes and replace V2 position delete files for new delete operations on V3 tables. When you run DELETE on a V3 Iceberg table from Spark, Spark records deleted row positions using deletion vectors.

For more information, see Prerequisites and limitations for using Iceberg in Spark and the Apache Iceberg Feature Support Matrix.

Spark Iceberg V3 row lineage support [Technical Preview]

Spark supports row lineage tracking on Iceberg format V3 tables. Row lineage is an Iceberg V3 capability that assigns a unique identifier to each row and records row provenance across table snapshots. This metadata supports change data capture (CDC) use cases. To use row lineage from Spark, create an Iceberg V3 table by setting 'format-version'='3'. When Spark writes to the table, Spark maintains the row lineage metadata required; you need not run new SQL syntax to enable row lineage.

For more information, see Prerequisites and limitations for using Iceberg in Spark and the Apache Iceberg Feature Support Matrix.

Iceberg V3 tables are not supported for replication
Replication Manager Iceberg replication policies support Iceberg V1 and V2 tables only. Iceberg V3 tables created or written by Spark are not supported for replication in Cloudera Runtime 7.3.2 SP1.

Cloudera Runtime 7.3.2.100 CHF 1

There are no new features introduced in this release.

Cloudera Runtime 7.3.2.0

Rebase Spark3 to Apache Spark 3.5.4 in Cloudera Runtime.

Spark 3.5.4 is the default Spark version in Cloudera Runtime. Refer to the Apache Spark documentation for more information on changes.

Spark 3 is the default in Cloudera Runtime

Spark 3 is the default Spark version in Cloudera Runtime. Spark 2 has been removed and no longer available in 7.3.2.0