Cloudera Data Services on Premises 1.5.5 SP3 Release Summary

This release summary for Cloudera Data Services on premises 1.5.5 SP3 introduces significant enhancements across the Cloudera Data Services on premises platform, focusing on the Cloudera Control Plane, Cloudera Data Warehouse, Cloudera Data Engineering, Cloudera AI, and Cloudera Observability.


Cloudera Control Plane

  • Hybrid Environment (Technical Preview): You can connect your on-premises Cloudera Control Plane with the cloud Cloudera Control Plane to set up cloud bursting. Cloud bursting helps you maximize existing on-premises investment by dynamically extending the private data center into the cloud when on-premises resource utilization approaches capacity. For more information, see the Hybrid Cloud documentation.
  • Tainted Node Support: Cloudera Embedded Container Service now automatically applies the Kubernetes taint and label when you select a node profile in Cloudera Manager. You do not have to run the kubectl commands manually to map profiles. Selecting a profile (such as GPU, NVMe, CDE, or CAI) in Host > Configuration and running Refresh ECS applies both the taint and the label in a single step. For more information, see Dedicating Cloudera Embedded Container Service nodes for specific workloads.
  • Administrator Policy Changes: Updated default administrator username and password policies enhance system security across Cloudera Control Plane systems. For more information, see Installing Cloudera Data Services on premises using Cloudera Embedded Container Service.
  • Logging and Diagnostic Bundle Scaling: This release resolves scaling issues in the logging and diagnostic bundle collection pipeline by introducing dynamic scalability to manage a large volume of logs. For more information, see Scaling issues pertaining to Logging and Diagnostic bundle collection.
  • Diagnostic Bundle Management: You can automatically delete old Apache Ozone logs in S3 based on retention, complete scheduled cleanup of diagnostics bundles, and collect bundle data. By deploying the diagnostics-api-service, you can use environment variables to schedule Ozone log clearing and configure thresholds. For more information, see Diagnostics bundle management.
  • Proxy Protocol for ECS: You can resolve client IP address masking in Cloudera Data Services on premises deployments on the Cloudera Embedded Container Service. For more information, see Enabling Proxy Protocol for Embedded Container Service installation.
  • S3-Compatible Credentials Management: Administrators can now centrally manage shared, S3-compatible object storage credentials under Shared Resources in the Cloudera Management Console. Workloads can securely reuse these credentials across your entire deployment. For more information, see Managing AWS S3 compatible credentials for Cloudera Data Services on premises.

Cloudera AI

General Enhancements

  • Cloudera AI Registry In-Place Upgrade: You can upgrade Cloudera AI Registry in place on the existing release within the same namespace in less than 15 minutes, preserving all Kubernetes resources. For more information, see Cloudera AI Registry upgrade.
  • Cloudera AI Inference Service Connecting to Multiple Registries: Cloudera supports deploying a Cloudera AI Inference service that connects to multiple Cloudera AI Registries. For more information, see Multiple Cloudera AI Registries connected to a single Cloudera AI Inference service.
  • API Endpoint for MLFlow Downloads: A new API endpoint enables you to download a model artifact for a specific model version from the Cloudera AI Registry for models using the MLFLOW repository type. For more information, see Downloading MLFlow models using API endpoint.
  • Fine-Grained Access Control: Fine-grained access control enables administrators to define specific access levels for Model Endpoints for individual users or groups. For more information, see Configuring fine-grained access control for Model Endpoints.
  • Quota Management (GA): Quota Management is generally available for fresh workbench installations to control resource allocation on user and team levels, introducing a new alert for insufficient quota.
    • New alert for insufficient quota: When creating a session, job, or application, an alert notifies the user if the configuration does not have sufficient quota.
  • User ID Decoupling: Because of dynamic pool assignment, UserIDs no longer map predictably to user namespace sequence numbers, and the ds-username label is excluded from new user namespaces.
  • Diagnostic Bundles: You can generate and download User logs from the Diagnostic page in the Site Administration settings. For more information, see Downloading diagnostic bundles for a workbench.
  • Automated Application Restart: Failed applications automatically restart up to three times with a five-minute delay between attempts to minimize downtime. Site Administrators can disable this feature globally. For more information, see Disabling global Application restarts.
  • Refined Project List Filters: The Projects list page includes a refined My Projects filter for explicitly owned projects, moving shared work to a My Team Projects filter for collaborators.
  • Self-Service “Run As” Service Account Assignment: Project Contributors can assign service accounts holding an Operator role to workloads without administrator intervention. For more information, see Creating a Workload as a Contributor.
  • Knox API Key Support: The Cloudera AI Inference service accepts Knox API keys to enable long-lived connectivity to Model Endpoints on Cloudera Runtime 7.1.9 SP2 and 7.3.2 versions.
  • Hardened Chainguard ML Runtimes: This release introduces new Hardened Edition ML Runtimes based on Chainguard images. These runtimes are designed to meet strict security standards, provide enhanced protection for your workloads, and are released behind a paywall. Hardened Runtime workloads do not use the Java version provided by the Hadoop Runtime add-on. Instead, they rely on the Java 17 installation included in the Runtime image. A new Hardened Edition option is added to the Runtime Catalog page.
  • Hadoop Runtime Add-on Upgrades: Hadoop Runtime add-on images are upgraded to Ubuntu 24.04 LTS.

Agent Studio

  • General Availability: Cloudera AI Agent Studio is now generally available. It provides a versatile low-code to high-code platform for building, testing, and deploying multi-agent workflows.
  • System Tools Framework and LLM-Native Agent Planning: Agent Studio now uses a system tools framework enabling agents to use structured, in-process tools that are not user-defined. The Agent now creates a To-do list before running a specific task.
  • Auditing of Agent Studio: Agent Studio now provides built-in auditing at two distinct levels:
    • For resource management, the platform automatically records all creation, modification, and deletion actions for workflows, agents, tools, and MCP servers.
    • These activities are captured from authenticated sessions and logged with the practitioner username in a structured audit log. Workflow changes are tracked specifically within the Application tab.
  • Guardrails in Agent Studio: Guardrails are now implemented on prompt inputs, LLM outputs, and tool or Model Context Protocol (MCP) server outputs to protect confidential data from being sent to the LLM.
  • Hardened Chainguard Runtimes: This release introduces new Hardened Edition ML Runtimes based on Chainguard images. These runtimes are designed to meet strict security standards, provide enhanced protection for your workloads, and are released behind a paywall. Hardened Runtime workloads do not use the Java version provided by the Hadoop Runtime add-on. Instead, they rely on the Java 17 installation included in the Runtime image.
  • Streamlining Agent Workflows by Task Planning: Task Planning is now implemented as a workflow feature that enables an agent to manage complex, multi-step projects by generating its own structured to-do list during a run. Task planning provides real-time visibility into agent progress.
  • Native Function Calling: Agent Studio now runs all agents and workflows on native function calling, also known as tool calling. This replaces the traditional text-based parsing methods to achieve deterministic, reliable execution of enterprise AI agents.
  • Shared Responsibility Security Model: The shared responsibility model for security in Cloudera AI Agent Studio details the security controls provided by Cloudera by default and the specific security configurations that administrators must manage to ensure a secure environment.

Cloudera Data Engineering

  • Tainted Node Support: Cloudera Embedded Container Service automatically applies the Kubernetes taint and label when you select a Cloudera Data Engineering node profile in Cloudera Manager, eliminating running the kubectl commands manually. For more information, see Dedicating Cloudera Embedded Container Service nodes for Cloudera Data Engineering.
  • Seamless Configuration Updates: Managing configuration changes for existing Services and Virtual Clusters (VCs) is simplified. When you update settings, such as proxy, LDAP, or Kerberos, an interactive banner is displayed within the affected service, letting you apply changes directly with a single click. This eliminates the need for manual configuration or service redeployment. For more information, see Updating and synchronizing stale configurations.
  • S3-Compatible Credential Management: You can register shared, S3-compatible object storage credentials that workloads can reuse across your deployment. For more information, see Connecting to S3 compatible accounts for Cloudera Data Engineering jobs or sessions.
  • Cloudera Base Component Support: Cloudera Data Engineering supports connections to Cloudera Base on-premises components, including Apache Kafka, HBase, Apache Phoenix, and Apache Ozone. For more information, see Adding Kafka data connector for Cloudera Data Engineering.
  • External IDE Connectivity (GA): External IDE connectivity through Spark Connect-based sessions is generally available for JVM clients such as Java and Scala. This release increases session timeouts up to a maximum of 90 days, enables dynamic resource allocation, and increases the allowed size for uploaded JAR artifacts up to 200 MB. For more information, see External IDE connectivity through Spark Connect-based sessions.
  • User Access Management for Airflow Jobs: Only users with administrator privileges, such as DEAdmin, Service Admin, and VC Admin, can manage connections, variables, and XCOM in the Airflow UI. All users can create or manage Airflow jobs. For more information, see Access roles in Cloudera Data Engineering.
  • Airflow Package Removals due to CVE Risks: The apache-airflow-providers-google (CVE-2026-27459) and apache-airflow-providers-snowflake (CVE-2025-50213) default packages are removed due to identified CVE risks. If your DAGs rely on these packages, you must either install the required packages manually using Airflow custom providers or update your DAG code to replace references. For more information, see Supported Airflow operators and hooks.
  • End of Support Announcement: CDS 3.3 (Cloudera Distribution of Apache Spark) in Cloudera Base on premises 7.1.9 is deprecated and ends support in August 2026. Cloudera recommends upgrading your platform to a version that includes Spark 3.5 support and migrating workloads to ensure continued support. For more information, see Cloudera Data Engineering Runtime end of support for Spark.

Cloudera Data Warehouse

  • Configuration Change Detection: Cloudera Data Warehouse introduces a new configuration change detection feature to enhance cluster stability, manageability, and security tracking. Administrators can proactively monitor and identify unexpected or impactful adjustments across essential services. This feature reduces the risk of errors, streamlines troubleshooting, and strengthens compliance and operational security. For more information, see Configuration change detection.
  • Automated Apache Ozone Storage Integration: This release introduces automated integration with Apache Ozone storage through the S3A protocol, eliminating manual configuration. Credentials and endpoints are managed centrally through the Cloudera Management Console user interface (UI) and automatically propagate to all relevant Virtual Warehouse components and query engines, including Apache Hive, Apache Impala, and Trino. For more information, see Automated Apache Ozone storage integration.
  • Rollback for Virtual Warehouse Upgrades: You can roll back a failed Hive or Impala Virtual Warehouse upgrade to restore it to its last known working state. If an upgrade fails and the Backup Virtual Warehouse namespaces before an upgrade is enabled in Advanced Configuration, a Rollback Upgrade option is displayed in the Virtual Warehouse actions menu. After you confirm the rollback, the Virtual Warehouse enters a RollingBack state and Cloudera Data Warehouse automatically reverts it to the pre-upgrade configuration. This helps reduce downtime and manual intervention, providing a reliable recovery path directly from the UI. For more information, see Rolling back of Virtual Warehouses.
  • Enhanced JVM Crash and Heap Dump Collection: A unified solution standardizes and expands diagnostic coverage by automatically capturing both Out-Of-Memory (OOM) heap dumps and Java Virtual Machine (JVM) crash dumps across all Hive, Impala, and Trino Virtual Warehouses. Collected files are automatically forwarded to the diagnostic storage location and included in generated diagnostic bundles. This ensures that vital diagnostic information is immediately available without manual pod inspection, reducing the time needed to troubleshoot JVM-related failures. For more information, see JVM crash and heap dump collection.
  • Migration from NGINX to Istio: For the RKE platform, Cloudera Data Warehouse transitions from the NGINX Ingress Controller to the Kubernetes Gateway API by using an Istio gateway controller. This change currently affects only the RKE platform; Red Hat OpenShift Container Platform (OCP) environments are not affected. For new and upgraded workloads on 1.5.5 SP3, the ingress layer is managed by Istio, utilizing gateway resources instead of traditional ingress resources. Consequently, workloads in 1.5.5 SP2 or lower versions cannot be deployed in these updated environments. The NGINX Ingress controller continues to manage ingress traffic for existing workloads until you upgrade them. Upgrading the workload automatically upgrades the underlying networking stack to the Istio Gateway API. If necessary, administrators can disable Gateway API networking on the server settings page to revert to the NGINX Ingress approach.
  • Removal of Unified Analytics: The Unified Analytics framework, including the Impala Virtual Warehouse implementation, is fully removed from Cloudera Data Warehouse. All remaining components, configuration paths, and UI flows are removed, and Virtual Warehouses automatically revert to the standard Impala Virtual Warehouse architecture. Saved queries in Cloudera Data Explorer (Hue) originally created as a Hive snippet type in Unified Analytics will fail with a configuration error. Migration steps are available to update and resolve these stored queries. For more information, see Migrating stored Hive queries after Unified Analytics removal.
  • UI Renaming Updates: The Edit option in the action menu for Environments, Virtual Warehouses, and Database Catalogs is renamed to Details. This change is a UI label update and has no functional impact.
  • Cloudera Data Visualization Upgrade: The bundled Cloudera Data Visualization component is updated to version 8.1.1-b65, bringing the latest improvements and bug fixes.

Cloudera Data Explorer (Hue)

  • Product Branding Update: The product component previously known as Data Explorer is renamed to Cloudera Data Explorer (Hue). This change reflects updated branding initiatives and rolls out in phases. You might notice a new logo, updated service names, and new documentation titles. Some UI references might still display the previous name until future updates are complete. This update has no functional impact.
  • Fact Support in SQL AI Assistant: You can define custom system instructions to guide the SQL AI Assistant in generating more accurate queries based on your specific business logic. This enhancement supports complex, cross-database workflows by letting you persist organizational context in the assistant settings. For more information, see Facts support for SQL query.
  • Boto3 SDK Support: Data Explorer supports the boto3 SDK for accessing AWS S3, replacing the legacy connector framework to improve performance and compatibility with AWS services. The system automatically converts your existing configurations. You can manually disable the feature flag if necessary. For more information, see Enabling the S3 File Browser for Cloudera Data Explorer (Hue) in Cloudera Data Warehouse with RAZ.

Hive

  • Small File Warnings: The MSCK and ANALYZE commands display a warning in the console if the average file size for a table or partition falls below the threshold, helping you identify small files that might affect performance. For more information, see Statistics generation and viewing commands in Cloudera Data Warehouse.
  • Performance Improvement for Column Changes: The ALTER CHANGE COLUMN command runs faster for tables with many partitions. This change prevents the command from performing a separate Metastore service call to update column statistics for every partition. For large partitioned tables, execution time reduces from hours to minutes (Apache Jira: HIVE-28346).

Iceberg

  • Table Repair Support: Impala introduces the repair_metadata() function for Iceberg tables to provide a self-service recovery path for tables that become inaccessible due to missing data files after manual file deletions in the underlying storage. For more information, see Table repair feature.
  • Partition File Inspection: Impala supports the SHOW FILES IN command with the PARTITION clause to list data files for specific partitions in Iceberg tables, enabling inspection of partition-level physical data directly from Impala. For more information, see Describe table metadata feature.
  • Additional Partition Transform Functions: Iceberg supports additional partition transform functions, including BUCKET, TRUNCATE, IDENTITY, and VOID, extending partitioning capabilities to support hashing, value truncation, direct partitioning, and null partition management. For more information, see Partition transform feature.
  • WHERE Clause Predicates for Compaction: Hive Iceberg compaction supports WHERE clause predicates on partition columns, letting you compact data selectively to improve efficiency and control. For more information, see Data compaction.

Impala

  • Caching Intermediate Query Results: Impala supports caching intermediate results to improve query performance and resource efficiency for repetitive workloads. By storing results within the SQL plan tree, the system reuses computations for similar queries when the underlying data and settings remain unchanged. For more information, see Caching intermediate results.
  • User Role Management: You can grant and revoke roles directly to and from individual users in Impala, providing granular security control. This feature supports GRANT ROLE, REVOKE ROLE, and SHOW ROLE GRANT USER statements, aligning Impala with Apache Hive functionality (Apache Jira: IMPALA-14085). For more information, see Impala role.
  • Native Geospatial Query Acceleration: This release introduces native implementations for specific geospatial functions to accelerate simple queries. This feature reduces processing overhead by avoiding transitions to the Java Virtual Machine and optimizing file-level filtering for Parquet and Iceberg tables. For more information, see Impala Geospatial query acceleration.
  • OpenTelemetry Integration: Impala provides OpenTelemetry support to help you monitor query performance and troubleshoot issues by collecting and exporting query telemetry data as traces to a central compatible collector (Apache Jira: IMPALA-13234). For more information, see OpenTelemetry support for Impala.
  • Filtering SHOW PARTITIONS Output: You can use the WHERE clause with the SHOW PARTITIONS statement to filter results based on partition column values, helping you manage tables with a large number of partitions (Apache Jira: IMPALA-14065). For more information, see the SHOW PARTITIONS statement.
  • Parallel JDBC External Table Queries: You can run queries on JDBC tables in parallel to improve performance for joins and aggregations. Impala estimates the number of rows by running a COUNT query during preparation, which helps the planner assign multiple scanner threads and produce efficient join orders. You can use the --min_jdbc_scan_cardinality backend flag to set a lower bound for these estimates. For more information, see Parallelizing JDBC External Table queries.
  • Recreating Tables with Statistics: You can use the WITH STATS clause in the SHOW CREATE TABLE statement to generate the SQL required to recreate a table with its column statistics and partition metadata (Apache Jira: IMPALA-13066). For more information, see SHOW CREATE TABLE WITH STATS statement.
  • Quoting Reserved Words: You can explicitly quote all column names projected in SQL queries generated for JDBC external table names using backticks (`) for Cloudera Runtime Hive, Impala, and MySQL, or double quotes (“) for other databases, supporting case-sensitive or reserved column names (Apache Jira: IMPALA-13066). For more information, see Quoting reserved words in column names.
  • New Catalogd Startup Flag: The disable_hms_sync_by_default catalogd startup flag sets a global default for the impala.disableHmsSync property, letting you skip event processing for all databases and tables by default while opting in specific elements (Apache Jira: IMPALA-14131). For more information, see Catalogd Daemon startup flag.
  • Specifying Compression Levels: You can specify compression levels for the LZ4, ZLIB, GZIP, and ZSTD codecs by using the compression_codec query option to achieve higher compression ratios, supporting high compression modes in LZ4 (levels 3 through 12) and negative levels for ZSTD (Apache Jira: IMPALA-10630, IMPALA-14082). For more information, see compression_codec query option.
  • Batch Processing for Reload Events: Catalogd supports batch processing of RELOAD events on the same table to load partitions in parallel and reduce duplicate reloads, improving coordinator performance and reducing query planning retries (Apache Jira: IMPALA-14082).
  • Consolidated Event Processing: Catalogd supports the ALTER_PARTITIONS event type, which consolidates multiple partition changes into a single event to synchronize metadata quickly and reduce the load on the Catalogd cache (Apache Jira: IMPALA-13593).

Trino

  • Improved Trino Connector UI Labels: The Trino connector UI is updated. Connectors previously marked as Optimized are now labelled Certified. The Data source column label is renamed to Catalog Name.
  • New Federation Connectors: This release introduces the Hive Engine and Impala Engine to the federation connectors list, providing dedicated JDBC-based connectivity options.
  • Teradata General Availability: The Trino-Teradata connector is generally available, supporting read-only SELECT operations on Teradata sources in ANSI Mode.
  • Trino Runtime Upgrade: The Trino runtime is upgraded from version 476 to 479 to introduce stability improvements and optimizations.

Cloudera Observability

  • Telemetry Export to Third-Party Destinations: You can configure the OpenTelemetry pipeline to duplicate and route your collected metrics and logs to external, third-party destinations. To manage these export destinations without modifying your core collector configuration, you must define your settings in the ObservabilityPipeline Custom Resource and Kubernetes Secrets.