Known Issues in Ozone

Known issues and technical limitations for Ozone are addressed in Cloudera Runtime 7.3.2, its service packs, and cumulative hotfixes.

Known issues identified in Cloudera Runtime 7.3.2.20000 SP2

There are no new known issues in this release.

Known issues identified in Cloudera Runtime 7.3.2.10000 SP1

CDPD-105709: Recon UI displays inconsistent Solr health status
7.3.2.10000
The Recon Overview page incorrectly displays SOLR is HEALTHY as an alert when the service functions normally. Additionally, other pages within the Recon UI report SOLR is UNHEALTHY despite Solr remaining in a healthy state.
None
Apache JIRA: HDDS-15571
CDPD-109005: Storage Container Manager (SCM) fails to update DatanodeDetails when only port configurations change
7.3.2.10000
When a DataNode re-registers with SCM or sends heartbeats after gaining new ports, SCM fails to update the stored DatanodeDetails in the node registry. SCM triggers an update only if the IP address, hostname, or software version changes. Consequently, SCM ignores port-only updates, which leads to stale metadata and potential failures in features that rely on the new ports.
None
CDPD-109002: Ozone DataStream writes fail after enabling Ratis DataStream on existing clusters
7.3.2.10000
When Ratis DataStream is enabled on an existing cluster, open pipelines retain stale DataNode details that lack the required DataStream port information. Consequently, Ozone client writes fail when they exceed the automatic threshold because the client cannot identify the necessary port for pipeline members. This occurs because SCM does not refresh member ports for existing open pipelines after the configuration change.
None
Apache JIRA: HDDS-15799
CDPD-99722: Container data checksum reverts to a stale value after successful reconciliation
7.3.2.0 and 7.3.2.10000
In Ozone, the BackgroundContainerDataScanner service reverts the container data checksum to a stale value even after a successful reconciliation repairs corrupt chunks. This occurs because the scanner builds a checksum tree in memory while reconciliation updates the container data on disk, causing the scanner to persist an outdated tree.
The issue resolves automatically during the next container scan cycle. To resolve the issue immediately, manually trigger container reconciliation from the command line.
Apache JIRA: HDDS-14936
CDPD-105115: DataNodes fail to communicate with SCM after restart
7.3.2.10000
In a SCM High Availability (HA) cluster, a restarting or failing-over follower SCM starts its DataNode protocol server before completing its Ratis log replay. This causes SCM to process DataNode container reports against an incomplete state, which leads to dropped replica locations. Consequently, if that follower becomes the leader, affected containers appear to have missing or empty replicas, and DataNode logs record errors such as Unable to clear disk, Failed to flush DB before close, and EOFExceptions.
This issue is temporary and typically resolves during the next report cycle when SCM rebuilds the replica list.
CDPD-119561: Ozone Manager (OM) becomes unresponsive during sustained snapshot operations
7.3.2.10000
The OM leader enters a deadlock state during heavy snapshot creation and deletion loads. This occurs when the Key Deleting Service holds a snapshot database lock while waiting for a transaction that cannot complete because the flush thread is blocked. As a result, the OM stops responding to client requests and snapshot operations time out.
  • Restart the affected OM to release the locks and restore service.
  • Limit snapshot deletion and heavy write activity to reduce the risk of recurrence.
  • Increase the ozone.om.unflushed.transaction.max.count configuration to lower the probability of the deadlock (monitor for increased memory usage).
Apache JIRA: HDDS-16122
CDPD-104229: Delay in deleting snapshot YAML file by OM Snapshot Local Data Manager
7.3.2.10000
A significant delay of approximately 20 to 30 minutes occurs in the deletion of snapshot YAML files by the OM Snapshot Local Data Manager after the system processes a snapshot deletion request. This delay happens because the current design requires the inDegree, which represents the number of descendant snapshots, to reach zero before the system removes the YAML file. This behavior impacts subsequent operations, such as bootstrap test cases, that rely on the timely cleanup of these files.
None
CDPD-119610: OM leader remains in LEADER_AND_NOT_READY state during Snapshot Defragmentation Bootstrap
7.3.2.10000
The Ozone Manager (OM) leader remains in the LEADER_AND_NOT_READY state after a follower node stops for a planned bootstrap. This issue occurs due to a deadlock between the snapshot installation thread and the double-buffer flush thread on the follower node. Consequently, the cluster loses its usable quorum, and all OM client RPC requests fail with an OMLeaderNotReadyException.
Restart the affected OM follower node to break the deadlock and restore normal cluster operations.
CDPD-119172: OM follower node bootstrap fails when OM leader changes
7.3.2.10000
In an Ozone Manager (OM) High Availability (HA) setup, the bootstrap process for a follower node fails if the OM leader node changes while the bootstrap is in progress. If the original leader stops (for example, through SIGTERM), the bootstrap thread attempts to download the database snapshot from the former leader and fails with a Connection refused error. When the thread subsequently attempts to connect to the newly elected leader, it receives a 503 Service Unavailable error because the new leader is not yet ready to serve requests.
To ensure a successful bootstrap, verify that two of the three Ozone Managers (OMs) are running before bootstrapping the third node.
CDPD-102216: Toggling to Old UI from the Cluster Capacity tab results in a 404 error
7.3.2.10000
The Cluster Capacity feature is available only in the new Recon UI. If you view the Cluster Capacity page in the new UI and click Switch to Old UI, the application attempts to load the same path in the legacy interface. Because this path does not exist in the old UI, a 404 Page Not Found error appears.
Manually navigate to the main dashboard if you encounter a 404 Page Not Found error after switching to the old UI. Use the new UI to view cluster capacity information, as the legacy interface does not support this feature.
Apache JIRA: HDDS-15148
CDPD-106805: Requests to Ozone S3 Gateway through HAProxy fail with 400 Bad Request during directory creation
7.3.2.10000
When using the S3 connector to access Ozone OBS buckets through an S3 Gateway behind HAProxy, requests fail with a 400 Bad Request error. This issue occurs because zero-byte directory-marker PUT requests (which include Expect: 100-continue headers) cause the S3 Gateway to skip reading the request body and return 200 OK immediately. This leaves unconsumed bytes on the wire and desynchronizes keep-alive connections. HAProxy then misinterprets subsequent requests, causing a parsing error.
Disable keep-alive connections in the HAProxy configuration by setting option httpclose and increasing the maxconn value (for example, to 4096).
CDPD-107197: The ozone admin datanode diskbalancer status command fails when passing a DataNode ID instead of an address
7.3.2.10000
The ozone admin datanode diskbalancer status command accepts only a DataNode network address or host string. When you use a valid datanodeUuid (DataNode ID) to identify the node, the command fails with an UnknownHostException or an invalid argument error. This behavior is inconsistent with other ozone admin commands that support both UUIDs and addresses for node identification.
To check the disk balancer status, use the DataNode network address or hostname instead of its UUID. Run the ozone admin datanode list command to locate the network address.
Apache JIRA: HDDS-15690
CDPD-100716: Reading keys results in a NO_REPLICA_FOUND error after SCM follower restart and leader transfer
7.3.2.10000
When a Storage Container Manager (SCM) follower restarts, it accepts DataNode container reports before catching up with the Ratis log. Processing these reports against a stale database triggers a NotLeaderException, causing the SCM to skip updating in-memory container replica locations. If this follower is later promoted to leader, reading keys fails with a NO_REPLICA_FOUND error due to missing replica metadata.
None
Apache JIRA: HDDS-14989

Known issues identified in Cloudera Runtime 7.3.2.100 CHF 1

There are no new known issues in this release.

Known issues identified in Cloudera Runtime 7.3.2

OPSAPS-76062: Exposing "ozone.replication" in Cloudera Manager configurations and assigning a default value
7.3.2
Exposing "ozone.replication in Cloudera Manager configurations and assigning a default value to it is causing it to override bucket replication configuration as a client side configuration even when you does not mean to set the client side configurations.
None
CDPD-73792: The new snapshot could not be found after renaming the old snapshot
7.3.2
The ozone sh snapshot rename command renames snapshots. It is a feature developed by the Apache Ozone community and is included in Cloudera Base on premises 7.3.1. However, it does not work properly, and Cloudera Base on premises does not support it.
None
Apache JIRA: HDDS-11384
CDPD-63350: Force deleting a FSO bucket and its contents while running rb --force from AWS S3 API is failing
7.3.2
Force deleting a File System Optimized (FSO) bucket and its contents while running the rb --force command from the AWS S3 API might fail with an error, because the S3 client sends individual delete requests for each key in the bucket. It might delete or fail to delete individual keys or directories, depending on the availability of leaf elements. It can completely delete the bucket only when all the keys or directories are cleaned up. If some keys are not deleted, the bucket will not be deleted.
Sample error message
# aws s3 rm s3://buck-fso --recursive
delete: s3://buck-fso/dir1/
delete: s3://buck-fso/dir1/dir2/
delete: s3://buck-fso/dir3/dir4/dir5/
delete: s3://buck-fso/dir3/dir4/
delete failed: s3://buck-fso/dir3/
# Rerun the same command again
# aws s3 rm s3://buck-fso --recursive
delete: s3://buck-fso/dir1/
delete: s3://buck-fso/dir3/
# To Confirm
# ozone sh key list s3v/buck-fso
[ ]
Run the rb --force command multiple times to completely clean up the keys and directories.
Apache JIRA: HDDS-9637
CDPD-98751: Mismatched Replicas tab in the Recon UI fails to display containers with inconsistent replica checksums
7.3.2
In the Ozone Recon UI, the Mismatched Replicas tab does not update or display containers when one or more replicas have differing checksums. Instead, they are displayed in the Under-Replicated tab.
The replica can be checked in the Under-Replicated tab, or can be cross-checked against API response.
CDPD-97512: Mismatch in Open Key count between Overview page and OM DB Insight in the Recon UI
7.3.2
In the Ozone Recon UI, the Open Key count on the Summary section of the Overview page might not match the count on the OM DB Insight > Open Key tab. Users might see different Open Key values for the same cluster across these two views.
Use the OM DB Insight > Open Key tab for the most accurate Open Key count.
CDPD-97376: Container replication counts mismatch in Recon UI
7.3.2
In the Ozone Recon UI, container replication counts, including Under-Replicated, Over-Replicated, and Mis-Replicated counts, differ between the updated and the legacy Container page for the same cluster.
Refer to the legacy Container page to view accurate replication counts.
CDPD-97311: Incorrect Creation Time and Modification Time displayed on the Namespace Usage page in the Recon UI
7.3.2
In the Ozone Recon UI, the Namespace Usage page displays incorrect Creation Time and Modification Time timestamps. These values do not accurately reflect the actual creation or last modification times or dates of the namespace, resulting in inaccurate metadata information.
Retrieve the correct timestamp values directly from the API response.
CDPD-97312: Mismatch between cluster State Container count and Container Summary totals
7.3.2
The Ozone Recon UI can display discrepancies between the Storage Container Manager (SCM) container count and the Recon container summary. This occurs because Recon does not synchronize all container states such as QUASI_CLOSED leading to inconsistent totals.
No workaround within the Ozone Recon UI. Use ozone admin container report CLI command to obtain the correct container counts for all states in the Ozone cluster.
CDPD-99248: After cdh upgrade, Ozone encounters failure while running the Finalize Upgrade for SCM on role Storage Container command
7.3.2
When finalizing an Ozone upgrade for the first time from Cloudera Manager, the Finalize Upgrade for SCM on role Storage Container command might fail with the following stderr message:
"Invalid response from Storage Container Manager.
Current finalization status is: FINALIZATION_IN_PROGRESS"

This error occurs even though finalization continues to run on the Storage Container Manager (SCM).

Ignore the failure in Cloudera Manager. Use the ozone admin scm finalizationstatus command to monitor progress and wait for the process to complete on SCM.
CDPD-98892: File Size Distribution bucket size range calculation is not correct
7.3.2
If large number of buckets exists within a volume, the Ozone Recon UI might display incorrect file size distribution bucket size range calculation in the File Size Distribution chart on the Insights page.
None
Apache JIRA: HDDS-14827
CDPD-93116: Ozone client hangs intermittently when disks are full
7.3.2

This hang is caused by continuous write retries that persist until the pipeline on the Datanode closes.

The Ozone client hangs for approximately five minutes when writing data to a Datanode if a disk full exception or other Datanode error occurs. This hang is caused by continuous write retries that persist until the pipeline on the Datanode closes.
Control the request retry behavior by setting the following configurations on the client side:
Table 1. Client-side retry request configurations
Configuration Recommended value
hdds.ratis.raft.client.rpc.request.timeout 30s
hdds.ratis.client.multilinear.random.retry.policy 1s, 1
hdds.ratis.client.exponential.backoff.max.sleep 5s
hdds.ratis.client.exponential.backoff.base.sleep 1s
hdds.ratis.client.exponential.backoff.max.retries 2
Apache JIRA: HDDS-14040

Known Issues identified before Cloudera Runtime 7.3.2

Known issues identified before Cloudera Runtime 7.3.2 include only unresolved issues from previous releases that continue to affect the Cloudera Runtime 7.3.2 base release.

CDPD-91562: test_validate_certs_configs configuration is failing with the maximum lifetime validation
7.3.2, 7.3.1.600
In daylight saving time zones, the autogenerated Ozone certificate duration might differ from the expected duration. This discrepancy is minor, because the default certificate duration is 365 days or five years, depending on the Ozone component.
CDPD-75954: The ozone debug ldb command and ozone auditparser fails with java.lang.UnsatisfiedLinkError
7.3.2, 7.3.1.400
For information on workaround, see Changing temporary path for Ozone services and CLI tools.
CDPD-54885: Ozone Prometheus does not work with TLS
7.3.2, 7.3.1 and its SPs and CHFs
The Prometheus service shipped by Ozone does not support TLS mode. So, Prometheus is not able to gather metrics from Ozone endpoints when TLS is enabled.
Go to Ozone > Configuration > Ozone Prometheus Endpoint Token and in the Ozone Prometheus Endpoint Token property enter any random string. This configuration generates a plaintext token in the Ozone endpoint process directory allowing Prometheus to authenticate and collect metrics despite the TLS limitation.
CDPD-56684: Keys and buckets get deleted without volume permission
7.3.2, 7.3.1 and its SPs and CHFs
When a volume deletion is initiated, the system recursively deletes all buckets and keys within the volume before attempting to delete the volume itself. Because the ACL check for volume deletion permissions occurs only in the end, all the data within the volume is deleted even without having delete permission on the volume.
CDPD-50610: Large file uploads are slow with OPEN and stream data approach
7.3.2, 7.3.1 and its SPs and CHFs
Hue file browser uses the append operation for large files. This API is not supported by Ozone in 7.1.9, therefore large file uploads can be slow or can time out in the browser.
Use native Ozone client to upload large files instead of the Hue file browser.
OPSAPS-66469: Ozone-site.xml is missing if the host does not contain HDFS roles
7.3.2, 7.3.1 and its SPs and CHFs
The client side /etc/hadoop/conf/ozone-site.xml file is not generated by Cloudera Manager if the host does not have any HDFS role. Because of this, issuing Ozone commands from that host fails because it cannot find the service name to hostname mapping. When this issue occurs, an error message is displayed: # ozone sh volume list o3://ozoneabc 23/03/06 18:46:15 WARN ha.OMProxyInfo: OzoneManager address ozoneabc:9862 for serviceID null remains unresolved for node ID null Check your ozone-site.xml file to ensure ozone manager addresses are configured properly.
Add the HDFS gateway role on that host.
CDPD-63144: hadoop.ozone.om.request.key.OMKeyRenameRequestWithFSO: Rename key failed with "failed to get parent dir" error
7.3.2 and its SPs and CHFs, 7.3.1 and its SPs and CHFs
Key rename inside the FSO bucket fails and discplays the Failed to get parent dir error. This happens when running impala workloads with ozone.
None.
CDPD-74331: Key put fails with "CompletionException: Failed to write chunk"
7.3.2 and its SPs and CHFs, 7.3.1 and its SPs and CHFs
Key put fails and displays the Failed to write chunk error when there is a volume failure during configuration.
None.
CDPD-74475: YCSB test with Hbase on Ozone degrades performance
7.3.2, 7.3.1 and its SPs and CHFs
HBase on Ozone is currently provided as a Tech Preview feature. Performance characteristics, including throughput, may not yet match those of HBase running on HDFS.
None.
CDPD-74884: Exclusive size of snapshot is always returning 0
7.3.2, 7.3.1 and its SPs and CHFs
The exclusiveSize and exclusiveReplicationSize statistics provided by the snapshot info command return a value of 0 even when the snapshot contains exclusive keys or files that do not exist in other snapshots.
None.
Apache JIRA: HDDS-11528
CDPD-75042: AWS Cli recursive delete only deletes the leaf element
7.3.2 and its SPs and CHFs, 7.3.1 and its SPs and CHFs
AWS CLI rm or delete command fails to delete all files and directories on the Ozone FSO bucket. It only deletes the leaf node.
None.
CDPD-75204: Setting file name length limit can cause NN shutdown with FSImage related error
7.3.2 and its SPs and CHFs, 7.3.1 and its SPs and CHFs
Namenode restart fails after dfs.namenode.fs-limits.max-component-length is set to a lower value and there is existing data present which exceeds the length limit.
Increase the value for the dfs.namenode.fs-limits.max-component-length parameter and restart the namenode.
CDPD-75635: Ozone write fails intermittently because SCM remains in safe mode.
7.3.2, 7.3.1 and its SPs and CHFs
Ozone write operations can fail intermittently after restarting the Storage Container Manager (SCM) leader node (or when stopping the OM leader and an SCM follower). This occurs because SCM may remain in safe mode after the restart, causing write/block allocation prechecks to fail (for example, SafeModePrecheck failed for allocateBlock).
Wait for SCM to exit safe mode automatically after it meets the required thresholds. Alternatively, manually force SCM to exit from safe mode using CLI options.