Known issues for Cloudera Data Services on premises 1.5.5 SP4
Review the known issues for Cloudera Data Services on premises, the impact or changes to the functionality, and the applicable workaround.
Known issues identified in 1.5.5 SP4 release
- OPSX-8420: (Use case 1): Longhorn upgrade failure due to longhorn-csi-plugin Pod stuck in a terminating state
- During Cloudera Embedded Container Service / Longhorn upgrades (for
example, from Cloudera Data Services on premises 1.5.5 SP3 to SP4), the
upgrade flow may stall or fail when the
longhorn-csi-pluginPod gets stuck in a terminating state (for example,longhorn-csi-plugin-fbnf5). Thelonghorn-csi-pluginPod remains in terminating state, preventing Longhorn components from fully updating and causeCSI provisioner/resizerPods to crash.
- OPSX-8420: (Use case 2): Recurring Longhorn upgrade failures
involving
helm-install-longhorn Recurring Longhorn upgrade failures involvinghelm-install-longhorn, Helm release errors, and CM/ecs-cli timeouts - Issues causing the failure could range from prior Helm upgrade left
the release in a failed state, stuck or failing
helm-install-longhornjob blocks retries, and deleting the job triggers RKE2 to spawn a new job.
- OPSX-8396: Longhorn Volume Replica rebuild stuck in Write-Only (WO) mode causing Volume degradation
- Longhorn health check fails with:
Longhorn Volume Status: ERROR - Volume health check failed: there are unhealthy volumesdue to a volume entering degraded state. A volume replica gets stuck in WO mode after a remount or rebuild, and Longhorn fails to recover automatically.
- OPSX-8396: (cluster upgrade issue): Longhorn Manager pod CrashLoopBackOff due to port 9501 binding conflict
- During the cluster upgrade process,
longhorn-managerPods fail to start and enterCrashLoopBackOffwith error:Error conversion webhook server failed: listen tcp :9501: bind: address already in use. The originallonghorn-managercontainer was not properly torn down during pod restart/upgrade process, remaining as an orphaned process holding ports - 9500-9503.
- OPSX-8551: During the Cloudera Embedded Container Service multihop Longhorn upgrade, hop verify can fail with a blocking workload warning while overall workload health checks pass
- Once the CP fluentd logs migration job is completed, Kubernetes
removes the job/Pod. Longhorn may retain a stale entry in
volumes.longhorn.io→status.kubernetesStatus.workloadsStatusfor a time period.
- OPSX-8374: Diagnostic bundle collection fails and new collections cannot be created when primary embedded DB Pod is deleted
- Diagnostic bundle collection fails if the primary embedded DB Pod is
force deleted while collection is in progress. After the failure, the Diagnostics UI is not
able to start a new diagnostic bundle collection, blocking users from collecting the support
data for the environment.
After Postgres restarts, TCP connections in Go’s pool become obselete on the server side while the client still reuses them, causing repeated connection reset by peer errors on the same socket.
- OPSX-8442: Control Plane authentication outage during Embedded primary DB OOM failover
- During a Pod-level OOM crash on the primary database
(
cdp-embedded-db-2), signing into the UI failed with authentication/LDAP errors for about 2–3 minutes until the failed Pod restarted and stabilized, despite CloudNativePG promoting a replica to primary.There are some services, which drop the stale DB connection (connection timout) and re-establishes the connection after some time.
- OPSX-8377: CNPG Embedded DB cluster stuck at 1/3 ready: disk-full crash on primary (db-3) and missing PVC on db-1 (Pending)
- A couple of issues were observed:
- Disk-full on primary (db-3): Filling the primary PVC to 100% caused Postgres to crash. CNPG failed over to db-2, but db-3 was not recovered and was removed.
- Missing PVC on (db-1): Instance
cdp-embedded-db-1was in a pending state because PVCcdp-embedded-db-1was not found.
- OPSX-8332: Cloudera Embedded Container Service cluster fails to initialize properly due to embedded-db un-hibernation race condition
- In very rare scenarios, a race condition can occur during a cluster
restart where the system attempts to un-hibernate the managed embedded-db High Availability
(HA) cluster before the
cnpg-operatorPod is fully operational. When this happens, the embedded database remains in a hibernated state. Since the system incorrectly marks this step as successful, it leads to an incorrect status report where the start operation completed but the Cloudera Embedded Container Service cluster failed to operate as expected.
