Known issues for Cloudera Data Services on premises 1.5.5 SP4

Review the known issues for Cloudera Data Services on premises, the impact or changes to the functionality, and the applicable workaround.

Known issues identified in 1.5.5 SP4 release

OPSX-8420: (Use case 1): Longhorn upgrade failure due to longhorn-csi-plugin Pod stuck in a terminating state
During Cloudera Embedded Container Service / Longhorn upgrades (for example, from Cloudera Data Services on premises 1.5.5 SP3 to SP4), the upgrade flow may stall or fail when the longhorn-csi-plugin Pod gets stuck in a terminating state (for example, longhorn-csi-plugin-fbnf5). The longhorn-csi-plugin Pod remains in terminating state, preventing Longhorn components from fully updating and cause CSI provisioner/resizer Pods to crash.
Force remove the stuck longhorn-csi-plugin Pod in the longhorn-system namespace:
kubectl delete pods -n longhorn-system --force

After force deleting the terminating Pod, resume the upgrade flow.

Incase of the Multi-hop Longhorn upgrade process and if the Resume button is not available, then you must trigger the multi-hop Longhorn upgrade API using the Cloudera Manager. You can trigger the API using the following path: /controlPlanes/commands/startLonghornMultihopUpgrade.
OPSX-8420: (Use case 2): Recurring Longhorn upgrade failures involving helm-install-longhorn Recurring Longhorn upgrade failures involving helm-install-longhorn, Helm release errors, and CM/ecs-cli timeouts
Issues causing the failure could range from prior Helm upgrade left the release in a failed state, stuck or failing helm-install-longhorn job blocks retries, and deleting the job triggers RKE2 to spawn a new job.
  1. Verify the current state:
    kubectl get settings current-longhorn-version -n longhorn-system
    helm history longhorn -n longhorn-system
            
    kubectl get job helm-install-longhorn -n longhorn-system
  2. Rollback to last deployed revision:
    helm rollback longhorn <LAST_DEPLOYED_REVISION> -n longhorn-system
  3. Delete stuck job:
    kubectl delete job helm-install-longhorn -n longhorn-system --ignore-not-found
  4. Wait for helm-install-longhorn job to finish.
  5. Resume upgrade process from Cloudera Manager.
OPSX-8396: Longhorn Volume Replica rebuild stuck in Write-Only (WO) mode causing Volume degradation
Longhorn health check fails with: Longhorn Volume Status: ERROR - Volume health check failed: there are unhealthy volumes due to a volume entering degraded state. A volume replica gets stuck in WO mode after a remount or rebuild, and Longhorn fails to recover automatically.
Make sure at least one healthy replica is available and delete the affected stuck replica:
kubectl delete replicas.longhorn.io <replica-name> -n longhorn-system

Longhorn recreates the replica and rebuilds from a healthy copy.

OPSX-8396: (cluster upgrade issue): Longhorn Manager pod CrashLoopBackOff due to port 9501 binding conflict
During the cluster upgrade process, longhorn-manager Pods fail to start and enter CrashLoopBackOff with error: Error conversion webhook server failed: listen tcp :9501: bind: address already in use. The original longhorn-manager container was not properly torn down during pod restart/upgrade process, remaining as an orphaned process holding ports - 9500-9503.
Restart the longhorn-manager DaemonSet:
kubectl -n longhorn-system rollout restart ds/longhorn-manager

Or delete the crash looping longhorn-manager pod to terminate the orphaned process.

OPSX-8551: During the Cloudera Embedded Container Service multihop Longhorn upgrade, hop verify can fail with a blocking workload warning while overall workload health checks pass
Once the CP fluentd logs migration job is completed, Kubernetes removes the job/Pod. Longhorn may retain a stale entry in volumes.longhorn.io → status.kubernetesStatus.workloadsStatus for a time period.
Use the API Explorer to retry the same Cloudera Manager API that started the upgrade (re-invoke the Longhorn multihop upgrade endpoint on the environment).

Before retrying, optionally confirm this is the ghost-migration case (not a real storage blocker):

  • kubectl get job,pod -n cdp | grep fluentd-logs-migration
    — no resources
  • Stale volume CR often already gone from longhorn-system
OPSX-8374: Diagnostic bundle collection fails and new collections cannot be created when primary embedded DB Pod is deleted
Diagnostic bundle collection fails if the primary embedded DB Pod is force deleted while collection is in progress. After the failure, the Diagnostics UI is not able to start a new diagnostic bundle collection, blocking users from collecting the support data for the environment.

After Postgres restarts, TCP connections in Go’s pool become obselete on the server side while the client still reuses them, causing repeated connection reset by peer errors on the same socket.

Restart cadence Pods to restart a new connection.
OPSX-8442: Control Plane authentication outage during Embedded primary DB OOM failover
During a Pod-level OOM crash on the primary database (cdp-embedded-db-2), signing into the UI failed with authentication/LDAP errors for about 2–3 minutes until the failed Pod restarted and stabilized, despite CloudNativePG promoting a replica to primary.

There are some services, which drop the stale DB connection (connection timout) and re-establishes the connection after some time.

If the issue persists for a longer period, restart the affected service, so that it re-establishes the connection and connects to actual primary.
OPSX-8377: CNPG Embedded DB cluster stuck at 1/3 ready: disk-full crash on primary (db-3) and missing PVC on db-1 (Pending)
A couple of issues were observed:
  • Disk-full on primary (db-3): Filling the primary PVC to 100% caused Postgres to crash. CNPG failed over to db-2, but db-3 was not recovered and was removed.
  • Missing PVC on (db-1): Instance cdp-embedded-db-1 was in a pending state because PVC cdp-embedded-db-1 was not found.
If the storage capacity is full or nearing the limits, increase the memory or have sufficent memory beforehand.

For example, perform the following steps to expand the PVC + patch cluster:

  1. kubectl patch pvc <pvc_name> -n cdp -p '{"spec":{"resources":{"requests":{"storage":"250Gi"}}}}'
    
  2. kubectl patch cluster cdp-embedded-db -n cdp --type merge -p '{"spec":{"storage":{"size":"250Gi"}}}'
OPSX-8332: Cloudera Embedded Container Service cluster fails to initialize properly due to embedded-db un-hibernation race condition
In very rare scenarios, a race condition can occur during a cluster restart where the system attempts to un-hibernate the managed embedded-db High Availability (HA) cluster before the cnpg-operator Pod is fully operational. When this happens, the embedded database remains in a hibernated state. Since the system incorrectly marks this step as successful, it leads to an incorrect status report where the start operation completed but the Cloudera Embedded Container Service cluster failed to operate as expected.
If the cluster start operation succeeds but the Cloudera Embedded Container Service cluster is unhealthy, you can manually un-hibernate the database by applying the following patch:
kubectl -n cdp annotate cluster.postgresql.cnpg.io cdp-embedded-db cnpg.io/hibernation=off --overwrite