Known issues for Cloudera AI on premises 1.5.5 SP4

You might run into some known issues while using Cloudera AI on premises 1.5.5 SP4.

Cloudera AI on premises 1.5.5, 1.5.5 SP1, 1.5.5 SP2 and 1.5.5 SP3 existing known issues are carried into Cloudera AI on premises 1.5.5 SP4.

For known issues fixed since Cloudera AI on premises 1.5.5, see:

Cloudera AI Workbench

DSE-59571: MIG variant profiles are not supported in Cloudera AI

Cloudera AI supports standard Multi-instance GPU (MIG) profiles only, MIG variant profiles are not supported, including profile variants such as +me, +me.all, +gfx, and -me (for example, 1g.10gb+me).

DSE-58791: Not supported GPU usage tracking for MIG MIX GPU in Grafana Quota Management Dashboard

When using Multi-Instance GPU (MIG) slicing with mixed GPU profiles in Cloudera AI Workbench, GPU usage metrics are not tracked or displayed within the Grafana Quota Management Dashboard. GPU usage tracking for Multi-Instance GPU (MIG) slicing and mixed GPU configuration profiles is not supported in the Grafana Quota Management Dashboard.

Workaround:

View GPU utilization through the standard Usage Dashboard or the Resource Usage Dashboard, that actively supports and displays resource tracking for MIG GPUs.

DSE-61563: MIG slice usage is displayed as GPU count in the GPU Usage graph
When using Multi-Instance GPU (MIG) slicing, each MIG slice is currently counted as one GPU in the GPU Usage graphs on the Usage and Resource Usage dashboards.
DSE-43884: Default timeout insufficient for large databases to be migrated from Cloudera Data Science Workbench to Cloudera AI

When migrating a Cloudera Data Science Workbench workspace containing large database tables, such as a dashboards table with several hundred thousand unstopped rows, the database update process inside the cml-db-migrate pod can take several hours to complete. Because the default job completion timeout is set to 10 minutes, the migration fails prematurely due to a timeout waiting for the cml-db-migrate job to finish.

DSE-58875: Workbench applications can remain indefinitely in the Starting state when the web server is unreachable

Previously, if a deployed application web server was unreachable on localhost:$CDSW_APP_PORT, for example, because of an incorrect port binding, a crashed framework, or HTTPS being served on an HTTP port, the application could remain indefinitely in the Starting state. Since the underlying Kubernetes pod remained healthy and in the Running state, the failure was not detected or reported to users or administrators.

A startup timeout reconciler has been added to address this issue. If an application remains in the Starting state beyond the configured timeout after the workload begins running, the reconciler automatically transitions the application to the Failed state and displays the APP_WEBSERVER_UNREACHABLE error code on the application card and in the logs.

DSE-57724: Sessions fail with argument list too long error at high session concurrency
If you scale Cloudera AI Workbench to high concurrent session counts (approximately 850 through 1,000 sessions), new session pods fail to start and enter an Init:Error state. The engine-deps init container displays the following error:
/deps/engine-deps-install: argument list too long
This issue occurs because Kubernetes automatically injects environment variables for all namespace Services into newly created pods. At high session concurrency or when stale runtime Service records accumulate, the volume of injected environment variables exceeds environment size limits, preventing container initialization scripts from running.

Workaround:

  • Limit concurrency: Keep the number of active concurrent sessions in any single user namespace below ~700.
  • Resource cleanup: Manually clean up any unremoved or stale cdsw-runtime-* Service objects in affected namespaces to reduce injected environment variable overhead.

Cloudera AI Registry

DSE-61353, DSE-61273: Model privacy and Administrator roles

Private models are visible only to their owners. Similarly, only the model owner can delete a model, regardless of whether the model is private or public.

This implies the following:

  • Users assigned elevated roles, MLAdmin and PowerUser, cannot override these ownership constraints.
  • Only the user who owns the model can perform the deletion.
DSE-61665: Post-upgrade relinking required for remote Cloudera AI Registries to deploy models in Cloudera AI Inference service
After upgrading to Cloudera AI 1.5.5 SP4, deploying a model from a Cloudera AI Registry in a different environment than the Cloudera AI Inference service fails. Deploying endpoint API requests, like /api/v1alpha1/deployEndpoint, returns an HTTP 403 Forbidden error, similar to:
{
  "err": "rpc error: code = PermissionDenied desc = user does not have access to the remote registry environment"
}

After the upgrade from Cloudera AI 1.5.5 SP3 CHF3 to 1.5.5 SP4, if the Cloudera AI Inference service and the Cloudera AI Registry are in different environments, they need to be relinked to allow Cloudera AI Registry access across environments.

Workaround:

To restore model deployment functionality from the remote Cloudera AI Registry, relink the registry from the Cloudera AI Inference service application:

  1. In the Cloudera console, click the Cloudera AI tile.

    The Cloudera AI page is displayed.

  2. Click AI Inference Services under ADMINISTRATION in the left navigation menu.

    The AI Inference Services page is displayed.

  3. For a selected Cloudera AI Inference service instance, click the icon from the Actions menu and select the Update Storage Configuration option.

    The Update Storage Configuration page is displayed.

  4. Select the Cloudera AI Registry from the Select AI Registry drop-down list that you want to relink to the Cloudera AI Inference service instance.

Once relinked, the secret is updated with the required roles and you can deploy models as expected.

Cloudera AI Inference service

Cloudera AI Inference service

In Cloudera AI 1.5.5 SP4, HGX support in the Cloudera AI Inference service is limited to single-node deployments.

DSE-61664: Cloudera AI Inference service instance appears disabled during endpoint creation in cross-registry environments

In cross-registry configurations where the Cloudera AI Inference service instance and Cloudera AI Registry are deployed in different environments, when you attempt to create a model endpoint through the Cloudera AI Inference service user interface (UI) by clicking Create Endpoint, the Cloudera AI Inference service instance appears disabled or unavailable because of a strict environment registry check. This issue occurs in both fresh installations and upgraded environments.

Workaround:

Deploy the model directly from the Cloudera AI Registry interface instead of the Cloudera AI Inference service user interface:

  • Go to Registered Models under Deployments in the left navigation menu.
  • Select the model that you want to deploy.
  • Click Deploy Model and select your target Cloudera AI Inference service instance.
DSE-49451: Tensor parallelism fails when multiple MIG slices are scheduled on the same host

Tensor parallelism is not supported when you schedule multiple Multi-Instance GPU (MIG) slices on the same host. If you attempt to deploy a model using tensor parallelism across multiple MIG slices on a single host, deployments might fail with an NVIDIA Collective Communications Library (NCCL) error: unhandled system error, including after pod restarts.

Workaround:

To work around this issue, perform one of the following actions:

  • Use a single MIG slice instead of multiple slices.
  • Select a larger GPU resource profile.
  • Deploy a smaller model, or reduce memory requirements by setting the --max_model_len parameter.
DSE-49451: Host displays as Invalid when GPUs use inconsistent Multi-Instance GPU profile types

When you configure a single Multi-Instance GPU (MIG) profile across GPUs on a host, each physical GPU must use one uniform profile slice type. If a GPU is configured with multiple slice types, the host status displays as Invalid, and available resource profiles do not appear in the user interface (UI).

Workaround:

Before deploying a model endpoint, reconfigure MIG on the host so that each GPU uses one consistent profile type:
  1. Go to Cloudera Management Console UI.
  2. Select Administration.
  3. Select the MIG Configuration tab.
DSE-61631: Knox API token authentication fails on Cloudera AI Inference service endpoints after upgrade

Following the upgrade from Cloudera AI 1.5.5 SP2 to 1.5.5 SP4, in OpenShift Container Platform environment with upgrades from Openshift 4.19 to 4.21, API requests to Cloudera AI Inference service model serving endpoints that use Knox-generated API tokens fail with an HTTP 401 error. Requests authenticated with standard JSON Web Token (JWT) credentials continue to work as expected.

This authentication failure originates within the Knox external authorization chain (cml-serving-remoteauth-remoteauthprovider) at the ExtAuthz layer before reaching the model endpoint. During the upgrade from version Cloudera AI 1.5.5 SP2 to 1.5.5 SP4, the Knox serving topology descriptor or token validation trust chain might become mismatched.

Workaround:

  1. Determine the Knox Authentication URL.
    1. Identify your Knox Gateway host in Cloudera Management Console. On Cloudera on premises, your target URL must follow the following exact structure:
      https://[***KNOX-GATEWAY-HOST***]:8443/gateway/cdp-proxy-token/auth/api/v1/pre

      Example

      https://ccycloud-10.cai-719.root.comops.site:8443/gateway/cdp-proxy-token/auth/api/v1/pre
  2. Patch Knox via kubectl.

    1. Set your specific URL as an environment variable, then apply the patch to update the Knox capabilities in the cml-serving namespace.

      REMOTE_AUTH_URL="https://ccycloud-10.cai-719.root.comops.site:8443/gateway/cdp-proxy-token/auth/api/v1/pre"
      
      kubectl patch knoxcapabilities -n cml-serving remoteauth --type=merge -p "{
        \"spec\": {
          \"capabilities\": [{
            \"type\": \"remoteauthprovider\",
            \"overrides\": {
              \"knox-auth-service.acl\": \"*;*;*\",
              \"remote.auth.url\": \"${REMOTE_AUTH_URL}\"
            }
          }]
        }
      }"
  3. Verify the update.

    1. Query the configuration to confirm the new URL was successfully written to the cluster.

      kubectl get knoxcapabilities -n cml-serving remoteauth -o yaml | grep remote.auth.url

      Expected outcome:

      kubectl get knoxcapabilities -n cml-serving remoteauth -o yaml | grep remote.auth.url
            remote.auth.url: https://ccycloud-10.cai-719.root.comops.site:8443/gateway/cdp-proxy-token/auth/api/v1/pre

Limitations

DSE-56558: Canary deployment

Canary deployment is supported only for MLflow models in Cloudera AI1.5.5 SP4 and higher releases. It is not supported for LLM workloads.

Spark External Shuffle Service for Spark on Kubernetes

Cloudera AI does not support Spark External Shuffle Service for Spark on Kubernetes.