Known issues for Cloudera AI on premises 1.5.5 SP4
You might run into some known issues while using Cloudera AI on premises 1.5.5 SP4.
Cloudera AI on premises 1.5.5, 1.5.5 SP1, 1.5.5 SP2 and 1.5.5 SP3 existing known issues are carried into Cloudera AI on premises 1.5.5 SP4.
Cloudera AI Workbench
- DSE-59571: MIG variant profiles are not supported in Cloudera AI
-
Cloudera AI supports standard Multi-instance GPU (MIG) profiles only, MIG variant profiles are not supported, including profile variants such as
+me,+me.all,+gfx, and -me(for example,1g.10gb+me). - DSE-58791: Not supported GPU usage tracking for MIG MIX GPU in Grafana Quota Management Dashboard
-
When using Multi-Instance GPU (MIG) slicing with mixed GPU profiles in Cloudera AI Workbench, GPU usage metrics are not tracked or displayed within the Grafana Quota Management Dashboard. GPU usage tracking for Multi-Instance GPU (MIG) slicing and mixed GPU configuration profiles is not supported in the Grafana Quota Management Dashboard.
Workaround:
View GPU utilization through the standard Usage Dashboard or the Resource Usage Dashboard, that actively supports and displays resource tracking for MIG GPUs.
- DSE-61563: MIG slice usage is displayed as GPU count in the GPU Usage graph
- When using Multi-Instance GPU (MIG) slicing, each MIG slice is currently counted as one GPU in the GPU Usage graphs on the Usage and Resource Usage dashboards.
- DSE-43884: Default timeout insufficient for large databases to be migrated from Cloudera Data Science Workbench to Cloudera AI
-
When migrating a Cloudera Data Science Workbench workspace containing large database tables, such as a
dashboardstable with several hundred thousand unstopped rows, the database update process inside thecml-db-migratepod can take several hours to complete. Because the default job completion timeout is set to 10 minutes, the migration fails prematurely due to a timeout waiting for thecml-db-migratejob to finish. - DSE-58875: Workbench applications can remain indefinitely in the Starting state when the web server is unreachable
-
Previously, if a deployed application web server was unreachable on
localhost:$CDSW_APP_PORT, for example, because of an incorrect port binding, a crashed framework, or HTTPS being served on an HTTP port, the application could remain indefinitely in theStartingstate. Since the underlying Kubernetes pod remained healthy and in theRunningstate, the failure was not detected or reported to users or administrators.A startup timeout reconciler has been added to address this issue. If an application remains in the
Startingstate beyond the configured timeout after the workload begins running, the reconciler automatically transitions the application to the Failed state and displays theAPP_WEBSERVER_UNREACHABLEerror code on the application card and in the logs. - DSE-57724: Sessions fail with
argument list too longerror at high session concurrency -
If you scale Cloudera AI Workbench to high concurrent session counts (approximately 850 through 1,000 sessions), new session pods fail to start and enter an Init:Error state. The
engine-depsinit container displays the following error:
This issue occurs because Kubernetes automatically injects environment variables for all namespace Services into newly created pods. At high session concurrency or when stale runtime Service records accumulate, the volume of injected environment variables exceeds environment size limits, preventing container initialization scripts from running./deps/engine-deps-install: argument list too longWorkaround:
- Limit concurrency: Keep the number of active concurrent sessions in any single user namespace below ~700.
- Resource cleanup: Manually clean up any unremoved or stale
cdsw-runtime-*Service objects in affected namespaces to reduce injected environment variable overhead.
Cloudera AI Registry
- DSE-61353, DSE-61273: Model privacy and Administrator roles
-
Private models are visible only to their owners. Similarly, only the model owner can delete a model, regardless of whether the model is private or public.
This implies the following:
- Users assigned elevated roles,
MLAdminandPowerUser, cannot override these ownership constraints. - Only the user who owns the model can perform the deletion.
- Users assigned elevated roles,
- DSE-61665: Post-upgrade relinking required for remote Cloudera AI Registries to deploy models in Cloudera AI Inference service
-
After upgrading to Cloudera AI 1.5.5 SP4, deploying a model from a Cloudera AI Registry in a different environment than the Cloudera AI Inference service fails. Deploying endpoint API requests, like
/api/v1alpha1/deployEndpoint, returns anHTTP 403 Forbiddenerror, similar to:{ "err": "rpc error: code = PermissionDenied desc = user does not have access to the remote registry environment" }After the upgrade from Cloudera AI 1.5.5 SP3 CHF3 to 1.5.5 SP4, if the Cloudera AI Inference service and the Cloudera AI Registry are in different environments, they need to be relinked to allow Cloudera AI Registry access across environments.
Workaround:
To restore model deployment functionality from the remote Cloudera AI Registry, relink the registry from the Cloudera AI Inference service application:
-
In the Cloudera console, click the Cloudera AI tile.
The Cloudera AI page is displayed.
-
Click AI Inference Services under ADMINISTRATION in the left navigation menu.
The AI Inference Services page is displayed.
-
For a selected Cloudera AI Inference service instance, click the
icon from the
Actions menu and select the Update Storage
Configuration option.The Update Storage Configuration page is displayed.
-
Select the Cloudera AI Registry from the Select AI Registry drop-down list that you want to relink to the Cloudera AI Inference service instance.
Once relinked, the secret is updated with the required roles and you can deploy models as expected.
-
Cloudera AI Inference service
- Cloudera AI Inference service
-
In Cloudera AI 1.5.5 SP4, HGX support in the Cloudera AI Inference service is limited to single-node deployments.
- DSE-61664: Cloudera AI Inference service instance appears disabled during endpoint creation in cross-registry environments
-
In cross-registry configurations where the Cloudera AI Inference service instance and Cloudera AI Registry are deployed in different environments, when you attempt to create a model endpoint through the Cloudera AI Inference service user interface (UI) by clicking Create Endpoint, the Cloudera AI Inference service instance appears disabled or unavailable because of a strict environment registry check. This issue occurs in both fresh installations and upgraded environments.
Workaround:
Deploy the model directly from the Cloudera AI Registry interface instead of the Cloudera AI Inference service user interface:
- Go to Registered Models under Deployments in the left navigation menu.
- Select the model that you want to deploy.
- Click Deploy Model and select your target Cloudera AI Inference service instance.
- DSE-49451: Tensor parallelism fails when multiple MIG slices are scheduled on the same host
-
Tensor parallelism is not supported when you schedule multiple Multi-Instance GPU (MIG) slices on the same host. If you attempt to deploy a model using tensor parallelism across multiple MIG slices on a single host, deployments might fail with an NVIDIA Collective Communications Library (NCCL) error:
unhandled system error, including after pod restarts.Workaround:
To work around this issue, perform one of the following actions:
- Use a single MIG slice instead of multiple slices.
- Select a larger GPU resource profile.
- Deploy a smaller model, or reduce memory requirements by setting the
--max_model_lenparameter.
- DSE-49451: Host displays as
Invalidwhen GPUs use inconsistent Multi-Instance GPU profile types -
When you configure a single Multi-Instance GPU (MIG) profile across GPUs on a host, each physical GPU must use one uniform profile slice type. If a GPU is configured with multiple slice types, the host status displays as
Invalid, and available resource profiles do not appear in the user interface (UI).Workaround:
Before deploying a model endpoint, reconfigure MIG on the host so that each GPU uses one consistent profile type:- Go to Cloudera Management Console UI.
- Select Administration.
- Select the MIG Configuration tab.
- DSE-61631: Knox API token authentication fails on Cloudera AI Inference service endpoints after upgrade
-
Following the upgrade from Cloudera AI 1.5.5 SP2 to 1.5.5 SP4, in OpenShift Container Platform environment with upgrades from Openshift 4.19 to 4.21, API requests to Cloudera AI Inference service model serving endpoints that use Knox-generated API tokens fail with an HTTP
401error. Requests authenticated with standard JSON Web Token (JWT) credentials continue to work as expected.This authentication failure originates within the Knox external authorization chain (
cml-serving-remoteauth-remoteauthprovider) at theExtAuthzlayer before reaching the model endpoint. During the upgrade from version Cloudera AI 1.5.5 SP2 to 1.5.5 SP4, the Knox serving topology descriptor or token validation trust chain might become mismatched.Workaround:
- Determine the Knox Authentication URL.
- Identify your Knox Gateway host in Cloudera Management Console. On Cloudera on premises, your target URL must follow
the following exact
structure:
https://[***KNOX-GATEWAY-HOST***]:8443/gateway/cdp-proxy-token/auth/api/v1/preExample
https://ccycloud-10.cai-719.root.comops.site:8443/gateway/cdp-proxy-token/auth/api/v1/pre
- Identify your Knox Gateway host in Cloudera Management Console. On Cloudera on premises, your target URL must follow
the following exact
structure:
-
Patch Knox via kubectl.
-
Set your specific URL as an environment variable, then apply the patch to update the Knox capabilities in the
cml-servingnamespace.REMOTE_AUTH_URL="https://ccycloud-10.cai-719.root.comops.site:8443/gateway/cdp-proxy-token/auth/api/v1/pre" kubectl patch knoxcapabilities -n cml-serving remoteauth --type=merge -p "{ \"spec\": { \"capabilities\": [{ \"type\": \"remoteauthprovider\", \"overrides\": { \"knox-auth-service.acl\": \"*;*;*\", \"remote.auth.url\": \"${REMOTE_AUTH_URL}\" } }] } }"
-
-
Verify the update.
-
Query the configuration to confirm the new URL was successfully written to the cluster.
kubectl get knoxcapabilities -n cml-serving remoteauth -o yaml | grep remote.auth.urlExpected outcome:
kubectl get knoxcapabilities -n cml-serving remoteauth -o yaml | grep remote.auth.url remote.auth.url: https://ccycloud-10.cai-719.root.comops.site:8443/gateway/cdp-proxy-token/auth/api/v1/pre
-
- Determine the Knox Authentication URL.
Limitations
- DSE-56558: Canary deployment
-
Canary deployment is supported only for MLflow models in Cloudera AI1.5.5 SP4 and higher releases. It is not supported for LLM workloads.
- Spark External Shuffle Service for Spark on Kubernetes
-
Cloudera AI does not support Spark External Shuffle Service for Spark on Kubernetes.
