What's new in 1.5.5 SP4
Cloudera AI on premises 1.5.5 SP4 delivers a set of new features for Cloudera AI.
Cloudera AI Workbench
NVIDIA GPUDirect Remote Direct Memory Access
NVIDIA GPUDirect Remote Direct Memory Access (RDMA) enables applications to directly access remote GPU memory without routing data through the operating system kernel. By minimizing CPU involvement in network transfers, GPUDirect RDMA reduces latency and overhead while freeing CPU resources for application processing. For more information see NVIDIA GPUDirect Remote Direct Memory Access for Cloudera AI.
Multi-instance GPU support in Cloudera AI Workbench
Multi-Instance GPU (MIG) support in Cloudera AI Workbench enables you to partition a physical GPU into isolated virtual instances to run multiple workloads, such as sessions, models, jobs or applications, simultaneously and increase hardware utilization. This enhancement allows multiple resource-isolated model inference workloads to run concurrently on the same GPU, improving GPU utilization and reducing the need to dedicate an entire physical GPU to a single deployment. For more information see Multi-instance GPU support in Cloudera AI.
Cloudera AI Workbench
The PostgreSQL database of Cloudera AI Registry is now upgraded to version 17 (PG17).
Cloudera AI Inference service
Persistent Volume usage for model storage in Cloudera AI Inference service
Persistent volume storage in Cloudera AI Inference service provides dedicated storage for model artifacts, eliminating dependency on the underlying node storage capacity.
You can use persistent volume storage to host large models when repeated downloads are impractical, such as when the model size exceeds the available root volume capacity. Cloudera AI Inference service downloads the model artifacts once and reuses them across workload restarts, avoiding the need to download the artifacts each time you recreate a workload. For more information see Persistent Volume usage for model storage in Cloudera AI Inference service.
Multi-instance GPU support in Cloudera AI Inference service
Cloudera AI Inference service supports NVIDIA GPU Operator and dynamic fractional GPU allocation (MIG), enabling users to run efficient, resource-isolated inference workloads on MIG-supported GPUs. This enhancement allows users to deploy models by selecting granular GPU memory profiles, helping smaller workloads avoid consuming entire high-cost GPU resources unnecessarily. The Model deployment panel provides predefined and standardized GPU resource profiles that can be selected based on the memory requirements of the model. For more information see Creating a Model Endpoint using UI.
Enhanced Replica monitoring for Model Endpoints
The Model Endpoint Details page header now features a real-time Replicas status row to improve endpoint monitoring and operational visibility. This status provides immediate breakdown counts for Total, Healthy, Pull Image, and Pending replicas across model deployments. By tracking active Kubernetes pod events alongside revision metadata, you can quickly identify deployment progress and image-pull delays directly from the UI without checking underlying cluster logs.
Low-Rank Adaptation adapters with Cloudera AI Inference service
Cloudera AI Inference service supports deploying LoRA (Low-Rank Adaptation) adapters in Cloudera AI 1.5.5 SP4 and higher releases.
This capability allows you to serve multiple fine-tuned model variants from one endpoint, significantly reducing GPU resource consumption compared to deploying each variant as an independent inference service. For more information, see Low-Rank Adaptation adapters with Cloudera AI Inference service.
Filter and discover certified models in Model Hub
Model Hub now provides advanced filtering capabilities using a new
Certification drop-down list and workload certification badges to
help you quickly identify optimized models. You can easily filter the catalog by
certification type, such as Agent Studio Certified, as well as model
family, machine learning task, and provider registry. Certified models include preverified
hardware configurations, compatible runtimes, and validated schemas to ensure smooth
deployment across Cloudera AI tools. Enterprise administrators can
also define custom certification labels directly within air-gapped
ModelHubCatalog YAML configurations.
Improvements
Agent Studio
Standardized prompts for Large Language Model (LLM) validation in Agent Studio
When testing a registered LLM in Cloudera AI Agent Studio, both the Test Model and Test Function Calling actions now use a standard, non-editable prompt. Previously, these actions accepted a user-supplied prompt; standardizing the prompt ensures consistent validation of model credentials, network connectivity, and function-calling support across environments. For more information, see Testing registered models in Agent Studio .
