Low-Rank Adaptation adapters with Cloudera AI Inference service
Cloudera AI Inference service supports deploying Parameter-Efficient Fine-Tuning (PEFT) adapters, specifically LoRA (Low-Rank Adaptation) adapters in Cloudera AI 1.5.5 SP4 and higher releases.
This capability allows you to serve multiple fine-tuned model variants from one endpoint, significantly reducing GPU resource consumption compared to deploying each variant as an independent inference service. A maximum of 10 adapters can be deployed per endpoint at the same time. The feature is enabled by default.
Adapter support is available for the following models:
-
HuggingFace base models served through the vLLM runtime
-
NVIDIA NIM base models
You can specify one or more adapters when creating or updating a Model Endpoint.
Using Low-Rank Adaptation adapters provides the following key benefits:
-
GPU efficiency — Multiple LoRA adapters share a single base model loaded on one set of GPUs, eliminating the need for N independent deployments for N task-specific variants.
-
Single endpoint — The base model and all attached adapters are accessible through one inference endpoint. Clients select the desired adapter using the standard OpenAI
modelfield. -
Deploy-time validation — Adapter-to-base-model compatibility is verified at deployment time, preventing mismatched configurations from reaching runtime.
-
Per-adapter metrics — Request metrics are available per adapter through the existing Prometheus scrape endpoint without additional scrape configuration.
