Low-Rank Adaptation adapters with Cloudera AI Inference service

Cloudera AI Inference service supports deploying Parameter-Efficient Fine-Tuning (PEFT) adapters, specifically LoRA (Low-Rank Adaptation) adapters in Cloudera AI 1.5.5 SP4 and higher releases.

This capability allows you to serve multiple fine-tuned model variants from one endpoint, significantly reducing GPU resource consumption compared to deploying each variant as an independent inference service. A maximum of 10 adapters can be deployed per endpoint at the same time. The feature is enabled by default.

Adapter support is available for the following models:

  • HuggingFace base models served through the vLLM runtime

  • NVIDIA NIM base models

You can specify one or more adapters when creating or updating a Model Endpoint.

Using Low-Rank Adaptation adapters provides the following key benefits:

  • GPU efficiency — Multiple LoRA adapters share a single base model loaded on one set of GPUs, eliminating the need for N independent deployments for N task-specific variants.

  • Single endpoint — The base model and all attached adapters are accessible through one inference endpoint. Clients select the desired adapter using the standard OpenAI model field.

  • Deploy-time validation — Adapter-to-base-model compatibility is verified at deployment time, preventing mismatched configurations from reaching runtime.

  • Per-adapter metrics — Request metrics are available per adapter through the existing Prometheus scrape endpoint without additional scrape configuration.