Configuring LoRA adapters for vLLM Model Endpoints
Configure Low-Rank Adaptation (LoRA) adapters on vLLM endpoints in Cloudera AI Inference service by setting environment variables and verifying rank compatibility.
Cloudera AI Inference service Model Endpoints use NVIDIA NIM with a vLLM backend
for serving fine-tuned models. When you attach adapters to a model endpoint, Cloudera AI Inference service automatically passes the
--enable-lora flag to the underlying server.
-
Check adapter compatibility:
Before configuring a model endpoint for vLLM adapters, ensure that your environment and adapter files meet the following prerequisites:
- Adapter weight compatibility: Ensure that your
adapter .safetensors files contain only LoRA decomposition
weights, such as
lora_Aandlora_Bmatrices. vLLM rejects adapters that include full, non-LoRA weights from modules such aslm_head.weightorembed_tokens.weightwith aValueError. - Tooling availability: Ensure that jq or a similar JSON parsing utility is installed to inspect adapter configuration files.
-
Check adapter rank:
The LoRA rank (
r) determines the minimum value required for the--max-lora-rankparameter on the endpoint. If an adapter's rank exceeds the endpoint's configured--max-lora-rank, the endpoint fails to load the adapter.To check the rank of an adapter, inspect adapter_config.json by running the following command:
jq '.r' adapter_config.json -
Ensure minimum rank requirements by adapter
When you attach multiple adapters to a single endpoint, set
--max-lora-rankto a value equal to or greater than the highest rank among all attached adapters.The following table provides examples of minimum rank requirements based on attached adapters:
Table 1. Minimum rank requirements by adapter Attached adapter Rank ( r)Minimum --max-lora-rankrequirementmedical-lora 32 32 sql-lora 128 128 Both medical-lora and sql-lora attached simultaneously - 128
- Adapter weight compatibility: Ensure that your
adapter .safetensors files contain only LoRA decomposition
weights, such as
vLLM LoRA parameters
You can pass additional vLLM configuration settings to the endpoint by using the NIM_PASSTHROUGH_ARGS environment variable.
The following table lists the key vLLM parameters for managing LoRA adapters:
| Parameter | Default value | Description |
|---|---|---|
--enable-lora |
off |
Enables LoRA support. Cloudera AI Inference service sets this flag automatically when adapters are configured on the endpoint. |
--max-lora-rank |
16 |
Specifies the maximum supported LoRA rank. This value must be equal to or
greater than the highest rank (r) among all attached
adapters. |
--max-loras |
1 |
Specifies the maximum number of LoRA adapters loaded into GPU memory simultaneously. |
--max-cpu-loras |
None | Specifies the maximum number of LoRA adapters cached in CPU memory for hot-swapping. |
