Configuring LoRA adapters for vLLM Model Endpoints

Configure Low-Rank Adaptation (LoRA) adapters on vLLM endpoints in Cloudera AI Inference service by setting environment variables and verifying rank compatibility.

Cloudera AI Inference service Model Endpoints use NVIDIA NIM with a vLLM backend for serving fine-tuned models. When you attach adapters to a model endpoint, Cloudera AI Inference service automatically passes the --enable-lora flag to the underlying server.

  • Check adapter compatibility:

    Before configuring a model endpoint for vLLM adapters, ensure that your environment and adapter files meet the following prerequisites:

    • Adapter weight compatibility: Ensure that your adapter .safetensors files contain only LoRA decomposition weights, such as lora_A and lora_B matrices. vLLM rejects adapters that include full, non-LoRA weights from modules such as lm_head.weight or embed_tokens.weight with a ValueError.
    • Tooling availability: Ensure that jq or a similar JSON parsing utility is installed to inspect adapter configuration files.
    • Check adapter rank:

      The LoRA rank (r) determines the minimum value required for the --max-lora-rank parameter on the endpoint. If an adapter's rank exceeds the endpoint's configured --max-lora-rank, the endpoint fails to load the adapter.

      To check the rank of an adapter, inspect adapter_config.json by running the following command:

      jq '.r' adapter_config.json
    • Ensure minimum rank requirements by adapter

      When you attach multiple adapters to a single endpoint, set --max-lora-rank to a value equal to or greater than the highest rank among all attached adapters.

      The following table provides examples of minimum rank requirements based on attached adapters:

      Table 1. Minimum rank requirements by adapter
      Attached adapter Rank (r) Minimum --max-lora-rank requirement
      medical-lora 32 32
      sql-lora 128 128
      Both medical-lora and sql-lora attached simultaneously - 128

vLLM LoRA parameters

You can pass additional vLLM configuration settings to the endpoint by using the NIM_PASSTHROUGH_ARGS environment variable.

The following table lists the key vLLM parameters for managing LoRA adapters:

Table 2. vLLM LoRA parameters
Parameter Default value Description
--enable-lora off Enables LoRA support. Cloudera AI Inference service sets this flag automatically when adapters are configured on the endpoint.
--max-lora-rank 16 Specifies the maximum supported LoRA rank. This value must be equal to or greater than the highest rank (r) among all attached adapters.
--max-loras 1 Specifies the maximum number of LoRA adapters loaded into GPU memory simultaneously.
--max-cpu-loras None Specifies the maximum number of LoRA adapters cached in CPU memory for hot-swapping.
  1. Determine the highest rank, that is r among all adapters that you plan to attach to the endpoint.
  2. In the Cloudera AI Inference service UI, open the model endpoint configuration settings.
    1. In the Cloudera console, click the Cloudera AI tile.

      The Cloudera AI Workbenches page displays.

    2. Click Model Endpoints under Deployments on the left navigation menu.

      The Model Endpoints landing page is displayed.

    3. Locate the model endpoint you want to modify, open the Actions drop-down menu, and select Edit.
    4. In the Edit Configuration window, update the model endpoint details across the required tabs.
    5. Go to the Review and Update tab and click Update to save your changes.
  3. In the environment variables section, add the NIM_PASSTHROUGH_ARGS variable.
  4. Set NIM_PASSTHROUGH_ARGS to include your required vLLM parameters.

    To configure a single argument, set the variable as follows:

    NIM_PASSTHROUGH_ARGS=--max-lora-rank 128

    To configure multiple arguments, separate each argument with a space:

    NIM_PASSTHROUGH_ARGS=--max-lora-rank 128 --max-loras 4 --max-cpu-loras 8
  5. Deploy or restart the Model Endpoint to apply the configuration.