Multi-Instance GPU support in Cloudera AI

Multi-Instance GPU (MIG) support in Cloudera AI enables you to partition a physical GPU into isolated virtual instances to run multiple workloads, such as sessions, models, jobs or applications, simultaneously and increase hardware utilization.

To avoid allocating an entire physical GPU to a single workload, you can use the NVIDIA MIG virtualization framework. You can partition a single high-performance physical GPU into as many as seven isolated virtual GPU instances. This partitioning enables you to run multiple smaller, resource-isolated model inference workloads simultaneously on the same physical GPU.

This fractional GPU allocation provides the following key advantages:

  • Improved hardware utilization: You can run multiple lightweight inference sessions on a single physical GPU to increase overall hardware efficiency and maximize your return on investment.

  • Multitenant isolation: Each MIG slice operates in an isolated environment with its own dedicated streaming multiprocessors (SMs), high-bandwidth memory (HBM), and L2 cache to ensure predictable performance and prevent workloads from interfering with each other.

  • Flexible resource allocation: You can select a Mixed NVIDIA Multi-Instance GPU (MIG) configuration slice profile in the resource profile from the standardized profiles available on your cluster based on your workload’s memory requirements.

Flexible resource allocation with slice profiles

Cloudera AI abstracts raw Kubernetes resource strings into standardized slice profiles.

This abstraction groups various physical memory capacities under functional profile names based on the compute instance size. For example, whether a compute instance requires 10 GB or 20 GB of video RAM (VRAM), both map directly to the small slice profile.

MIG support in Cloudera AI Workbench is dynamic, based on the hardware, you can select different slice profiles.The following table displays the slice profile examples and their recommended use cases:

Table 1. Slice profile examples and recommended use cases
Compute instance size Slice profile Recommended use case
1g / 5 GB Small slice profile Lightweight tasks or small models
2g / 10 GB Medium slice profile Standard inference sessions
3g / 20 GB Medium-large slice profile Moderate-scale or large-scale LLMs
7g / 40 GB Full maximum slice profile Heavy text or embedding models, and high-concurrency workloads

Single and mixed multi-Instance GPU configurations in Cloudera AI

In Cloudera AI 1.5.5 SP4 and higher releases, both single and mixed NVIDIA MIG configurations are supported.

With a single GPU configuration strategy, GPUs are provisioned with uniform, identical MIG slice profiles across the instance.

Mixed NVIDIA MIG configurations allow a single physical GPU or GPU node to be divided into multiple virtual GPU instances with different profiles and resource allocations.

In Cloudera AI, a mixed MIG strategy supports heterogeneous GPU slicing and granular resource allocation, helping optimize utilization of high-performance GPU hardware across diverse workloads.

The mixed MIG strategy provides the following key characteristics:

  • Heterogeneous slicing: You can configure different MIG profiles on the same physical GPU or across different GPUs on a single host. For example, you can slice a single NVIDIA H100 GPU to simultaneously host the following profiles:

    • A single 3g.20gb slice for medium-large workloads, such as embeddings or smaller Large Language Models (LLMs)

    • Multiple 1g.5gb slices for lightweight inference or development tasks

  • Granular resource allocation: The mixed MIG strategy abstracts raw hardware into standardized, user-facing memory profiles in the Cloudera AI Inference deployment panel.

  • GPU-level flexibility: In the Cloudera AI user interface, selecting a mixed strategy enables you to clear the Apply to All GPUs check box. This flexibility enables you to define profiles and slice sizes on a GPU-by-GPU basis instead of applying a single blanket profile to the entire host.