Troubleshooting
GPU Memory Limit Not Enforced
If a container exceeds its nvidia.com/gpumem limit, check the following causes:
-
CUDA_DISABLE_CONTROL=trueis set - disables HAMi-core enforcement entirely. Remove it from production workloads. -
Docker-in-Docker (DinD) - inner containers do not inherit the
/etc/ld.so.preloadhostPath mount. HAMi enforcement does not apply inside DinD. -
Direct driver API usage - workloads calling NVML or the CUDA Driver API directly bypass
libvgpu.so. -
nvidia-container-runtimenot set as default - verify with:containerd config dump | grep default_runtime_nameThe output must show
nvidia. If not, follow the Prerequisites guide. -
If you don’t explicitly request vGPUs when using the device plugin with NVIDIA images, all GPUs on the host may be exposed to your container.
-
Currently, A100 MIG can be supported in only "none" and "mixed" modes.
-
Tasks with the "nodeName" field cannot be scheduled at the moment; please use "nodeSelector" instead.
-
Only computing tasks are currently supported; video codec processing is not supported.
-
Since v2.3.10, HAMi has changed the
device-pluginenvironment variable name fromNodeNametoNODE_NAME. If you are using an image version earlier than v2.3.10, thedevice-pluginmay fail to start.To resolve this issue, you have two options:
-
Manually edit the DaemonSet using
kubectl edit daemonsetand update the environment variable fromNodeNametoNODE_NAME. -
Upgrade the
device-pluginimage to the latest version using Helm:helm upgrade hami hami/hami -n kube-systemThis will apply the fix automatically.
-
NVIDIA containers fail with GPU Operator 25.10+
Use this section when the HAMi Device Plugin or a HAMi-scheduled NVIDIA workload stops starting after GPU Operator is installed or upgraded.
Problem 1: The HAMi Device Plugin fails to start
Identify the cause
Check the Device Plugin logs:
kubectl logs -n kube-system \
-l app.kubernetes.io/component=hami-device-plugin \
--all-containers --tail=200
Match the output to one of these errors:
| Error in the log | Cause |
|---|---|
Incompatible strategy detected auto | The Device Plugin cannot discover NVML devices because its container did not receive the NVIDIA driver and devices. |
invalid device discovery strategy | The Device Plugin could not initialize NVIDIA device discovery, usually for the same runtime-injection reason. |
failed to locate libcuda.so or failed to locate libnvidia-ml.so | HAMi cannot find the driver libraries under the configured driver root while generating a CDI specification. |
Confirm the runtime and CDI configuration:
kubectl get clusterpolicy -o yaml | grep -A 5 'cdi:'
kubectl get runtimeclass nvidia
kubectl get pods -n kube-system \
-l app.kubernetes.io/component=hami-device-plugin \
-o custom-columns=NAME:.metadata.name,RUNTIMECLASS:.spec.runtimeClassName
With GPU Operator 25.10+, CDI is normally enabled, the nvidia RuntimeClass must exist, and the HAMi Device Plugin must show nvidia in the RUNTIMECLASS column.
Solution
Configure the nvidia RuntimeClass for HAMi and restart the Device Plugin:
helm upgrade hami hami-charts/hami \
--namespace kube-system \
--reuse-values \
--set devicePlugin.runtimeClassName=nvidia
kubectl rollout restart daemonset/hami-device-plugin -n kube-system
kubectl rollout status daemonset/hami-device-plugin -n kube-system
If the logs report missing driver libraries while HAMi CDI is enabled, use the GPU Operator paths:
devicePlugin:
runtimeClassName: nvidia
deviceListStrategy: cdi-annotations
nvidiaDriverRoot: /run/nvidia/driver
nvidiaHookPath: /usr/local/nvidia/toolkit/nvidia-ctk
Wait until the NVIDIA driver and Toolkit DaemonSets are ready, then restart the HAMi Device Plugin. For host-installed drivers, set devicePlugin.nvidiaDriverRoot to / instead.
Problem 2: A HAMi-scheduled Pod fails to start
Identify the cause
Inspect the Pod events and its assigned RuntimeClass:
kubectl describe pod <pod-name> -n <namespace>
kubectl get pod <pod-name> -n <namespace> \
-o custom-columns=NAME:.metadata.name,RUNTIMECLASS:.spec.runtimeClassName
Use the error text to select the correct path:
| Error in the Pod events | Cause |
|---|---|
libcuda.so.1: cannot open shared object file | The container started without the NVIDIA driver libraries. |
unresolvable CDI devices management.nvidia.com/gpu=GPU-... | The NVIDIA runtime selected a GPU management CDI device, but the corresponding device could not be generated or resolved. |
unresolvable CDI devices k8s.device-plugin.nvidia.com/gpu=GPU-... | HAMi returned a CDI device, but the runtime cannot find a matching HAMi-generated CDI specification. |
Solution
First determine which HAMi injection mode is configured:
helm get values hami -n kube-system | grep -A 5 'devicePlugin:'
- For the default
devicePlugin.deviceListStrategy=envvarmode, setdevicePlugin.runtimeClassName=nvidiaby using the Helm command from Problem 1. This makes the NVIDIA runtime process the UUID returned throughNVIDIA_VISIBLE_DEVICES. - For
devicePlugin.deviceListStrategy=cdi-annotations, apply all four CDI values shown in Problem 1. Then inspect/var/run/cdi/k8s.device-plugin.nvidia.com-gpu.jsonon the node and verify that it contains the allocated GPU UUID. - For a host-installed Container Toolkit, confirm that the
nvidiaruntime is present in the active container runtime configuration. Restart the container runtime after correcting its configuration.
Do not mix HAMi CDI annotations with a missing or stale HAMi CDI specification. See NVIDIA CDI support for the complete setup and verification procedure.
Why this happens
Starting with GPU Operator 25.10.0, CDI is enabled by default and the Operator no longer makes the nvidia runtime the default runtime.
Before 25.10.0, GPU Operator normally configured the NVIDIA runtime as the default. As a result, the NVIDIA runtime hook processed every Pod and injected the devices and driver libraries selected through NVIDIA_VISIBLE_DEVICES.
With 25.10.0 and later, runc remains the default runtime and the container runtime uses native CDI for standard Device Plugin workloads. GPU management containers that access GPUs through NVIDIA_VISIBLE_DEVICES, including the HAMi Device Plugin, must explicitly use runtimeClassName: nvidia.
HAMi supports two device-injection paths:
| HAMi mode | Allocation result | Runtime requirement |
|---|---|---|
envvar (default) | HAMi writes the allocated GPU UUID to NVIDIA_VISIBLE_DEVICES. | On GPU Operator 25.10+, the Pod must use the nvidia RuntimeClass. |
cdi-annotations | HAMi returns a CDI device named k8s.device-plugin.nvidia.com/gpu=GPU-... and generates its CDI specification on the node. | The container runtime must have CDI enabled and be able to read the current HAMi specification. |
The HAMi chart applies devicePlugin.runtimeClassName both to the Device Plugin and to NVIDIA workloads mutated by the HAMi scheduler. This is why setting it to nvidia fixes the management container and keeps the workload runtime path consistent.
For new clusters, GPU Operator is recommended because it provides one entry point for configuring and upgrading the driver, Container Toolkit, and monitoring components. If these components are already installed on the hosts and you maintain their runtime configuration yourself, GPU Operator is optional; follow Prerequisites.
Disable the GPU Operator Device Plugin when using HAMi. Both plugins advertise nvidia.com/gpu and must not run on the same nodes.
devicePlugin:
enabled: false
Pod Stuck in Pending
If a Pod requesting GPU resources stays in Pending, check the following before assuming the cluster lacks capacity.
- Confirm the Pod is using HAMi's scheduler - HAMi's mutating webhook only rewrites
schedulerNamefor Pods whose resource requests it recognizes as HAMi-manageable. If it doesn't recognize the request, the Pod falls through to the default Kubernetes scheduler silently, and none of the failure reasons below will apply. Check with:
kubectl get pod <pod-name> -o jsonpath='{.spec.schedulerName}'
If this does not return HAMi's scheduler name, confirm your resource requests use HAMi's expected resource names (e.g. nvidia.com/gpu), and confirm the webhook itself is running with kubectl get pods -n kube-system | grep hami-scheduler.
-
Confirm the correct device plugin is installed - HAMi requires its own customized device plugin per accelerator vendor. The stock/official vendor device plugin is not compatible and will produce unexpected behavior rather than a clean error. Check with
kubectl get pods -n kube-system -o wide | grep device-pluginand confirm the image matches HAMi's documented device plugin for your vendor. -
Read the failure reason in Pod events - as of HAMi v2.7.0, the scheduler reports a specific reason for each rejected node directly through Pod events, instead of only a generic "no available node" message. Check with
kubectl describe pod <pod-name>and look under Events. If you're running an earlier version, Pod events show only the generic message; check the scheduler logs directly instead withkubectl logs -f <hami-scheduler-pod-name> -n kube-system -c vgpu-scheduler-extender.
Common failure reasons and what they mean:
NodeInsufficientDevice- the Pod requested more devices (by count) than the node has at all.CardTypeMismatch- the Pod's requested device type doesn't match the candidate card's actual type.CardInsufficientMemory- the card doesn't have enough free memory for the request. HAMi computes free memory asTotal memory - Used memoryand compares it against the requested amount (set viagpumem, or as a percentage of total device memory viagpumem-percentageifgpumemis unset).NumaNotFit- NUMA-aware scheduling is enabled, and the candidate device's NUMA node doesn't satisfy the Pod's topology requirement.
Worked example
apiVersion: v1
kind: Pod
metadata:
name: gpu-pod
spec:
containers:
- name: worker01
image: ubuntu:22.04
command: ["bash", "-c", "sleep 86400"]
resources:
limits:
nvidia.com/gpu: "1"
nvidia.com/gpumem: "3000"
nvidia.com/gpucores: "30"
Replace 3000 with a value higher than the free memory on any candidate node in your cluster, so the request genuinely exceeds capacity. Apply this Pod, then check its events:
kubectl apply -f gpu-pod.yaml
kubectl describe pod gpu-pod
You should see a FilteringFailed event whose message includes CardInsufficientMemory. For the specific total/used memory numbers behind the failure, check the scheduler logs:
kubectl logs -f <hami-scheduler-pod-name> -n kube-system -c vgpu-scheduler-extender
Adjust the request to fit within available capacity and reapply; the Pod should transition to Running.