Version: v2.9.0
Quick Start
Get HAMi up and running in minutes by deploying the Helm chart and submitting your first shared GPU workload.
Prerequisites
Before deploying HAMi, ensure your GPU nodes meet the following prerequisites:
- Helm v3+
- kubectl v1.23+
- CUDA v10.2+
- NVIDIA Driver v440+
- NVIDIA Container Toolkit (with
nvidia-container-runtimeset as default runtime)
1. Label your nodes
Label the target GPU nodes with gpu=on. Nodes without this label will not be managed by HAMi:
kubectl label nodes <node-name> gpu=on
2. Deploy HAMi using Helm
Add the official HAMi Helm repository and deploy the chart:
helm repo add hami-charts https://project-hami.github.io/HAMi/
helm repo update
helm install hami hami-charts/hami -n kube-system
Verify that the hami-scheduler and hami-device-plugin pods are running:
kubectl get pods -n kube-system | grep hami
3. Submit a vGPU Workload
Save the following manifest as gpu-pod.yaml to request 1 vGPU with 10240 MiB of GPU memory limit:
apiVersion: v1
kind: Pod
metadata:
name: gpu-pod
spec:
containers:
- name: ubuntu-container
image: ubuntu:22.04
command: ["bash", "-c", "sleep 86400"]
resources:
limits:
nvidia.com/gpu: 1
nvidia.com/gpumem: 10240
Apply the manifest and wait for the Pod to become ready:
kubectl apply -f gpu-pod.yaml
kubectl wait --for=condition=Ready pod/gpu-pod --timeout=120s
4. Verify GPU Memory Isolation
Execute nvidia-smi inside the running container:
kubectl exec -it gpu-pod -- nvidia-smi
Expected output showing HAMi-core hard memory limit (10240MiB):
[HAMI-core Msg]: Initializing.....
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 550.54.15 Driver Version: 550.54.15 CUDA Version: 12.4 |
|-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
|=========================================+========================+======================|
| 0 Tesla V100-PCIE-32GB On | 00000000:3E:00.0 Off | 0 |
| N/A 29C P0 24W / 250W | 0MiB / 10240MiB | 0% Default |
+-----------------------------------------------------------------------------------------+
Cleanup
Delete the test Pod:
kubectl delete pod gpu-pod
Next steps
- Verify your setup in detail with Verify HAMi Installation.
- Learn how to customize your deployment parameters in the Configuration Guide.
- Understand how GPU sharing works under the hood in Device Sharing.