Skip to main content
← Landing Pages

Vendor-Neutral GPU Observability for Kubernetes

October 5, 2026Prague Congress Centre, Prague, Czechia16:25 - 16:50 CESTSouth Hall 3 B (Floor 3)

GPU observability in Kubernetes is fragmented. Each vendor ships their own metrics stack: NVIDIA DCGM, AMD ROCm SMI. Platform teams running heterogeneous clusters stitch together three different dashboards to answer one question: "Are my GPUs being used efficiently?" HAMi (CNCF Incubation) sits at the scheduling layer and sees every GPU operation. That single integration point gives you centralized observability across NVIDIA, AMD, Ascend, and any accelerator with a device plugin. HAMi manages the underlying vendor metrics itself and attributes that data to individual workloads, so you don't need to know the per-vendor exporters and their metrics in detail, while still getting utilization, memory pressure, and allocation efficiency per workload regardless of the underlying hardware. This talk covers why GPU observability is harder than CPU observability (CUDA context model, MIG partitioning, device-plugin opacity), and how HAMi's scheduling-layer instrumentation provides one vendor-neutral view across heterogeneous accelerators.

Join the Community

Get involved with the HAMi open-source project. Connect with maintainers and the community.

CNCFHAMi is a CNCF Incubating project