NVIDIA Lepton
NVIDIA Lepton
GPUd
Container
NVIDIA Lepton
NVIDIA Lepton
GPUd

GPUd is designed to ensure GPU efficiency and reliability by actively monitoring GPUs and effectively managing AI/ML workloads. License: Apache License 2.0

Overview

GPUd is an open-source daemon for monitoring NVIDIA GPU health and the surrounding host and container runtime. It collects telemetry, detects hardware and software errors, and provides diagnostic data for individual nodes and large GPU fleets.

GPUd can run directly on a host, as a systemd service, or as a Kubernetes DaemonSet through the official Helm chart.

Core features

AreaCapabilities
GPU telemetryMonitors utilization, memory, power, temperature, clock speed, processes, and remapped rows.
Error detectionDetects Xid, SXid, ECC, hardware slowdown, and other errors through kernel logs, NVML, nvidia-smi, and optional DCGM checks.
GPU fabricMonitors NVLink, NVSwitch, InfiniBand, Fabric Manager, NCCL, and peer-memory status where available.
Host monitoringReports CPU, memory, disk, operating system, PCI, network, and file-descriptor health.
Runtime monitoringTracks Docker, containerd, and kubelet status alongside GPU workloads.
ObservabilityExposes health information and metrics for local diagnostics and integration with centralized monitoring systems.

See the component reference for the full list of checks.

Typical use cases

  • Monitor GPU health across bare-metal, virtual-machine, and Kubernetes fleets.
  • Detect Xid, ECC, thermal, power, and fabric problems before assigning new workloads to a node.
  • Distinguish software faults that may be resolved by a restart from hardware faults that may require repair.
  • Validate node health after driver changes, reboots, or hardware maintenance.
  • Correlate GPU failures with host, container-runtime, and Kubernetes state.
  • Export GPU and host telemetry to an existing observability platform.

GPUd complements tools such as node_exporter and dcgm-exporter by combining GPU health checks, host monitoring, and operational diagnostics in one agent. See the project's design rationale and comparisons.

Deployment options

  • Linux host: Install the standalone binary and run GPUd with systemd.
  • Kubernetes: Deploy GPUd as a DaemonSet using the official Helm chart.
  • Container environments: Monitor Docker, containerd, and kubelet from the same agent.
  • DGX Cloud Lepton: Connect GPUd to the platform with a token for centralized health reporting. GPUd is used in DGX Cloud Lepton production infrastructure.

System requirements

AreaMinimum requirement or limitation
Operating systemThe documented installation path targets Linux on amd64. The installer also supports Linux arm64 and macOS, although some Linux-specific checks are unavailable on macOS.
GPU hardwareAn NVIDIA GPU is required for GPU-specific telemetry and health checks. GPUd can still collect supported host metrics when no GPU is present.
NVIDIA softwareGPU checks require a working NVIDIA driver with nvidia-smi and NVML. Optional checks may also use DCGM, NCCL, Fabric Manager, InfiniBand, or nvidia-peermem.
Driver versionGPUd does not publish one universal minimum NVIDIA driver version. Use a driver supported by the installed GPU and any optional NVIDIA components enabled on the host.
KubernetesKubernetes 1.19 or later and Helm 3.16.1 or later.
Lepton-connected nodeWhen joining DGX Cloud Lepton with --token: at least 3 CPU cores and 4 GiB of memory. The recommended allocation is 4 or more CPU cores and 8 GiB or more of memory. Standalone GPUd has lower requirements.

See the installation documentation for current platform details.

Project and support

GPUd is maintained by Lepton AI with contributions from the open-source community. Copyright is held by Lepton AI Inc., as recorded in the project's NOTICE.

License

GPUd is released under the Apache License 2.0.

Publisher
NVIDIA Lepton
NVIDIA Lepton
Latest Tag0.12.15
UpdatedJuly 23, 2026 UTC
Compressed Size1.51 GB
Multinode SupportNo
Multi-Arch SupportYes

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.