NVIDIA
DCGM Exporter
Container
NVIDIA
DCGM Exporter

Monitor GPUs in Kubernetes using NVIDIA DCGM. This is an exporter for a Prometheus monitoring solution in Kubernetes.

  • Overview

    Monitoring stacks usually consist of a collector, a time-series database to store metrics and a visualization layer. A popular open-source stack is Prometheus used along with Grafana as the visualization tool to create rich dashboards. Prometheus is deployed along with kube-state-metrics and node_exporter to expose cluster-level metrics for Kubernetes API objects and node-level metrics such as CPU utilization.

    NVIDIA DCGM

    NVIDIA DCGM is a set of tools for managing and monitoring NVIDIA GPUs in large scale linux based cluster environments. It's a low overhead tool that can perform a variety of functions including active health monitoring, diagnostics, system validation, policies, power and clock management, group configuration and accounting.

    DCGM Exporter

    DCGM-Exporter is an exporter for Prometheus to monitor the health and get metrics from GPUs. It leverages DCGM using Go bindings to collect GPU telemetry and exposes GPU metrics to Prometheus using an http endpoint (/metrics). DCGM-Exporter can be used either standalone or deployed as part of the NVIDIA GPU Operator.

    Government Ready: STIG/FIPS Hardening

    This ensures the highest level of security and compliance for regulated environments, the container image is:

    • hardened

    Learn more about NVIDIA's hardened image in the AI Software for Regulated Environments White Paper.

    Usage

    For using the DCGM-Exporter, visit the user guide

    License Agreements

    Suggested Reading

    Get Help

    Enterprise Support

    Get access to knowledge base articles and support cases or submit a ticket.

    Publisher
    NVIDIA
    LicenseNVIDIA proprietary
    Latest Tag4.8.3
    UpdatedJuly 14, 2026 UTC
    Compressed Size56.87 MB
    Multinode SupportNo
    Multi-Arch SupportYes