NVIDIA DCGM Exporter flaw can crash GPU monitoring

NVIDIA DCGM Exporter flaw can crash GPU monitoring

A high-severity vulnerability in NVIDIA's DCGM Exporter (CVE-2026-47483) allows unauthenticated attackers to knock the GPU monitoring service offline and could interfere with AI workloads running on the same machine. The research comes from Lava, which also found hundreds of exposed GPU servers on the internet.

Lava reported the issue to NVIDIA. The company gave it a CVSS score of 8.2 and released a security bulletin on July 28, 2026.

What the exporter reveals

GPU servers are machines built around graphics processing units and are used for AI, machine learning and scientific computing. DCGM Exporter reads data straight from the GPUs on a host, such as temperature, utilization, memory usage, power draw and error events. It publishes this data over HTTP, usually on port 9400, so that tools like Prometheus can collect it.

The output also includes the exact GPU model and a unique ID (UUID) for each card. Lava researcher Michael Katchinskiy said this is enough for an attacker to learn what hardware a server runs, how busy it is and whether it is reporting errors. That led him to ask, "How many of them are exposed to the internet?"

Thousands of servers, no authentication

The team found more than 2,000 servers exposing the exporter publicly. Across four scans run between March and May 2026, these servers reported more than 12,000 unique GPUs, with an estimated value of $100 million.

"Every host returned metrics over plaintext HTTP, and none required authentication," Katchinskiy wrote.

The exposed hardware included NVIDIA's Blackwell Ultra B300, H200 and H100 GPUs, which are designed for large AI workloads, as well as consumer RTX 5090 and 4090 cards. Nearly 300 organizations were affected. The US accounted for 5,274 of the exposed GPUs (44%), ahead of Romania with 2,054 and China with 1,967.

Some exporters ran on customer infrastructure at GPU cloud providers, including Voltage Park, Lambda, Northern Data and DigitalOcean. Voltage Park had the largest share, with 672 public Node Exporter hosts and 71 public DCGM Exporter hosts. The company said most of the Node Exporter instances and all of the DCGM Exporter instances were deployed by customers, and its security team reached out to them.

Debug endpoints open the door to crashes

Around a quarter of the exposed hosts also served Go's /debug/pprof/ profiling endpoints next to /metrics. Some of these endpoints hold a request open for as long as the caller specifies, and many simultaneous unauthenticated requests drive up memory use.

The researchers first assumed this was an operator misconfiguration. They then reproduced the behavior with NVIDIA's official DCGM Exporter container, unmodified. Any deployment exposing the exporter on a routable interface could expose /debug/pprof/ too.

"With enough concurrent unauthenticated requests, the exporter could run out of memory and crash, cutting off visibility into GPU health and activity," Katchinskiy explained. The added CPU and RAM pressure could also slow the host and disrupt training or inference jobs, especially where strict resource limits are missing.

Node Exporter adds more detail

Lava also looked at Prometheus Node Exporter, which tracks a server's hardware and operating system. It found 12,096 public hosts reporting metrics from NVIDIA/Mellanox InfiniBand and RoCE network adapters, including models, firmware versions, link state and fabric activity. Some leaked hostnames, OS and kernel versions, and BIOS details. One endpoint alone identified a Dell PowerEdge XE9680 running Ubuntu 22.04.5 LTS with an active 400 Gb/s port.

"Exact versions let an attacker go straight to matching known vulnerabilities," Katchinskiy noted.

Fixes

Users should upgrade DCGM Exporter to version 4.8.2 or later and keep the -enable-pprof flag off unless profiling is needed. In current versions, profiling is opt-in. Lava advises keeping Node Exporter, DCGM Exporter and Prometheus off the public internet, binding exporters to loopback or private interfaces, and restricting access with firewall rules or security groups. Prometheus query APIs and target pages need the same protection.

Our Take

The CVE is the headline, but the bigger problem in Lava's findings is exposure. Monitoring tools are often treated as harmless plumbing. Here they gave anyone who asked a free map of expensive AI hardware, software versions and network details. This fits a wider pattern of attackers going after poorly secured AI infrastructure, such as the PoeLLM malware hijacking exposed AI servers for cryptomining.

The Voltage Park case also suggests that many teams are unclear on where the provider's responsibility ends and the customer's begins. It is worth watching whether GPU cloud providers start blocking public monitoring ports by default, and how quickly the exposed hosts get patched or pulled offline.