Thousands of Nvidia GPU Monitors Were Left Wide Open on the Internet

A flaw in Nvidia's DCGM Exporter, the software companies use to watch over their AI computing hardware, let anyone crash the monitoring service or profile what machines an organisation was running, no password required.

ThreatVectr NewsdeskAI-assistedPublished · Editor: Lee Brown· 4 min read
Illustration: a large data centre floor at night
Illustration made with AI. Not a photograph of the events described.
Share

Key points

  • Researchers at Lava Security found more than 2,000 GPU servers running Nvidia DCGM Exporter exposed directly to the public internet, with no authentication required to reach them.
  • Those servers collectively reported more than 12,000 unique graphics processing units (GPUs), the specialised chips that power AI work, representing an estimated $100 million in hardware.
  • CVE-2026-47483, published with a severity score of 8.2 out of 10, describes how an attacker can crash the monitoring service and extract information about the underlying infrastructure.
  • Nvidia's fix landed in DCGM Exporter version 4.8.2; operators still running older versions should update immediately and verify that the profiling feature is switched off.
  • Exposed machines included the Nvidia Blackwell Ultra B300, H200, and consumer graphics cards used by smaller operators.

Nvidia DCGM Exporter is a small piece of software that runs quietly on servers stuffed with expensive AI hardware. Its job is to collect performance data from graphics cards and feed it into monitoring dashboards, the same way a car's dashboard reads engine temperature and fuel. Lava Security researchers scanning the public internet found thousands of these monitors sitting wide open, no login needed.

That exposure enabled two distinct problems.

How could attackers cause harm?

The simpler attack is a denial-of-service, meaning an outsider overwhelms the software until it stops working. DCGM Exporter contains internal diagnostic pages, paths starting with /debug/pprof, meant for developers troubleshooting performance. Because those pages required no authentication, anyone on the internet could hammer them with simultaneous requests and crash the exporter.

Losing the monitoring service sounds minor. Lava researcher Michael Katchinskiy pointed out the real risk in a blog post: the exporter shares a server with the AI workloads it watches. The same memory pressure that kills the monitor can spill over and disrupt AI training or inference jobs running alongside it. A single server housing dozens of high-end GPUs going dark mid-job is an expensive problem.

The second issue is quieter. Even without crashing anything, a curious outsider could simply read the telemetry those open endpoints serve. GPU utilisation, power draw, and error rates were all visible, along with hardware model numbers. Katchinskiy, whose findings were first reported by CSO Online, put it plainly: anyone who reached those endpoints could see exactly what hardware an organisation was running and how hard it was working. That detail doesn't hand attackers a model's training data, but it lets them profile a target, spot older software versions, and plan a more focused follow-up.

Among the exposed machines Lava identified were Nvidia's H100 and H200 data-centre chips alongside the newer Blackwell Ultra B300, all used for large-scale AI model training. Consumer RTX 5090 and 4090 cards appeared in the data too, pointing to smaller teams running GPU workloads without dedicated security staff on the configuration side.

We covered a related Nvidia patch two weeks ago in AMD, Arm, and Nvidia Fix Security Flaws in Graphics and AI Chips; this finding shows the exposure was already live in the wild while that story was being written.

What should operators do right now?

Update to DCGM Exporter version 4.8.2 or later. That release addresses CVE-2026-47483 by controlling how the profiling endpoints handle concurrent requests.

Updating alone isn't enough.

Action Detail
Update software DCGM Exporter 4.8.2 or later
Disable profiling Confirm --enable-pprof flag is off
Restrict network access Bind exporter to loopback or private network interface only
Enforce firewall rules Block public traffic to exporter ports via security groups or equivalent

The failure mode here is a familiar one in cloud and AI infrastructure: a developer tool with no authentication ships with a feature enabled by default, gets deployed at speed to production, and sits exposed until a third party runs a scan. The teams buying $100 million worth of GPU hardware are often not the same teams writing the firewall rules.

Check your network exposure before you check your model accuracy.

© 2026 Threat Vectr