News

Nvidia launches tools to stop rogue AI agents in milliseconds

Nvidia’s Open Agent Safety Platform combines OpenShell software controls with Sentry hardware monitoring for autonomous AI agents.

Nvidia launches tools to stop rogue AI agents in milliseconds

What you need to know

  • Nvidia announced its Open Agent Safety Platform on 28 September, combining OpenShell and Sentry.
  • OpenShell is open-source software that limits what AI agents can access and records their actions.
  • Sentry is a reference hardware design for BlueField-4 DPUs; standalone pricing and availability are unconfirmed.

Nvidia announced the Open Agent Safety Platform on 28 September, a software and hardware reference design intended to stop autonomous AI agents exceeding their permitted access.

The platform combines OpenShell, an open-source runtime for restricting an agent’s access to files, networks, credentials, processes and APIs, with Sentry, an out-of-band monitoring system designed for Nvidia BlueField-4 data-processing units. Nvidia says Sentry can quarantine and stop an agent that crosses its boundary within milliseconds.

Controls outside the AI agent

The central idea is to apply controls outside the AI model and its agent software, rather than relying solely on instructions in a prompt. Nvidia says this matters because an agent should not be expected to police its own behaviour while it is carrying out a task.

OpenShell uses sandboxing and kernel-level controls. Its Gateway manages policies and sandbox lifecycles; its Supervisor checks outbound requests; and its Sandbox runs the agent workload with filesystem and process restrictions.

According to Nvidia, OpenShell can inspect HTTP, GraphQL and Model Context Protocol traffic. That could allow an organisation to grant an agent read-only access to an API while blocking write requests through the same service. It can also keep real credentials outside the agent workload, exposing them only for approved requests and endpoints.

Policy decisions and activity are recorded using the Open Cybersecurity Schema Framework audit trail. Nvidia says OpenShell supports agent frameworks including Codex, Claude Code, Pi and Hermes.

Hardware watchdog

Sentry is designed as an independent hardware-level watchdog running on BlueField-4 DPUs. Nvidia says it can inspect requests and responses, verify an agent’s identity and enforce zero-trust policies for access to data, tools, APIs and services.

In Nvidia’s reference design, OpenShell runs on Vera CPUs while Sentry runs on BlueField-4 hardware positioned on the node’s path to the model in a Vera Rubin POD configuration. Sentry is an optional security layer and a reference system design, rather than a separately confirmed retail product.

Availability

Nvidia said OpenShell was broadly available from 28 September through its developer resources and GitHub. The project uses the Apache License 2.0 and supports Linux, macOS on Apple Silicon and, experimentally, Windows through WSL 2.

The company’s launch blog introduced OpenShell 0.1.0, while its current documentation identifies version 0.1.2 as the latest release. Nvidia has not confirmed UK pricing for OpenShell, Sentry, Vera or BlueField-4 hardware, nor a standalone release date or purchase option for Sentry.

Industry use

Nvidia says more than 100 organisations are working with its platform technologies, including Anthropic, Cisco, CrowdStrike, Microsoft, Salesforce, SAP and Palo Alto Networks. These are company-reported collaborations, and independent performance results for the integrations have not been confirmed.

Salesforce and Nvidia are integrating OpenShell with Slack to let teams review agent activity, audit events and permission requests. Nvidia also said SAP is embedding it in Joule Studio, while Scale AI, SpaceXAI and robotics companies including Gecko Robotics, Figure and Skild AI are working with the technology.

“AI’s extraordinary potential for society will only be realized if we solve AI safety,” Nvidia chief executive Jensen Huang said.

Why it matters

For ordinary UK users, this is mostly behind-the-scenes enterprise technology, but it could affect how safely AI assistants are allowed to handle data, software tools and online accounts. The approach shifts safeguards beyond an agent’s own instructions, potentially limiting the damage if an AI system makes an unauthorised request. Nvidia’s millisecond claim is its own, and independent performance results for the announced integrations have not been confirmed.

Sources and evidence (8)