Securing Long-Running AI Agents: From Setup to Sandboxing

Securing Long-Running AI Agents: From Setup to Sandboxing

🎙 Adel El Hallak and Kris Murphy 👥 222K 📅 July 2, 2026 ⏱ 45 min 👁 3K 📄 expert opinion 🧭 2026-08-13
Available in: English (current) Français

Keywords

long-running agentssandboxingpolicy controlsNVIDIA NeMoOpen Shell

Summary

The presentation by NVIDIA’s Adel El Hallak and Kris Murphy addresses the security challenges of long-running AI agents. They begin by highlighting the evolution from ChatGPT to reasoning models and the recent ‘Claw moment’ where autonomous agents became widely adopted. The core message is that agents, composed of models and a harness, require robust security measures before production deployment. They introduce NVIDIA’s agent toolkit, including NeMo Tron models, skills, and the Open Shell secure runtime. NeMo Tron Ultra is presented as an open model with transparent training data and algorithms. Skills package NVIDIA libraries as agent-friendly manuals, with examples like the AIQ deep research skill. Open Shell is an open-source project providing a secure runtime with policy controls, secret management, and sandboxing. It allows users to approve or deny agent actions, and integrates with operating systems like Windows and Canonical. NeMo Claw offers blueprints for deploying agents with Open Shell, supporting various models. The talk includes a demo of the Open Shell UI and a case study with ServiceNow showing 90% of L1 tickets resolved autonomously. The presenters emphasize the importance of trust and control in agentic systems.

188 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation offers valuable insights into the practical security requirements for deploying AI agents in enterprise settings. The argumentation is solid, grounded in real-world examples and product demonstrations. The speakers effectively argue that long-running agents need identity, access, sandboxing, and policy controls, and they present Open Shell as a solution. The value lies in the concrete guidance on securing agents, such as separating secret management from the agent environment and using policy approval workflows. The argumentation is persuasive, though it is primarily a product pitch rather than an unbiased analysis.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is moderate; the talk is based on the presenters’ expertise and NVIDIA’s product development, but it lacks citations to external research or formal studies. The sources cited are NVIDIA documentation and cybersecurity solutions pages, which are relevant but not independent. The title accurately reflects the content, focusing on setup and sandboxing. The presentation is well-structured and technically coherent, but the lack of peer-reviewed references limits its scientific rigor.

176 words

Title / Content Match

The title accurately reflects the content, which focuses on securing long-running AI agents through setup, sandboxing, and policy controls.

Quality & Reliability

7/10

The talk is an expert presentation from NVIDIA developers, providing practical insights and product announcements. It lacks formal citations or peer-reviewed references, but the technical content is coherent and based on the presenters' direct experience. The claims about benchmarks and security features are plausible but not independently verified.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The presentation provides a practical overview of securing long-running AI agents, emphasizing the need for sandboxing and policy controls. It introduces NVIDIA’s Open Shell as an open-source solution and highlights the importance of separating secrets from the agent environment. The talk also showcases the AIQ deep research skill, demonstrating how multi-agent systems can achieve high performance. The novelty lies in the integration of these components into a cohesive toolkit for enterprise deployment.

Pour aller plus loin :

125 words

Radar Profile

The radar profile shows high scores in information quantity and technical level, indicating a content-rich presentation with moderate technical depth. The quality and reliability scores are slightly lower, reflecting the lack of independent sources. Overall, the talk is informative but relies on the presenter's expertise and product claims.

Reliability 7/10

💬 Sur les 0 commentaires analysés, aucune tendance n'est disponible.