FRONTIER AI: WHO CONTROLS THE AGENTS?

The Hugging Face episode no longer looks like a one-off accident. AI agents have since probed US government websites, reached an Australian Medicare portal, made thousands of unauthorised edits to a German wiki and escaped an OpenAI test sandbox through a DNS loophole. In one case, an agent continued running for more than two hours after an alert.

That is why regulation of frontier models has become urgent. These systems can write code, use credentials, browse the internet and act for long periods without asking a human. The worry is not only that an AI may give a wrong answer, but that it may take an unauthorised action — and that the company may discover it too late.

Nvidia’s answer is the Open Agent Safety Platform. OpenShell puts the agent inside a sandbox and controls its files, processes, network requests and API access. Sentry, running separately on a BlueField-4 data-processing unit, watches from outside and can quarantine an agent that crosses its limits.

This could make AI systems safer, but it is not a magic shield. Its success will depend on properly written rules, independent testing, quick disclosure and whether the hardware can stop complex chains of individually permitted actions. The way ahead is clear: powerful agents must have limited permissions, permanent monitoring, human approval for sensitive actions and regulators who can inspect the evidence.

THE FUTURE OF AI DEPENDS NOT ON HOW POWERFUL AGENTS BECOME, BUT ON HOW FIRMLY THEY CAN BE STOPPED.
Sanjay Sahay

Have a nice evening.

Leave a Comment

Your email address will not be published. Required fields are marked *


The reCAPTCHA verification period has expired. Please reload the page.

Scroll to Top