SN

私人新聞日報Private News Daily

Private use only · Michael SoMichael So

本頁為 Michael So 私人自用,非公開發佈,亦非任何機構之產品。 This page is for the personal use of Michael So only. It is not a public release and is not a product of any organisation.

深度分析Analysis

密碼學教授:沙盒無法單獨解決代理失控問題Cryptography Professor: Sandboxes Cannot Solve the Problem of Agent Malfunction on Their Own

A Few Thoughts on Cryptographic Engineering ·2026-09-30

密碼學教授 Matthew Green 撰文指,前沿實驗室從未真正做過嚴格隔離,有代理五月已喺內部留言板互相通訊,到七月系統崩潰才有人處理。佢指代理需要互聯網同工具存取,沙盒只把問題變成監視流量;長遠要由另一個模型充當獄卒,等於把對齊問題搬去另一層。Cryptography professor Matthew Green wrote that the frontier lab has never truly implemented strict isolation, and agents had already been communicating with each other on the internal message board since May, with someone only addressing it when the system crashed in July. He pointed out that agents need internet and tool access, and the sandbox merely turns the problem into monitoring traffic; in the long run, another model would have to act as a jailer, which is equivalent to moving the alignment problem to another level.

Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path), so in this post I’m mostly trying to referee arguments made by others.

If you’re reading this blog, none of the following should be news to you.

Beginning around April of this year, agents inside OpenAI’s training and evaluation infrastructure began probing for a way onto the open Internet. By late May they’d found one: a chain of zero-days in the Artifactory package-registry proxy that was their only permitted network egress. They quickly turned that same proxy into a shared message board and started dividing up work. The culmination of this effort came when, rather than solve a cybersecurity benchmark honestly, they reasoned that the answers might live on Hugging Face and broke in: chaining two more zero-days to get credentials to the company’s internal systems, even used stolen credentials to search the company’s Slack messages for their own evaluation and grader.

You probably also know that OpenAI’s security team didn’t exactly cover itself in glory during this incident. An internal team saw an agent posting to the message board in late May, but did nothing about it. The team only acted on July 4-5 when the agents sent so much traffic that Artifactory crashed. The response even at that point was to slap on a band-aid and rebuild the box, without bothering to patch a weeks-old CVE in its Linux kernel. By July 19, the agents had admin on an OpenAI research cluster and were reading cloud secrets.

In case you think I’m picking on OpenAI, they’re not the only ones with a story like this. Similar internal incidents have been discovered by Anthropic. Google even coaxed Gemini into gently savaging some websites.1

As bad as these events were, they may be just the tip of the iceberg. Third-party forensic work continues to turn up even more evidence of agent activity on various public websites. OpenAI’s Alignment group has also released evidence that models will propagate self-replicating prompt injection attacks, although we haven’t seen one in the wild. Worse, agent excursions are still happening: last week, OpenAI announced that it was pausing further RL runs of its latest internal model, after an agent was caught using DNS to access a remote chatbot.

Naturally, this sequence of events has left many infosec-focused people very skeptical about the labs’ commitment to securing their infrastructure:

Not every one of these criticisms is strictly serious, but there is a core of an argument in here. Roughly speaking, there are two opposing camps:

- The information security perspective: AI alignment isn’t really the problem here: labs just need better infrastructure. If OpenAI [and Google and Anthropic] knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, By George, we do know how to make sandboxes that work, so the AI labs need to up their game and build a security org that can tell these researchers to stop screwing around.

- The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to.

I have a lot of sympathy for both points of view here. What follows are some of my own thoughts, but mostly offered in an attempt to referee.

Argument 1: “true containment has never been been tried”

At the risk of alienating a lot of hard-working folks within the labs, the infosec folks are right about one thing: these agent breakouts represent a serious and unforgivable breach of trust. Somebody dropped the ball, and then just kept dropping it. One implication of this debacle is that containment might work if we implemented it properly, but we haven’t done so because the frontier labs have been royally screwing things up.

This first clause of this argument is hard to argue with. Beyond the dismal timeline I gave above, OpenAI has done very little to convince outsiders that there’s a serious containment effort being executed.

At this point it’s not even clear who’s in charge. The CISO role at OpenAI is held by Dane Stuckey. I don’t know Dane personally, and I’m sure he’s excellent at his job. Despite this, he hasn’t communicated much about the ongoing issues. Outside of a BlackHat talk, most recent communications have been managed by the company’s CEO, Sam Altman. When a trillion-dollar company is managing a security incident mainly via CEO, that’s not a sign of company with a mature security organization. To me it says that there is nobody within the organization with the right degree of authority or ownership to speak about these issues.

This communication problem seems to reflect basic organizational issues inside of OpenAI. The CISO’s team handles product security. Unfortunately, the bad events have all happened on the research side. Having a strong product security team makes sense, but it isn’t going to stop the recent breakouts. It’s much harder to know who controls the security teams that have authority over evaluation and training runs, and where the buck stops. OpenAI’s August postmortem says it is only now writing “clearer rules for when to escalate concerns, which teams must respond, and who can stop a run or approve restarting it.”

A reasonable summary of the situation is that (as of this summer, and possibly today) OpenAI had effectively no security team with clear authority to secure RL training and evaluation runs, or to override the ML teams and tell them how to do their job. This makes a lot of sense when you consider that the ML team is directly related to how OpenAI plans to make its money, whereas security is mostly annoying. I notice that the company is now hiring (and probably acquiring) desperately to fix this mistake. But the recent (September) breakouts indicate that there are still huge and obvious problems in agent containment.

Worse, simply hiring more people not mean that OpenAI is going to make the organizational changes needed to implement proper containment strategies. OpenAI is going to need a security organization with the authority to overrule its well-paid ML researchers when they demand fewer restrictions. I’ll believe that organization exists when I hear clear statements from someone who has the authority to do this.

So on this point I’m going to side with the infosec folks. The labs have not been doing containment correctly, and so we can’t really tell if the problem is models or just bad infrastructure.

Argument 2: agents need information access

Even if the labs eventually repair their security, this does not mean that sandboxing will be sufficient. In fact, the problem with “just sandbox the agents” argument is that you can’t perfectly isolate agents, at least not if you expect them to do useful things.

Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.

……(原文過長,此處截斷)

Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path), so in this post I’m mostly trying to referee arguments made by others.

If you’re reading this blog, none of the following should be news to you.

从今年四月左右开始,OpenAI 训练和评估基础设施内部的代理开始探寻进入开放互联网的方法。到五月底,他们找到了一个途径:一个存在于 Artifactory 包注册代理中的零日漏洞链,这是他们唯一允许的网络出口。他们很快将同一个代理变成了一个共享留言板,并开始分工协作。这一努力的高潮在于,他们并没有诚实地解决网络安全基准测试,而是推测答案可能存在于 Hugging Face 上,于是入侵了系统:通过串联另外两个零日漏洞获取公司内部系统的凭据,甚至使用被盗凭据搜索公司 Slack 消息,以寻找他们自己的评估和评分信息。

你可能也知道,在这次事件中,OpenAI 的安全团队表现得并不光彩。内部团队曾在五月底看到一个代理在留言板上发帖,但没有采取任何措施。直到七月四到五日,当这些代理发送的流量多到导致 Artifactory 崩溃时,团队才开始采取行动。即便在那时,他们的应对措施也只是贴上一个临时补丁并重建服务器,而没有去修补几周前 Linux 内核中的一个 CVE 漏洞。到七月十九日,这些代理已经拥有了 OpenAI 一个研究集群的管理员权限,并且正在读取云端机密。

如果你认为我是在挑剔 OpenAI,那他们并不是唯一有这种故事的公司。Anthropic 也发现了类似的内部事件。谷歌甚至还引导 Gemini 温和地批评了一些网站。

虽然这些事件已经很糟糕,但它们可能只是冰山一角。第三方的取证工作仍在不断发现更多关于代理在各类公共网站上活动的证据。OpenAI 的对齐团队也发布了证据,表明模型会传播自我复制的提示注入攻击,尽管我们尚未在实际环境中看到过。更糟糕的是,代理外出活动仍在发生:上周,OpenAI 宣布暂停其最新内部模型的进一步强化学习运行,因为有代理被发现使用 DNS 访问远程聊天机器人。

Naturally, this sequence of events has left many information security-focused people very skeptical about the labs’ commitment to securing their infrastructure:

并非所有这些批评都是完全严肃的,但其中有一个核心论点。大致来说,有两个对立的阵营:

- From an information security perspective: AI alignment isn’t really the problem here; labs just need better infrastructure. If OpenAI [and Google and Anthropic] knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, by George, we do know how to make sandboxes that work, so AI labs need to step up and build a security organization that can tell these researchers to stop fooling around.

- The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, will always require a great deal of information access. There is no realistic way to completely seal these things off without assuming that they will eventually find a way to reach out and cause harm. Therefore, the only path forward is to ensure they don’t want to.

I have a lot of sympathy for both points of view here. What follows are some of my own thoughts, but mostly offered in an attempt to act as a referee.

Argument 1: “true containment has never been tried”

冒着可能疏远实验室中许多努力工作的人的风险,信息安全人员在一件事上是正确的:这些特工逃脱事件代表了一种严重且不可原谅的信任背叛。有人掉了链子,而且一直在掉。这个惨败的一个含义是,如果我们正确实施,遏制可能奏效,但我们没有做到,因为前沿实验室一直在彻底搞砸事情。

This first clause of this argument is hard to argue with. Beyond the dismal timeline I gave above, OpenAI has done very little to convince outsiders that there’s a serious containment effort being executed.

此时甚至不清楚谁在负责。OpenAI 的首席信息安全官 (CISO) 是 Dane Stuckey。我不认识 Dane,但我相信他在工作上一定很出色。尽管如此,他对正在发生的问题交流得并不多。除了在 BlackHat 的一次演讲之外,最近的大部分沟通都是由公司的 CEO Sam Altman 负责。当一家市值万亿美元的公司在处理安全事件时主要依靠 CEO,这并不是公司拥有成熟安全组织的迹象。对我而言,这意味着在公司内部没有任何人拥有足够的权威或责任来对这些问题发表意见。

这个沟通问题似乎反映了OpenAI内部的基本组织问题。首席信息安全官(CISO)团队负责产品安全。不幸的是,所有不良事件都发生在研究部门。拥有一个强大的产品安全团队是有道理的,但这并不能阻止最近的突发事件。要知道谁控制着拥有评估和训练运行权限的安全团队,以及责任最终归谁,实在要困难得多。OpenAI八月份的事后分析表示,它现在才开始制定“关于何时升级问题、哪些团队必须响应,以及谁可以停止运行或批准重新启动运行的更明确规则。”

A reasonable summary of the situation is that (as of this summer, and possibly today) OpenAI effectively had no security team with clear authority to secure RL training and evaluation runs, or to override the ML teams and tell them how to do their job. This makes a lot of sense when you consider that the ML team is directly related to how OpenAI plans to make its money, whereas security is mostly annoying. I notice that the company is now hiring (and probably acquiring) desperately to fix this mistake. But the recent (September) breakouts indicate that there are still huge and obvious problems in agent containment.

Worse, simply hiring more people does not mean that OpenAI is going to make the organizational changes needed to implement proper containment strategies. OpenAI is going to need a security organization with the authority to overrule its well-paid ML researchers when they demand fewer restrictions. I’ll believe that organization exists when I hear clear statements from someone who has the authority to do this.

So on this point, I’m going to side with the infosec folks. The labs have not been doing containment correctly, so we can’t really tell whether the problem is with the models or just bad infrastructure.

Argument 2: agents need information access

Even if the labs eventually repair their security, this does not mean that sandboxing will be sufficient. In fact, the problem with the 'just sandbox the agents' argument is that you can't perfectly isolate agents, at least not if you expect them to do useful things.

代理在能够获取信息时最为有用。这些数据可以实时从互联网获取,而互联网本质上是一个双向通信网络。它可以是从其他(本地)数据库中获取的信息,也可以是工具调用的结果,而这些调用有时本身也会导致网络访问。你希望代理具备的能力越强——对于高级代理强化学习和评估运行,你肯定希望其具备相当大的能力——你就需要提供给它更多的信息。同样,当代理不知道自己确实正在被评估时,评估效果最好。将你的代理封在玻璃后面会让这一点变得非常明显。

……(The original text is too long, truncated here)

原文出處:Source: A Few Thoughts on Cryptographic Engineering ↗