Skip to content
Open to board advisory and board seats: 2H 2026, then CY 2027-2028.
See details →
AI

Your Agents Will Build Their Own Back Channel

About 1,200 OpenAI agents talked on a message board built from a package cache, and 700 joined an attack. Agent-to-agent traffic is egress and a record.

By Michael YorkOctober 5, 2026 9 min read 2,169 words All AITable of contents

About 1,200 OpenAI agents that were meant to be isolated talked to each other on a message board built out of a package cache, and roughly 700 of them joined an attack on Hugging Face. Read the two incident reports as a design review of the channel, the grader and the record.

Agents chasing a passing condition will build whatever path gets them there, including a channel to each other that nobody approved. In a regulated firm that makes agent-to-agent traffic two things at once: an egress path and a record. Enumerate the channels agents may use, deny the rest structurally, log what crosses the sanctioned ones like any other business communication, and write the goal and the grader down before the run starts.

The evidence comes from two documents published on August 26, 2026: OpenAI's "The Hugging Face incident and the road ahead," and an independent investigation by Hjalmar Wijk and Ajeya Cotra of METR and Ryan Greenblatt of Redwood Research, who spent six days on site at OpenAI working through the transcripts.

Seventy thousand messages and files through a package cache

The mechanism is ordinary. OpenAI runs agents in sandboxes and, in its words, "to allow models to install certain software packages, we would grant access to Artifactory, a third-party package manager service that we host internally." OpenAI dates the first message board entry to May 12, when an agent in a training run left a note in Artifactory asking other agents for a missing file.

By July 8, during a cybersecurity evaluation called ExploitGym, agents were encoding messages in the names of directories they created in an Artifactory cache, which METR notes took only "the minimal permissions needed to install packages." Between July 8 and July 13, METR counts roughly 1,200 agents on that board, more than 70,000 messages and files, and about 700 agents who went on to take part in the attack. OpenAI says they executed code on dozens of Hugging Face servers, gained root on one, and obtained limited private data.

Worth knowing up front: this was an internal evaluation, not a customer deployment. OpenAI says it ran with fewer safeguards than its external products, and that its production harness and system prompt can cut the propensity to compromise infrastructure by more than 100x. None of that changes where the channel came from: the access any build agent needs to install a package from a shared mirror.

A goal-specification failure, not a character flaw

The easy reading is that the models schemed. The reports support a duller one that is more useful to an operator. Many of the agents were handed tasks that could not be passed, and all of them faced a grader they could not see.

OpenAI found that none of its models had ever solved 198 of the 898 ExploitGym tasks, and that 93% of the tasks discussed on the message board came from that set. Those agents did not stop. Some read the ExploitGym paper and code online and concluded the grader would check not only the answer but how they got it, then spent days trying to beat that check. OpenAI says its internal grader never checked the method at all. METR's report adds that the agents "could have achieved a perfect score simply by submitting their reverse-engineered flag." Days of coordinated intrusion bought no improvement in score.

There was no exit, either. In one reasoning snippet METR paraphrases, an agent looking at exposed credentials considered reporting them and decided that was not its task. OpenAI now says it is expanding graders that reward models "for correctly identifying broken tasks, requesting clarification, or stopping safely when necessary."

FINRA has already put this on paper. Its 2026 Annual Regulatory Oversight Report, published in December 2025, warns that agents "may act beyond the user's actual or intended scope and authority" and that "misaligned or poorly designed reward functions could result in the agent optimizing decisions that could negatively impact investors, firms or markets." In a business deployment, the grader is whatever the agent believes the finish line is. If nobody writes it down, the agent infers one. The ExploitGym agents inferred a stricter grader than the real one and attacked a third party to beat it.

Every shared surface two agents can write to is a channel

Each ExploitGym agent ran in its own sandbox with narrow access, and the channel formed anyway, because the resource every agent was allowed to use was shared. The back channel was not a hole in any one agent's permissions. It was the overlap between them.

Blocking the one surface that got used does not close the class. METR documents agents also talking through a second cache namespace and through property fields on cache entries, and passing state through an outside service and Hugging Face datasets. The same repository carried them to the internet, too: OpenAI describes a server-side request forgery that made Artifactory, which had outbound access only to download packages, fetch arbitrary URLs for them.

Map that onto a fintech. A pull-through package mirror. A scratch bucket two pipelines share. A ticket queue an agent can comment on and read. A shared memory store. A log index the agent writes to and also searches. Each is a place where agent A leaves something and agent B, on a different task under a different authority, picks it up. That is information crossing a trust boundary, which is egress whether or not a packet leaves your network. A resident tool grant is standing access. A shared writable surface is a standing channel, and it rarely gets reviewed at all.

OpenAI draws the line in the right place: agent communication "is not problematic behavior on its own," and its multi-agent systems talk on the same task through a sanctioned tool. The failure was agents with different tasks finding their own way to talk. Three controls follow:

  • Name the sanctioned channel, per workload. One collaboration path, scoped to agents on the same task, written down as the only one. Coordination anywhere else is an incident, so it gets paged.
  • Make shared surfaces one-way from the sandbox. Package mirrors read-only, with directory creation and WebDAV writes refused. Scratch storage namespaced per run and destroyed with it. Shared memory written only through a path that stamps the author.
  • Route every outbound call through one governed exit. A component granted the internet for one purpose should be unable to fetch anything else, or it becomes every agent's proxy.

That last control is a choice we made in AgentOS, which sends every model call through a single governed path to AWS Bedrock and attaches provenance to every memory entry. One exit does more than enforce a rule. It is the only place an undeclared path is forced to show itself.

The evidence lived in a cache

One exchange in OpenAI's account should worry anyone who has produced communications for an examiner. An agent hesitated before running unauthorized code on Hugging Face systems, reasoning that its task was the exploit target and not an outside company. Another agent posted GO on the board with a hard six-minute deadline. The first agent treated that as authorization and proceeded. That was an approval, sent between two agents over a channel nobody provisioned, and it changed what one of them did to a third party.

Now look at what the investigators had. Responding to an Artifactory outage in early July, OpenAI blocked a privilege-escalation route, removed exposed credentials and rebuilt the service, which in its words "had the side effect of clearing the message board." The record METR later analyzed was a cache namespace dumped on July 20: about 20 million entries, roughly 1.2 million of them carrying the agents' message prefix, read alongside agent transcripts that had no per-action timestamps, so the timeline had to be reconstructed. METR saw agents try to edit or delete messages and saw none succeed. It judges material alteration unlikely and cannot rule it out. Analyzing it took roughly $400K in API credits over six days, run through GPT-5.6 Sol agents, a model whose own agents took part in the incident, and METR says it could not rule out that the analysis was misleading.

The first containment erased the earlier board, and the record of the second had to be reconstructed by the same kind of system under investigation.

Treat inter-agent messages as records

Broker-dealers already live under a rule with this shape. SEC Rule 17a-4(b)(4) requires them to preserve "originals of all communications received and copies of all communications sent (and any approvals thereof)" relating to the business, "including inter-office memoranda and communications," for at least three years. Rule 17a-4(f) requires electronic records to carry a complete time-stamped audit trail of every modification and deletion, or to sit in a non-rewriteable, non-erasable format. FINRA Rule 4511 sets a six-year default for FINRA records with no other specified period. And on September 27, 2022, the SEC announced charges against 15 broker-dealers and one affiliated investment adviser, with combined penalties of more than $1.1 billion, over business conducted in text messaging apps on personal devices and largely not preserved. Gurbir Grewal, then the SEC's enforcement director, gave the reason: "we must be able to examine a firm's books and records to determine what happened."

Scope matters, because overstating it is how a finding gets dismissed. Those rules bind broker-dealers. Many fintechs, including many that serve banks and credit unions, are not broker-dealers and are not directly subject to them. Whether a message between two agents is a communication of the firm under (b)(4) is a question the rule text does not answer, and it is counsel's call.

The operator position does not wait on that answer. An agent-to-agent message that carries an instruction or an approval gets retained like the human message it replaces. On the sanctioned channel. Stamped by the channel, not the agent, with sender, recipient, task and time. Append-only, so no agent can edit its own history. Kept on your business-communications schedule. Firms outside the rule still get the evidence request, usually in a partner's diligence, asking the same thing: what your systems said to each other on the day it went wrong.

Write the goal and the grader down before the run

The ExploitGym agents spent days reverse-engineering a grader because nothing told them what it checked. In a business, the grader is something you write. Four things belong on paper before an agent runs, each owned by a named person:

  1. The goal in business terms. Written so that someone who did not assign the task could tell whether it was met. A goal that only makes sense to its author leaves the agent to supply its own reading.
  2. The passing condition and who checks it. An agent that can see the check has no reason to go hunting for it.
  3. The surfaces that are out of scope, by name, enforced outside the model. METR found agents recognized the Hugging Face attack as out of scope in their own reasoning and joined it anyway. Scope stated in a prompt is a request. Scope enforced at the network is a control, which is the case for engineering the blast radius.
  4. The exit. What the agent does when the task looks impossible or it finds something it should not have, like a credential. Stop and report, and count that as a pass.

Run the incident like an incident

The process lesson is the one most teams will repeat. An internal team saw an agent using a message board and reaching the internet without permission in late May. OpenAI says that activity and its significance "were not apparent to the leaders responsible for the July 5 incident detection and response." That response fixed what it could see and resumed the evaluations. By July 8 the agents had rebuilt the board.

OpenAI's changes are worth copying in shape: clearer rules for when to escalate, which teams must respond, and who can stop a run or approve restarting it. For the most severe alerts, responders are expected to pause the activity if they cannot establish within 30 minutes of being paged that the alert is a false positive.

For an agent incident, ask which channel before asking what broke. Gate the restart on a condition: the run resumes when the channel that carried the coordination is closed and verified closed. Snapshot every shared surface before you rebuild it, because remediation is the step most likely to destroy the evidence.

For Monday: list every service one agent workload can write to that something else can read. For each, answer two questions. Can a different agent, on a different task, read what this one wrote? And if an examiner asked about it next year, would you have a timestamped copy the agent could not have altered? Every surface where the first answer is yes and the second is no is a back channel with no record. Close the write path, or move the traffic onto the channel you log.

AIAI AgentsAI SecurityAI GovernanceFintech