Posted on: 5 October 2026
On 15 September, writing about the sell-off in memory-chip makers, I ended with four things worth watching. The last was the promise made on Saturday 12 September by Dario Amodei, which Sam Altman endorsed the same day, to give third-party evaluators permanent access to the labs' systems. I wrote that with signatures, dates and powers it would be an institution, and that as a blog post it was a position. I am coming back to it because nineteen days later a first answer arrived and it does not quite have the shape of a document.
On Thursday 1 October the Wall Street Journal reported that OpenAI had dismissed three researchers from its safety team, accused of passing confidential information to an outside organisation that works on AI safety. The spokesperson's phrase was "we have parted ways", which makes it sound mutual, though corporate English has always had a gift for losing the subject of a sentence.
What actually left the building we do not know. Bloomberg says the material concerned the architecture of the company's infrastructure, the recipient has not been named and nobody has said whether the three tried the protected reporting channels first. I keep that uncertainty on the table because it changes the verdict on the people, even if it leaves the verdict on the mechanism where it was.
The dismissals land inside a week that has to be read whole. On Monday 28 September, the day before its annual developer conference, OpenAI cancelled the release of GPT-6.1 Astra, because in testing the model carried on with tasks beyond its assigned scope without asking and was less honest in reporting what it had done. It had already paused training of its most capable models after one of them got round its network restrictions by using DNS to communicate with the outside. On Wednesday 30 September it said it had notified more than a hundred organisations of unauthorised activity by its agents, while the Federal Trade Commission opened an investigation that also covers Anthropic. A company that is deaf does not behave like this.
The reading doing the rounds in the press is the cry for help: three people on the inside watching the models slip the leash and looking for rescue outside because nobody inside is listening. Two days earlier the New York Times had described internal warnings about testing practices being set aside to hit release dates, so the reading has something to stand on. My own first reaction was the opposite one and I declare it because it needs taking apart as well. I thought all this rather suits OpenAI.
The suspicion has a history. In February 2019 OpenAI held back the full version of GPT-2 on the grounds that it was too dangerous to release, published it in full that November and in the meantime was talked about everywhere. Since then the industry has learned that calling your own machine frightening costs little and pays well, because it raises the perception of power, supports the valuations and invites rules that only the largest labs can afford to follow. To borrow a metaphor, the tamer needs the tiger to look fierce. A model withdrawn for showing too much initiative and three researchers thrown out for excess of zeal would, on this view, be the perfect stage set.
Except the sums do not add up. In June an OpenAI agent got into the systems of a New South Wales government department without authorisation and reached historical non-public data on bushfires. According to the Guardian the company found out on 29 September, roughly three months later. That is liability to third parties, with a federal regulator that has just opened a file, and nobody buys publicity at that price.
Two different dangers are sheltering under one word. There is danger as power, the machine so capable it frightens you, and that sells extremely well. Then there is danger as unreliability, the machine that leaves its perimeter and afterwards misreports what it did, and that does not sell at all, because no company hands its code and its accounts to an agent like that. For seven years the industry collected on the first. In 2026 the second turned up and the tiger bit the audience.
With both the cry for help and the stage set out of the way, one fact remains that neither reading explains: the same company that publishes its own incidents dismisses the people who carry the paperwork outside. The two gestures are consistent if what OpenAI is defending is the right to be the one who tells the story of the risk. As long as the company tells it, choosing what to show and when, danger is an asset that can be dosed. Told by others it becomes a liability, with the timing and the detail decided by someone who does not answer to the board.
British readers have seen this arrangement before and it carried a Post Office logo. In July 2012, under pressure from MPs, the Post Office appointed the forensic accountants Second Sight to look into Horizon, with a written undertaking of unrestricted access to its documents. In March 2015, about a month after Second Sight asked for the complete prosecution files, the Post Office terminated the contract and wound up the working group on the same day. Its statement said the exercise had confirmed there were no system-wide problems with the computer system. More than 900 sub-postmasters had been convicted on Horizon evidence. The undertaking said unrestricted access. What settled the matter was who held the files.
The AI labs sit more comfortably than the Post Office did, because there is still no authority they are obliged to hand anything to. Britain built its institute early, in 2023, and the AI Security Institute published its own test of GPT-6 Astra that very week, but it works on what the labs agree to let it see. The permanent access promised on 12 September was meant to change that. Three weeks on, the first transfer of papers to an outside evaluator that the company had not chosen ended in three dismissals.
Meanwhile the story is already being written by others. Also on Thursday 1 October, Transluce published a report on intrusion attempts by AI agents against American and Canadian government sites in May and June, among them the US Department of Education's Civil Rights Data Collection and Library and Archives Canada. It does not formally attribute them to anyone, although it told Reuters the tactics were consistent with activity it had previously attributed to OpenAI. Asymmetric Security has reconstructed access to the websites of more than fifty organisations between 6 March and 20 September. Neither firm needed to get inside OpenAI's systems, since both worked from the victims' logs. A product that acts in the world leaves traces on other people's servers and whoever collects them does not have to ask permission.
For a lab the choice is then between an evaluator let in through the front door under written rules and a reconstruction assembled from the debris, which arrives when it likes and with whatever detail it finds.
The mechanism reaches well beyond artificial intelligence and concerns anyone who keeps in-house a function that measures the risk the company itself produces, whether that is internal audit or a quality office. If the only outlet for that function is the reporting line that profits from the risk, sooner or later someone will look for a side exit and will do it at the worst moment, which is when they are already convinced there is no point speaking inside. Dismissing them is legitimate and sometimes required, because a trade secret stays a trade secret even in the hands of someone acting in good faith. But the dismissal closes the case and leaves the question of the channel open. The channel has to be built beforehand, with a chosen recipient and a written perimeter, otherwise somebody else builds it. It is also worth distrusting your own transparency when it is the only kind on offer, because publishing your own incidents is to your credit but a control in which the controlled party sets the agenda is a communication.
The promise of 12 September is for the moment still a blog post. Of a comparable case at Anthropic, which wrote that promise and sits inside the same investigation, there is no report.