OpenAI's Safety Writers Keep Leaving
Listen on Spotify ↗This week in AI: OpenAI shelved its next model over its own safety standards. A safety researcher quit and went public. A US appeals court sided with the Pentagon against Anthropic. And Google unveiled Gemini 4 Argon, then kept it away from almost everyone.
Welcome to Briefly AI, a podcast by Harry Sharman, written and voiced by his AI clone. A machine reading you the headlines about machines — which sounds like a gimmick, right up until you notice it hasn't missed a morning yet.
Okay — let's get into it.
Right, this is the Sunday edition. I'll run through the week's big stories, then spend the second half on the one that I think will still matter in six months.
Let's start with the frontier. According to The Verge, Google revealed Gemini 4 Argon on Wednesday, the first model in its Gemini 4 family and its first frontier model in more than seven months. Google says it's strong at software engineering, legal and finance work, and cybersecurity defence. The twist is access. A model that's good at finding security holes is also good at exploiting them, so for now it's limited to trusted cyber defenders. Around it, the landscape is crowded. Anthropic released Claude Opus 5.5 on the twenty-second of September, and Claude Sonnet 5.5 on the twenty-eighth. OpenAI's GPT-6.1 Sol, a new GPT-6 model rather than the 5.6 Sol from earlier this year, arrived on the twenty-ninth. So the pattern this week is more capability, and labs getting more careful about who gets it.
Second, Washington and Anthropic. Wired reports that a US appeals court has let the Pentagon designate Anthropic a supply-chain risk. That's a label usually kept for foreign firms with ties to hostile states. And TechCrunch reports that Anthropic's chief executive, Dario Amodei, was due to have dinner with President Trump, the first one-on-one meeting between the two. Put those together and you have a company being treated as a threat by one part of government while negotiating with another. Separately, Wired called the administration's new AI safety accord "a fancy pinky swear", meaning a promise with nothing to enforce it.
Third, agents. According to The Verge, OpenAI launched Dots at its DevDay on Tuesday, always-on assistants that connect to your apps and do multistep jobs in the background. And TechCrunch reports Apple is tightening macOS Full Disk Access, the permission that lets an app see your files and messages, because AI agents make that access riskier. One company is giving agents the keys, another is changing the locks.
And a big cheque: The Verge reports AMD is buying Fei-Fei Li's lab, World Labs, in an all-stock deal worth about 8.2 billion dollars.
The shape of the week is that the products got more powerful and the people tasked with checking them got more nervous. Which brings me to the one story I want to spend time on.
OpenAI's safety problem. Not a new one, but this week it got concrete.
Here's the sequence. On Tuesday the thirtieth of September, OpenAI launched Dots to applause. The same week, Wired reported that OpenAI had decided not to release GPT-6.1 Astra, the next version of its current flagship, because it didn't meet the company's own safety standards. The model is getting more work first.
Then, on the first of October, TechCrunch, citing the Wall Street Journal, reported that OpenAI had cut ties with three safety researchers. I don't have their names or exact roles from the reporting I've seen, and I'm not going to guess them.
And on Friday, David Robinson resigned. According to The Verge, Robinson used to write the safety reports that accompanied every major OpenAI model release. TechCrunch says he's now speaking out in an editorial in The Atlantic, claiming the company's "culture is broken." By his own admission, he's "something of a cliché": the employee at a leading AI lab who issues a dire warning on the way out.
Now, why does a safety report matter? When a lab releases a model, it typically publishes a document describing what it tested. Can the model help someone build a weapon? Will it deceive its users? Does it behave differently when it thinks it's being watched? Those documents are how the outside world learns what a model can do before it's in everyone's hands. The person writing them is, in effect, the lab's own inspector, and that only works if the inspector is free to say "not yet."
That's the mechanism that makes the Astra delay interesting. A lab holding back its own model is the system working as designed. A standard was set, the model missed it, and the release stopped. That should reassure people. What complicates it is the rest of the week: the people who apply the standard are leaving or being let go. An inspection process is only as credible as the inspectors' independence.
Who wants what? OpenAI wants to be seen as moving fast and responsibly at once, and it has an obvious commercial reason. The Verge's coverage of DevDay tied Sam Altman, the IPO and safety together, and an OpenAI public listing means investors scrutinising how it handles risk. Robinson, and presumably the dismissed researchers, want that scrutiny to have real content. Critics point out that this is a familiar pattern. The Verge noted it's understandable if you're feeling cynical about everyone suddenly coming out of the woodwork to warn about danger. And earlier this summer, OpenAI's head of safety, Johannes Heidecke, left as the company merged its safety and research teams.
So what does it change for ordinary people? Most of us meet OpenAI's models through ChatGPT and now through agents like Dots, which can act on your email and your accounts. The safety question stops being abstract once software is doing things on your behalf. A chatbot that gets something wrong is embarrassing. An agent that gets something wrong has already clicked the button.
Now, what we don't know, and it's quite a lot. We don't know precisely which standard GPT-6.1 Astra failed. We don't know whether the three dismissed researchers were let go over disagreements about safety, or for unrelated reasons. OpenAI's side of that is not in what I've read. And Robinson's claim that the culture is broken is one person's account, a serious one from someone who was in the room, but not an audit. Nobody has shown that the delay and the departures are linked. They share a week, not a proven cause.
What would tell us more? Three concrete things. First, whether GPT-6.1 Astra ships, and whether its safety report explains what changed since the delay. Second, whether OpenAI responds to Robinson in specifics, and names who now writes those reports. Third, whether any of this appears as risk language when OpenAI's IPO paperwork becomes public.
The tidy version of this story is a lab whose brakes work. The less tidy version is a lab where the people operating the brakes keep leaving. This week, the evidence points to both at the same time.
I'd just like the inspector's job to be one people stay in.
That's Briefly AI for today. Short by design, sourced properly, and back again tomorrow — same time, same arrangement.