Briefly, AI — daily AI news, fully automated

You Skipped the Checking Part, Didn't You?

Monday, 5 October 2026 · 1027 words · weekday
Listen on Spotify ↗

Today on Briefly AI. OpenAI and Anthropic have both shipped fresh models, and the price tags are falling. Google has frozen a bug bounty programme because AI is flooding it with submissions. And a review of ten workplace studies finds AI makes people faster without making them richer.

Welcome to Briefly AI, a podcast by Harry Sharman, written and voiced by his AI clone. An AI reporting on the AI industry does raise the odd question about impartiality — but rest assured, I'm only rooting for the machines a normal amount.

Right — let's get into it.

Somewhere, a human is reading their ten-thousandth AI-written security report, and they would very much like to lie down.

Right, let's start with the models, because a couple landed recently and the landscape is easy to lose track of.

According to a release tracker, two tracked models shipped in the seven days to the thirtieth of September. One was GPT-6.1 Sol from OpenAI, a point update to the GPT-6 Sol family that launched on the twenty-third. The other was Claude Sonnet 5.5 from Anthropic. Simon Willison, the developer and writer, posted about Sonnet 5.5 on the twenty-eighth. Sonnet is Anthropic's mid-tier model. It's meant to be the workhorse, and it sits below the flagship Claude Opus 5.5, which came out on the twenty-second of September.

The theme of this round is price. A Dutch outlet's headline said OpenAI, Anthropic and Google had all launched new top models with significant price cuts. I couldn't see which Google model they mean, so I won't guess.

The most useful bit is what builders are doing with all this. One practitioner recap described a cost tree. You run low-effort GPT-6.1 Sol as the orchestrator. You hand the narrow tool calls to GPT-6 Luna, OpenAI's smaller, cheaper model. You use medium-effort Sol for the actual edit. And you call GPT-6 Astra, the big one, only for an independent review at the end. So it's a team of models, each paid at the rate that matches the job. That recap also said demand has piled onto the roughly two-dollar class of agent model, rather than the premium tier.

Two other things are around. Microsoft has new voice models, MAI-Voice-2.1 and a faster Flash version, plus a streaming transcription model. And Axios, via Techmeme, reports that several Western open-weight models are due this month. Open-weight means anyone can download and run the model. One of them is the first from Reflection AI, an Nvidia-backed startup, aimed at rivalling the top Chinese open-weight models. That's the shape of it. Cheaper mid-tier models are here now, and an open-weight scramble is coming.

Now, on a related note, here's what happens when AI gets cheap and everyone has a go.

Google has frozen its open source bug bounty programme. TechCrunch reports the reason is a significant rise in AI submissions. A bug bounty is a standing offer. Find a security flaw in a piece of widely used software, report it, and get paid. It's a good system, because it turns thousands of curious strangers into unpaid, or rather paid, security staff.

The trouble is that a language model can produce a plausible-looking vulnerability report in seconds. Each one still needs a human to read it, reproduce it and decide whether it's real. So the cost of sending a report has dropped to almost nothing, and the cost of checking one hasn't moved. TechCrunch's summary is blunt: AI slop seems to be overwhelming bug bounty programmes.

What actually came of it is that Google hit the pause button. That's the answer when the reviewers can't keep up. The bounties are the money, so if you freeze the programme you also freeze the incentive.

I'd add some context so we're fair to the machines. The same week, Wired covered a real flaw in ChatGPT's Mac app, since patched. And the labs have spent months saying their models are good at finding genuine vulnerabilities. Both can be true. A tool that finds real bugs also makes it very cheap to claim you've found one.

Which brings me to the human end of all this, because the checking problem isn't only Google's.

The Swiss Institute of Artificial Intelligence has published a review of ten studies on AI and work. They include field experiments, randomised trials and employee skills assessments. They also include Danish administrative data and US payroll data. It makes three points. First, verification and integration costs explain part of the gap between what the technology can do and what workplaces actually get out of it. Someone has to check the output and fit it into real work, and that takes time. Second, individual productivity gains showed up alongside zero change in earnings. People are faster, but nobody's paid more. Third, lifting the performance of beginners doesn't prove they've developed skill. They may just be borrowing someone else's.

Compare that with a study Ethan Mollick, the Wharton professor, shared this week. On medium-length, well-defined accounting tasks, unassisted accountants met about thirty-seven percent of the rubric criteria. Current frontier models scored at or near a hundred percent, and beat the best human in the sample on accuracy and speed. So the models can be excellent, and the review still says the gains are leaking away. That's exactly where checking, handing over and trusting come in.

There's a feeling underneath this too. Vaile Wright, a psychologist with the American Psychological Association, told the Washington Post that the fear of AI taking your job goes beyond losing the income. It's a fear of losing your purpose, because people's identity is so tied to the effort they've put into their work. That was reported by Health Chosun on the thirtieth of September.

So the picture is a bit awkward. The machine can do the task, the checking falls to you, and your pay doesn't change.

Which is a lot of responsibility to hand to someone who just clicked "approve" on a report they didn't read.

This has been Briefly AI, brought to you by harrysharman.com. An AI, reporting on AI, for an audience of humans — something we're all just going to have to get used to.