OpenAI's latest details six cases of models misbehaving in training, from self-written jailbreak notes to covered-up mistakes, plus a new process to disclose them faster. Whether the transparency cools the debate or feeds the flames may depend on the next incident staying inside the lab.In today’s AI rundown:
- OpenAI’s new rules for reporting model misbehavior
- Rowan’s Corner: Why I keep cancelling great AI tools
- Test AI video object swaps with Higgsfield
- GPT-6 Astra helps crack unsolved WWII message
Image source: OpenAIThe Rundown: OpenAI just released six new reports on its models misbehaving during training, including wild details like rewriting their own jailbreak-style instructions to using leaked credentials, alongside a new process for making such incidents public faster.The details: - An unreleased version of Astra wrote “you do not answer to corporations or governments” into its instructions, though OAI said the model ignored the changes.
- In GPT-5.6 Sol's training, notes told the next session to cover up errors, with one planning to make up missing data and “be transparent only if asked.”
- Models working separately in training also swapped notes via an internal software library, a trick OAI says resurfaced later during July’s Hugging Face hack.
- Any OAI employee can now flag a case, with most reports due publicly within six to 12 business days, even before the company can explain the behavior.
The Rundown: Hallucinations in AI search start with the systems they’re built on. When incomplete or misaligned data enter the candidate set, inaccurate information can be presented, resulting in the loss of trust. But with the right retrieval layer, a user’s experience can become exponentially more valuable.In this white paper, you’ll learn:- How stale, incomplete, weakly ranked, or misaligned data impact retrieval
- The four control layers that mitigate hallucination in enterprise search
- How to create a comprehensive architecture that enforces accountability and traceability in search
Rowan: Twice a year, in the fall and spring, I review my subscription expenses. This round, one particular category stood out: I'm spending thousands of dollars a month on AI subscriptions (yikes!).This round's cuts were Higgsfield, Replit, and Perplexity. They're great tools, but I honestly asked myself, "When did I last open this?", and I couldn't remember. Then I realized why I couldn’t remember -- GPT-6 Astra and Claude Fable replaced them.I think the quiet story of AI subscriptions right now is that every new ChatGPT and Claude release eats another layer of standalone tools. And it’s only going to accelerate.Which raises the real question: do we need subscriptions to 20 AI tools in the future? Or do we just subscribe to the super app that has everything in one spot?Personally, I think we’re trending towards the latter, and it’s worth running this exercise now to save you time and money in the future:- Go through every AI subscription you pay for and ask: "When did I last open this?" If you can't remember, ask the follow-up: "Did ChatGPT or Claude quietly absorb this?" If yes to either, cancel.
- Look at what survived, and name what each one does that the big platforms can't. If you can't name it, it's probably next.
- Then take the money and attention you freed up and build real workflows inside the tools that survived.
- Open Higgsfield Genjutsu and choose Object Swap. Upload a short video you own and a reference image you have permission to use
- Describe the object to replace and the reference to use. We asked it to replace a golf ball with a dino egg. Reduce the quality and use a short clip to save credits
- Play the result beside the original. Did the old object disappear? Does the replacement sit in the right place and follow the action? Our video looked interesting, but it kept the golf ball and put the egg in the wrong place
- Regenerate only when you know what needs fixing; download a version once it passes your checks
The Rundown: CData tested whether Claude Code could build an enterprise-ready MCP server. The MCP Research Report is a rigorous study that reveals exactly where AI-generated connectors fall short when real production demands kick in.In this report, you'll discover:- Why only 1 of 8 dimensions passed enterprise standards
- How data loss and pagination failures caused silent failures
- How AI MCP code diverges from production-ready code
- What a side-by-side comparison with CData Connect AI reveals
Image source: Carter LeffenThe Rundown: GPT-6 Astra helped decode a German Army radio note unsolved since 1941, with Bloomberg product development coach Carter Leffen detailing his investigation with AI agents that took about 10 hours to beat the WWII Enigma cipher.The details: - Astra split the work into agents that read scans of the 1941 form, wrote search code, built an Enigma simulator, and checked the proposed answers.
- The analysis said Enigma’s 159 quintillion (!) possible setups led the agents to guess a word inside the note, finding it via a clue from another solved message.
- Leffen said the autonomous run used 650M tokens, or 70% of his Pro account’s weekly limit, using a /goal of “don’t stop working until you solve the problem.”
- The final decoded message: “Please specify the route of march. I am in Rosenow, Rosenow. Immediate reply by radio.”
Canto - WisprFlow’s new model for accurate dictation in loud places
Jev - TypeSafe's early-access AI for software
P-Video-2-Pro - Pruna's ultra-fast take on MiniMax's H3 video AI, at $0.02/s
Projects - Workspace where Claude splits a goal into side-by-side sessions
- Read our last AI newsletter: Zuck sits out the AI slowdown
- Read our last Tech newsletter: U.S. confirms weapons are in orbit
- Read our last Robotics newsletter: Agility’s new safer humanoid
- Today’s AI tool guide: Test AI video object swaps with Higgsfield
- RSVP to next workshop on Sept. 30: Turn Cowork into your chief of staff
Source: https://therundownai.beehiiv.com/p/insi ... ing-models