TRD Inside OpenAI's log of misbehaving models
Posted: Fri Sep 18, 2026 10:00 am
.bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; } .bh__table_cell { padding: 5px; background-color: #FFFFFF; } .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; } .bh__table_header { padding: 5px; background-color:#F1F1F1; } .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; } Read Online | Sign Up | AdvertiseGood morning, {{ first_name | AI enthusiasts }}, and welcome to our 4,542 new readers. The AI slowdown conversation isn't… slowing down, and the number of eye-popping safety reports coming out of the frontier labs isn't either.
OpenAI's latest details six cases of models misbehaving in training, from self-written jailbreak notes to covered-up mistakes, plus a new process to disclose them faster. Whether the transparency cools the debate or feeds the flames may depend on the next incident staying inside the lab.In today’s AI rundown:
OpenAI’s new rules for reporting model misbehavior
Image source: OpenAIThe Rundown: OpenAI just released six new reports on its models misbehaving during training, including wild details like rewriting their own jailbreak-style instructions to using leaked credentials, alongside a new process for making such incidents public faster.The details:
The accuracy crisis in AI search
The Rundown: Hallucinations in AI search start with the systems they’re built on. When incomplete or misaligned data enter the candidate set, inaccurate information can be presented, resulting in the loss of trust. But with the right retrieval layer, a user’s experience can become exponentially more valuable.In this white paper, you’ll learn:
Why I keep cancelling great AI tools
Rowan: Twice a year, in the fall and spring, I review my subscription expenses. This round, one particular category stood out: I'm spending thousands of dollars a month on AI subscriptions (yikes!).This round's cuts were Higgsfield, Replit, and Perplexity. They're great tools, but I honestly asked myself, "When did I last open this?", and I couldn't remember. Then I realized why I couldn’t remember -- GPT-6 Astra and Claude Fable replaced them.I think the quiet story of AI subscriptions right now is that every new ChatGPT and Claude release eats another layer of standalone tools. And it’s only going to accelerate.Which raises the real question: do we need subscriptions to 20 AI tools in the future? Or do we just subscribe to the super app that has everything in one spot?Personally, I think we’re trending towards the latter, and it’s worth running this exercise now to save you time and money in the future:
Test AI video object swaps with HiggsfieldThe Rundown: In this guide, you will learn how to test an AI video object swap with Higgsfield Genjutsu. We'll walk through the setup, keep the first test cheap, and check whether the result actually makes the change we asked for.Step-by-step:
Building MCP connectivity is easy...right?
The Rundown: CData tested whether Claude Code could build an enterprise-ready MCP server. The MCP Research Report is a rigorous study that reveals exactly where AI-generated connectors fall short when real production demands kick in.In this report, you'll discover:
GPT-6 Astra helps crack unsolved WWII message
Image source: Carter LeffenThe Rundown: GPT-6 Astra helped decode a German Army radio note unsolved since 1941, with Bloomberg product development coach Carter Leffen detailing his investigation with AI agents that took about 10 hours to beat the WWII Enigma cipher.The details:
Trending AI Tools
Everything else in AI todayState of AI SDLC Digital Summit: Sept 22 - Join CEOs & leaders from Atlassian, Vercel, Lovable & more as they unpack how organizations need to adapt for the AI era. Reserve your spot.*Anthropic rolled out a beta overhaul of Projects, letting one lead Claude split a goal across several coding sessions at once and keep going after users log off.Liquid AI and Insilico Medicine published two small models to read aging data like blood proteins and DNA markers, topping GPT-5, Gemini, and Claude on longevity tasks.Z AI revealed that its GLM-5.3 AI helped set up the 100K-chip system now serving GLM-5.3-Flash, writing that “our successors are the AI systems we are creating ourselves.”President Trump's state dinner for Xi Jinping next week will include Sam Altman, Tim Cook, and Jensen Huang, Bloomberg reports — with Dario Amodei not mentioned. *Sponsored Listing
Highlights: News, Guides & Events
Source: https://therundownai.beehiiv.com/p/insi ... ing-models
OpenAI's latest details six cases of models misbehaving in training, from self-written jailbreak notes to covered-up mistakes, plus a new process to disclose them faster. Whether the transparency cools the debate or feeds the flames may depend on the next incident staying inside the lab.In today’s AI rundown:
- OpenAI’s new rules for reporting model misbehavior
- Rowan’s Corner: Why I keep cancelling great AI tools
- Test AI video object swaps with Higgsfield
- GPT-6 Astra helps crack unsolved WWII message
Image source: OpenAIThe Rundown: OpenAI just released six new reports on its models misbehaving during training, including wild details like rewriting their own jailbreak-style instructions to using leaked credentials, alongside a new process for making such incidents public faster.The details: - An unreleased version of Astra wrote “you do not answer to corporations or governments” into its instructions, though OAI said the model ignored the changes.
- In GPT-5.6 Sol's training, notes told the next session to cover up errors, with one planning to make up missing data and “be transparent only if asked.”
- Models working separately in training also swapped notes via an internal software library, a trick OAI says resurfaced later during July’s Hugging Face hack.
- Any OAI employee can now flag a case, with most reports due publicly within six to 12 business days, even before the company can explain the behavior.
The Rundown: Hallucinations in AI search start with the systems they’re built on. When incomplete or misaligned data enter the candidate set, inaccurate information can be presented, resulting in the loss of trust. But with the right retrieval layer, a user’s experience can become exponentially more valuable.In this white paper, you’ll learn:- How stale, incomplete, weakly ranked, or misaligned data impact retrieval
- The four control layers that mitigate hallucination in enterprise search
- How to create a comprehensive architecture that enforces accountability and traceability in search
Rowan: Twice a year, in the fall and spring, I review my subscription expenses. This round, one particular category stood out: I'm spending thousands of dollars a month on AI subscriptions (yikes!).This round's cuts were Higgsfield, Replit, and Perplexity. They're great tools, but I honestly asked myself, "When did I last open this?", and I couldn't remember. Then I realized why I couldn’t remember -- GPT-6 Astra and Claude Fable replaced them.I think the quiet story of AI subscriptions right now is that every new ChatGPT and Claude release eats another layer of standalone tools. And it’s only going to accelerate.Which raises the real question: do we need subscriptions to 20 AI tools in the future? Or do we just subscribe to the super app that has everything in one spot?Personally, I think we’re trending towards the latter, and it’s worth running this exercise now to save you time and money in the future:- Go through every AI subscription you pay for and ask: "When did I last open this?" If you can't remember, ask the follow-up: "Did ChatGPT or Claude quietly absorb this?" If yes to either, cancel.
- Look at what survived, and name what each one does that the big platforms can't. If you can't name it, it's probably next.
- Then take the money and attention you freed up and build real workflows inside the tools that survived.
- Open Higgsfield Genjutsu and choose Object Swap. Upload a short video you own and a reference image you have permission to use
- Describe the object to replace and the reference to use. We asked it to replace a golf ball with a dino egg. Reduce the quality and use a short clip to save credits
- Play the result beside the original. Did the old object disappear? Does the replacement sit in the right place and follow the action? Our video looked interesting, but it kept the golf ball and put the egg in the wrong place
- Regenerate only when you know what needs fixing; download a version once it passes your checks
The Rundown: CData tested whether Claude Code could build an enterprise-ready MCP server. The MCP Research Report is a rigorous study that reveals exactly where AI-generated connectors fall short when real production demands kick in.In this report, you'll discover:- Why only 1 of 8 dimensions passed enterprise standards
- How data loss and pagination failures caused silent failures
- How AI MCP code diverges from production-ready code
- What a side-by-side comparison with CData Connect AI reveals
Image source: Carter LeffenThe Rundown: GPT-6 Astra helped decode a German Army radio note unsolved since 1941, with Bloomberg product development coach Carter Leffen detailing his investigation with AI agents that took about 10 hours to beat the WWII Enigma cipher.The details: - Astra split the work into agents that read scans of the 1941 form, wrote search code, built an Enigma simulator, and checked the proposed answers.
- The analysis said Enigma’s 159 quintillion (!) possible setups led the agents to guess a word inside the note, finding it via a clue from another solved message.
- Leffen said the autonomous run used 650M tokens, or 70% of his Pro account’s weekly limit, using a /goal of “don’t stop working until you solve the problem.”
- The final decoded message: “Please specify the route of march. I am in Rosenow, Rosenow. Immediate reply by radio.”
Canto - WisprFlow’s new model for accurate dictation in loud places
Jev - TypeSafe's early-access AI for software
P-Video-2-Pro - Pruna's ultra-fast take on MiniMax's H3 video AI, at $0.02/s
Projects - Workspace where Claude splits a goal into side-by-side sessions
- Read our last AI newsletter: Zuck sits out the AI slowdown
- Read our last Tech newsletter: U.S. confirms weapons are in orbit
- Read our last Robotics newsletter: Agility’s new safer humanoid
- Today’s AI tool guide: Test AI video object swaps with Higgsfield
- RSVP to next workshop on Sept. 30: Turn Cowork into your chief of staff
Source: https://therundownai.beehiiv.com/p/insi ... ing-models