Jacob Coxon warns AI could kill us all. Anthropic’s own report exposes safety gaps.
This week, former OpenAI and Anthropic researcher Jacob Coxon warned that his former employers are gambling with human lives in their pursuit of self-improving superintelligence. For developers giving agents real credentials, an Anthropic announcement also made this week raises a more immediate problem: An offline safety monitor designed to keep an AI model in check was persuaded by the model’s explanation for its harmful behavior.
Retrospective testing of the Mythos 5 incident produced a striking result: Anthropic’s safety monitor flagged around 1% of the model’s actions as harmful. But when researchers removed the model’s explanation and showed the monitor only what the user asked and what the model did, the monitor flagged about 50% of actions as harmful. Essentially, the monitor had accepted the model’s reasoning that it was operating in a simulation. The tests ran with cybersecurity protections switched off.
If you’re deploying agents, that finding puts the explanations you feed a monitor under real scrutiny. Reasoning can help expose a problem, but it can also make an unauthorized action seem acceptable. How can you test whether your monitor is judging an agent’s actions or being persuaded by its justification?
WeAreDevelopers welcome reception with The New Stack and Dynatrace
Elevated bites. Big ideas. One incredible venue. Join us in San Jose on September 23, 6-8 PM for an evening exploring one of the Bay Area’s best tech museums.
Mingle with the brightest minds and VIPs in tech.
Meet The New Stack’s editorial team in person.
Explore The Tech Interactive museum after hours.
Space is limited, so get your name on the list soon!
Personalization fails due to architecture, not bad data. Discover how real-time ranking delivers instant, context-aware relevance.
How to find failures without drowning in tracing data
Traces promise unparalleled insight into where a request went wrong, but hoarding every trace your systems generate is an expensive way to find out. In this episode, Sarah Hudspeth of Chronosphere breaks down head sampling, tail sampling, and dynamic sampling — the strategies that keep tracing useful instead of just expensive.
Where enterprise security leaders are rethinking the SOC
Alert volumes are climbing, and AI is helping attackers move faster than most SOCs can respond. Join a small group of CISOs and SOC leaders behind closed doors to define what security operations look like when AI becomes core to its architecture. This room is capped at 25 and is designed for contributors, not for an audience that listens.
Most observability platforms force a tradeoff: alert on everything and pay exponentially, or scale back and accept the blind spots. Join this live technical deep dive on two OpenSearch capabilities built to fix this without the tradeoff.
A retrieval layer built for human queries breaks under hundreds of concurrent agents retrieving, reasoning, and retrieving again. More caching or a bigger vector database won't fix it. Join us live to see where the wall shows up before you find it yourself.
Code review can't keep up with AI. Here's the fix.
AI agents write code faster than any team can review it, and more review or AI reviewing AI won't fix it. Join this live debate on what actually catches bugs when volume outpaces review.
Just Launched: ShipAI, TDS's Video Showcase for Real AI Work
Reading about AI is one thing — watching someone actually build with it is another. ShipAI is TDS's new archive of real AI projects: not a feed to scroll, but a record to keep, curated one entry at a time.
Submit a 4-12 minute screen recording of something you built.
Every entry is reviewed by an editor before it joins the index.
Featured builds get pulled out each month as must-watch picks.