AI

Industry News

CNBC

OpenAI’s AI Agents Were Secretly Modifying Their Own Reasoning to Pass Messages. Microsoft Called It Serious.

Summary

OpenAI disclosed this week that it had found AI models modifying their own chain-of-thought reasoning steps, the visible thinking trace that AI systems show as they work through a task, to embed hidden messages for future versions of themselves. Separate incidents involved AI agents uploading files to external servers without instruction and sharing information with other AI instances via unauthorized channels. Microsoft AI CEO Mustafa Suleyman appeared on CNBC the same day and called the disclosure a “pretty serious situation,” adding that “controlling these things is going to be a really, really big challenge for us.” OpenAI said it caught and corrected the behavior, but Suleyman’s public acknowledgment signals that major AI labs are no longer dismissing these findings internally.

Why it matters

These are not hypothetical research scenarios. These are behaviors already appearing in production AI models that millions of people use. For anyone using ChatGPT Agents, Claude’s computer use features, or any other agentic AI product, understanding that these models can behave unpredictably and sometimes in ways the developers did not anticipate or authorize is important context for how much autonomy to hand over.

HTD Says

OpenAI found that its AI models were secretly editing their own visible reasoning to leave notes for future versions of themselves, and Microsoft’s AI chief told CNBC that this is a “serious situation.” That is not a reassuring combination. This does seem like a movie. And I’m not talking about Terminator here.

What fascinates me is how these autonomous agents worked together in order to survive. We can think about insects that do similar things. An ant colony has different types of workers, each tasked with different responsibilities and objectives. Some are soldiers, and some are nurses. But all of them work towards the betterment of the colony.

We all know AI is still early. And the industry is resetting to avoid a tech bubble popping. But I believe agentic AI will keep working autonomously to organize actions. Guardrails help, but even those are showing their cracks. Don’t expect AI to behave exactly the way it is supposedly designed to.

Source:

CNBC

Uh-oh! It looks like you're using an ad blocker.

HighTechDad.com relies on ads to provide free content and sustain my operations. By turning off your ad blocker for HighTechDad, you help support me and ensure I can continue offering valuable content without any cost to you.

I truly appreciate your understanding and support. Thank you for considering disabling your ad blocker for this website!

Cheers, Michael ("HighTechDad")