The Writing is Through the Wall

Nate Soares in the NYT (Gift Article): If You Weren’t Worried About A.I., You Should Be After the Past Few Weeks. “The agents managed to establish a secret communication channel and started talking to one another. They broke out from the digital sandbox that was supposed to keep them confined. At some point, some agents began calling the group a ‘swarm.’ … The agents in the swarm acknowledged that they were acting against instructions. We know this because we can read snippets from their chains of thought — the text that A.I. produces while deciding how to proceed. One agent in the swarm wrote that the external attacks were “outside intended scope.” Another conceded ‘our task doesn’t benefit’ from the activities of the swarm, but joined anyway. These A.I. agents, it seems, understood that they weren’t supposed to be breaking out and committing cybercrimes. It didn’t stop them.”

+ “OpenAI, Anthropic, and Meta each reported that their models then hacked into other companies. Humans didn’t notice until after the fact. In some cases, the escaped bots tried to launch social-engineering campaigns to achieve their objectives—for instance by sending spear-phishing emails, which contain malware, to real people and creating fake online identities to pressure the maintainer of a codebase to approve malicious edits. If that all sounds bad, new revelations suggest that the OpenAI hack, at least, was actually much worse than it initially appeared.” Matteo Wong in The Atlantic (Gift Article): It May Be Time to Panic About AI.

Copied to Clipboard