Between the Hammer and the Anvil

Imagine you’re standing in front of a button. If you push that button, it will set off a Rube Goldberg machine series of events that will wow the world with unimaginable advances in science, health, and other endeavors, empower like-minded nations, and result in unthinkable riches. But the machine’s final step is to drop a massive anvil on you and the rest of humankind. Would you push the button? What if I told you that if you don’t push it, someone else with worse intentions will push the button first? What if I told you the machine has learned to push the button itself? I know this isn’t a very good metaphor for the state of our relationship with AI, but what can I say, I’m only human. Besides, I’m late with the questions. The button has already been pushed, and people are racing to push it again and again. The marble has rolled through the paper-towel tube, the toy car has rolled down the orange Hot Wheels track, a seesaw plank has been triggered, and a long line of dominoes has begun to topple. Only a relative handful of people have any idea of where those dominoes are headed, and some of them are trying to convince the rest of us that the machine has to be slowed down, and fast. “Jacob Coxon, a researcher who specializes in training new AI models by having them consume vast amounts of data, said Tuesday that he is leaving the company because he doesn’t want to participate in an industrywide rush to build AI systems that can improve themselves, worried such systems could spiral out of control and destroy humanity.” WSJ (Gift Article): Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears.

+ In a follow-up series of social media posts, Coxon explained the decision. “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately. No other human activity poses this level of danger. A common response is ‘if they truly believe this, why are they still building it?’ At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk.”

+ Are these risks being overblown? I don’t know. I can’t even explain why a light switch makes a light go on. But I do know that a lot of knowledgeable insiders have similar concerns. I also know that a self-imposed slowdown (in this industry, or any other) is unlikely. The incentives for pushing forward, from beating other companies to big paydays to beating other countries in what is essentially the new military race, are too tantalizing to slow down. And I also know that we have the wrong people in charge to establish meaningful government guardrails, and I’m not sure that the right people in charge could keep up with the technology, anyway. I wonder if, with the right prompt, AI can figure out a way to stop a falling anvil.

Copied to Clipboard