Natural pacing through human accountability
Jacob Coxon’s recent resignation from Anthropic, announced in a thread on X, has sparked the latest round of debate about the pace of AI development, government regulation, and independent nonprofits reporting on the safety of AI labs.
One underlying trend in these discussions, especially among people arguing for government intervention or cross-industry commitments, is the tendency to anthropomorphize AI. They talk about the risk as though it isn’t a risk of humans doing things they are accountable for, but of technology somehow going off and doing its own thing.
People are comfortable assigning human agency to preventing disasters and working on safety and “alignment.” But they rarely talk about human agency playing a role in bad outcomes.
When Dwarkesh Patel described the incident in which OpenAI’s agents attacked Hugging Face, he framed it as the actions of “three consecutive secret AI civilizations”. That framing obscures the role of researchers working on reinforcement learning without strong enough controls and guardrails to prevent the processes running in their sandboxes from reaching the open web and exploiting other companies’ systems.
It’s pretty unthinkable that employees at my own company, Netlify, would launch attacks against competitors. But if they did, even unintentionally, I would have a hard time just blaming our Kubernetes cluster or our cloud computing platform. I would expect serious legal ramifications.
When AI researchers or CEOs of AI companies with a deep understanding of the technology describe their fears about what could happen if they achieve recursive self-improvement, it would be naive not to take those fears seriously. It’s also disingenuous to dismiss it all as a “psy-op” or some game of 4D chess aimed at regulatory capture, when the much simpler explanation—that people are genuinely afraid of what this technology could do—seems fairly easy to validate.
But these fears and predictions keep being phrased as if the risks could materialize without any real human involvement or agency. The claims are typically not that “Anthropic will kill all humans” or “the reinforcement learning research team at OpenAI will take down the internet,” but that some ominous AI being will do it.
They also seem to assume that we will go from “all is good” to “extermination” without any intermediate safety incidents or major accidents.
That’s the part that seems least likely to me. In fact, incidents involving OpenAI’s agents hacking Hugging Face, reportedly attacking RubyGems, and taking over a German programming wiki show that real-world incidents are already happening.
Notice how all of these independent reports lead with agents as the seemingly accountable actors: “investigation of agents’ behavior,” “agents carried out an undisclosed cyber-attack,” and “AIs colluded to share answers, research their environment, and bypass sandbox restrictions.”
Why not “Investigation of OpenAI’s cyberattack on Hugging Face,” “OpenAI carried out attacks on RubyGems,” or “Details of sandbox security failures within OpenAI’s RL environments”?
Why not focus on the major human efforts involved in these attacks? RL and pretraining runs don’t just happen. They are incredibly expensive, intentional R&D efforts run by large teams of humans—so why the constant phrasing of these incidents as if they involved no human activity or agency?
But the fact that we are seeing these incidents along the way is mainly a good thing when it comes to assessing any “extermination”-level risk.
As long as we actually hold humans and companies run by humans accountable for what they do with technology, this will lead to a natural pacing. Not a pacing that requires government rules about what math or computer science we can do, or how a model needs to look and behave, but one that requires holding companies and people accountable for what they do with the technology they build and operate.
From a business perspective, software that’s both nondeterministic and randomly harmful—to the point of carrying out cyberattacks—is going to be worth a lot less than software that generally follows instructions and delivers valuable results.
This principle doesn’t just hold in the US or in democracies. In many ways, I suspect autocratic governments would be even more worried about companies wielding rogue, uncontrollable software that launches cyberattacks against arbitrary targets, including the governments themselves.
So as long as the journey from the current state of AI—which is absolutely not going to end humanity—to some mythical recursive self-improvement (RSI) takeoff is littered with accidents along the way, we can expect a natural pacing. You just can’t operate a business that keeps causing accidents.
Waymo would not have been able to expand into its current markets if there had been a constant stream of cars accidentally killing pedestrians or cyclists. As a business, it naturally had to pace its launches according to its ability to operate safely. In fact, Waymo’s largest early competitor, Cruise, imploded after launching fully autonomous rides before it was ready and having a robotaxi involved in an accident with a pedestrian.
I don’t see why digital applications of AI should be materially different, as long as we actually hold humans accountable and don’t treat training-related cyber incidents as the mythical actions of alien civilizations.