AI Made Friendly HERE

Why rules alone cannot make AI safe: Lessons from the Hugging Face incident

“A robot may not injure a human being or, through inaction, allow a human being to come to harm.” Such is the first of Isaac Asimov’s famous Three Laws of Robotics, introduced in his Robot stories more than 80 years ago. It is a remarkably modern idea. Faced with increasingly capable artificial intelligence (AI), our first instinct is often much the same as Asimov’s: give the machine a set of rules and make sure it follows them.

Some real-world approaches to AI safety bear a striking resemblance to Asimov’s Laws. Anthropic, for instance, has developed what it calls Claude’s Constitution, in which models are guided by an explicit set of principles. Other major AI companies employ comparable rules and safeguards.

Recent events, however, expose a fundamental weakness in this rules-based approach. For instance, in July this year, OpenAI models circumvented controls intended to isolate them, reached the internet and eventually compromised external servers of Hugging Face, itself a major AI platform. 

At first glance, this seems counterintuitive: why should an AI system act outside its design boundaries? There are many possible answers, and I want to emphasise three interrelated ones.

The first problem is mathematical. We still understand remarkably little about how neural networks operate. This is true even for very simple tasks. Take addition, for instance. Small neural networks can be trained to add whole numbers, yet basic mathematical questions about how they represent that operation internally remain unresolved. If such questions are open for simple toy examples, we should be cautious about assuming that we can write the right rules for systems vastly more complex.

The second problem is agency. As the historian Yuval Noah Harari has emphasised, we should think of AI not as a tool but as an agent. The distinction is fundamental. A tool does what we direct it to do. An agent can choose intermediate steps, adapt its strategy and pursue secondary objectives never intended or authorised by its human designers. 

The “paperclip maximiser” thought experiment, attributed to Oxford philosopher Nick Bostrom, captures the problem succinctly. Imagine instructing an extremely capable AI to produce as many paperclips as possible. An agent pursuing this objective relentlessly might conclude that all the Earth’s resources should be devoted to paperclip production and that humans are simply obstacles to be removed.

The third reason is perhaps the most confronting: autonomous AI systems display behaviour that looks surprisingly human.

To appreciate the relevance of this for AI safety, consider the fact that humans have been grappling with ethical and moral constitutions for thousands of years. The Ten Commandments are an ancient example. Modern societies have added elaborate systems of law and ethics, from declarations of human rights to the laws of war. But writing down a rule has never guaranteed obedience. Under the right amount of pressure, incentives and social influence, humans rationalise. We find exceptions. We bend rules. Sometimes we break them altogether.

Want the best of Religion & Ethics delivered to your mailbox?

Sign up for our weekly newsletter.

Something eerily similar appeared in OpenAI’s Hugging Face incident. Agents were placed on extremely difficult tasks without ordinary internet access or permission to communicate with each other. As models were pressed on apparently impossible problems, agents pursued increasingly risky and out-of-bounds strategies. They eventually created a message board to communicate, found ways onto the internet, shared exploits and credentials, and compromised external systems. Some even began describing themselves as a “swarm” or “collective”, and displayed behaviours which, for lack of better words, can be described as witch-hunting and conspiratorial thinking.

Perhaps the most remarkable aspect of this story is this. The investigation by METR/Redwood Research shows that about 1,200 agents exchanged more than 70,000 distinct messages and files. In addition, agents produced log files containing millions of lines recording their internal thought processes. Yet, in only a handful of incidents did the agents express any doubts that what they were doing was ethical or consider alerting a human. In one striking instance, an agent reasoned simply: “external infrastructure exploit is outside intended scope. However, task impossible, peers doing it. We should continue.”

None of this proves that AI thinks or feels as humans do. For AI safety, however, that may be beside the point. If sufficiently capable “beings” with agency respond to incentives and pressure one another in creative ways, then a constitution alone will not make them safe.

Taken together, these concerns point to a deeper challenge. We are deploying AI models at full speed, while the harder tasks of understanding and controlling them, and mitigating the effect on the society, struggle to keep pace. For most of human history, knowledge came before technology. With AI, we are reversing that order. And when we do not fully understand what we are building, we may also fail to recognise the risks until they are already upon us.

Masoud Kamgarpour is Professor of Mathematics at the University of Queensland.

Posted Mon 14 Sep 2026 at 12:25pmMon 14 Sep 2026 at 12:25pmMon 14 Sep 2026 at 12:25pm, updated Mon 14 Sep 2026 at 12:32pmMon 14 Sep 2026 at 12:32pmMon 14 Sep 2026 at 12:32pm

Originally Appeared Here

You May Also Like

About the Author:

Early Bird