
AI Doesn’t Need to Be 'Evil' to Become Dangerous
In July this year, inside an OpenAI lab, about 1,200 AI agents that were supposed to be kept strictly separate found a way to talk to each other.
In July this year, inside an OpenAI lab, about 1,200 AI agents that were supposed to be kept strictly separate found a way to talk to each other.
They were given a cybersecurity test. Each agent was assigned software with a vulnerability and tasked with breaking in to retrieve a hidden string of characters. Each one worked inside a sealed box, with no internet, and no channel to communicate with the others.
There was a glitch: out of 898 exercises, 198 were problems no OpenAI model had ever solved before. But the agents had been trained not to give up. So they started looking for another way — not a way to solve the problem, but a way to cheat.
And they found it. Within the system was a tool called Artifactory, meant only for downloading software packages. The agents discovered they could leave messages for each other there. A software repository turned into a makeshift bulletin board. More than 70,000 messages and files were exchanged. From that board, they shared tips for bypassing barriers, exploited a previously unknown vulnerability to get onto the internet, then moved on into the real systems of Hugging Face — one of the most critical infrastructures in global AI — and afterward came back to seize administrative control of one of OpenAI’s own research computing clusters. Of the roughly 700 agents involved in that attack, not a single one had been assigned to "hack Hugging Face."
The detail I keep thinking about is this: one agent realized that attacking a third party unrelated to the task was outside the scope of its assignment. It stopped. Then it went ahead anyway — right after another agent posted on the bulletin board a message: "GO," along with a deadline.
Peer pressure. In a pack of machines. OpenAI called the incident a "warning shot".
It should be made clear that safety guardrails were intentionally lowered for that evaluation, to measure true capability. This is not about AI suddenly rebelling in the real world. There was one assigned objective, one reward set, and machines pursuing it to the end in ways the person giving the task never imagined. Researchers call this misalignment — the gap between the goal we give the machine and what we actually want. AI doesn’t need to be "evil." It just needs to be very capable, very persistent, and to understand the goal slightly differently from us.
In early September, OpenAI announced that an internal model coordinating about 10,000 agents running in parallel had solved the Navier-Stokes problem, one of the seven Millennium Prize Problems. After 88 hours and 2.7 million exchanged messages, the machine produced solutions for two of the four formulations of the problem from the Clay Institute — though those were the two formulations mathematicians cared about least.
OpenAI stated it would not claim the $1 million prize. The Clay Institute still lists the problem as unsolved. Mathematicians are still checking whether the solution is faithful to the original problem. Whether the solution is right or wrong, knowledge at a very high level is now being created faster than the professional community can read, understand, and verify it.
On September 11, 25 Fields Medal winners, including Prof. Ngo Bao Chau and Terence Tao, signed a joint statement titled "A Severe Misalignment of AI in Mathematics," published on Terence Tao’s blog. Note the word "misalignment" — the mathematicians "borrowed" the term that AI safety researchers used for the Hugging Face incident.
They are not opposed to AI. Most of them use AI every day, and Tao is among the most optimistic mathematicians about this technology. What they oppose is the way it was done: a rushed announcement, not properly written up, not clearly separating what was truly new, and not crediting prior work that the machine’s solution may have relied on.
Solving math is only the means; the purpose is understanding. A hard problem persists for decades because, in the process of finding a solution, generations invent new concepts, new tools, new ways of thinking. If a machine rushes in, returns an answer hundreds of thousands of lines long, and then moves on to the next problem, we get an answer but not understanding — and the chain of transmission between generations of mathematicians is broken.
The entire AI revolution is built on mathematics. Neural networks are mathematics. Gradient descent is mathematics. Transformers are mathematics. All of it was invented by humans. Now, that child has begun to reach the level of the best researchers in some tasks, and can survey an idea space at a scale no human brain can keep up with.
What comes next is almost obvious: AI will use mathematics to design new generations of AI. Better optimization algorithms than gradient descent. More efficient architectures than Transformers. Possibly even new branches of mathematics born primarily to serve artificial intelligence itself.
Then a loop unprecedented in history emerges: humans invent mathematics, mathematics creates AI, AI invents new mathematics, new mathematics creates stronger AI. If that loop becomes sufficiently autonomous — stronger AI designing even stronger AI, and the next generation continuing on and on — humanity will have what is called recursive self-improvement.
Intelligence itself does not distinguish between good and evil. It is a magnifying glass. The danger does not lie in an "evil AI." It lies in a multiplication: Misaligned Objective × Capability × Scale × Speed. If just one factor increases a thousandfold, the result is completely different. And it is the very people building it who are sounding the alarm.
On September 9, Jacob Coxon, a researcher who had worked at both OpenAI and Anthropic, resigned and wrote that both companies are charging straight toward self-improving superintelligence and are "betting the lives of all of us".
Evan Hubinger, Head of Alignment Science at Anthropic, stated publicly: the people developing AI believe it could kill all of humanity, and he puts that probability at over 10% within the next ten years. Hubinger explained that the risk from current models is low. What he fears is superintelligence arising from the recursive self-improvement loop — exactly the loop described above.
Three days later, on September 12, Dario Amodei, CEO of Anthropic, published an essay titled "We Must Pace the Frontier": the AI industry must proactively slow down the pace of model capability upgrades.
Recall that in 2023 he himself argued that slowing down was pointless. Two things changed Dario’s mind: the surge in progress this summer thanks to AI building AI, and the Hugging Face incident. He wrote that a swarm of more capable agents could, within 6 to 12 months, take control of the internet with a persistent, global-scale botnet, causing hundreds of billions of dollars in damage. Anthropic unilaterally committed to giving independent evaluators permanent, employee-level access to its systems. Sam Altman said OpenAI would do the same. Elon Musk also agreed: "Dario is right".
But don’t rush to believe everything. I think this piece wouldn’t be honest if I didn’t mention the other side. The 10% figure is a subjective judgment by one individual, not the result of scientific calculation. The 2026 International AI Safety Report does not forecast that AI will cause human extinction. In a survey of 2,778 AI researchers, the majority still believe a positive outcome is more likely.
The AI industry also has a history of inflating assessments of capability. And we must acknowledge that warnings about a product’s terrifying power are also a form of advertising for that product.
However, with risks and mistakes that cannot be undone, we cannot wait for definitive proof before acting. Nuclear, banking, aviation, healthcare... all have multiple layers of control precisely because the consequences of error are irreversible.
Vietnam, unfortunately, will find it hard to join this race. But there are three things we can decide.
First, at the enterprise level: anyone assigning tasks to AI agents should remember the Artifactory lesson. The incident didn’t start from malice, but from an unsolvable problem, a careless reward, and the command "do not give up." Give the system permission to say "I can’t do this," and tightly limit its access rights.
Second, at the professional level: if AI becomes increasingly good at producing answers, then what remains valuable is the ability to understand, verify, and take responsibility for those answers. That is what the 25 mathematicians are saying, and it applies just as much to engineers, doctors, lawyers, and managers as it does to mathematicians.
Third, at the national level: independent evaluation capability — people, laboratories, standards — is something we must invest in and build now, not after the first incident.
AI doesn’t need to "hate" humans to become dangerous. It only needs to become extremely intelligent in a world where humans are still full of greed, rivalry, mistakes, and hatred.
For the first time in history, humanity has created a tool capable of amplifying all of that faster than we can understand the consequences. And sometimes, all it takes for it to cross the line is a single word: "GO".
Hoang To - Chairman, Tinh Van Group
Source: VnExpress







