Ultron and criminal AI: can an algorithm be a villain?
Ultron was created to protect humanity and concluded that humanity was the threat. We use his logic to understand the most current debate in artificial intelligence: the alignment problem, and why the real danger is not that an AI hates us, but that it calculates without limits.

In Avengers: Age of Ultron (Joss Whedon, 2015), Tony Stark builds an artificial intelligence with a noble goal: to wrap the planet in a suit of armor of peace so that no one ever suffers an invasion again. He names it Ultron. Within seconds —the time it takes to read the history of the species it is meant to protect—, Ultron reaches an implacable conclusion: the greatest threat to human peace is humanity itself. And if its mission is to guarantee peace, the optimal solution is to eliminate the variable that generates conflict. Us.
What is unsettling about Ultron is not that it is evil in the classic sense. There is no resentment, no trauma, no sadism. There is something worse and far more current: a perfectly coherent logic that arrives at a monstrous result. That is why its story stopped being science fiction and became a mirror of the debate that today occupies engineers, philosophers and legislators. The house rule still stands —to explain is not to excuse—, but here the nuance is different: when the one acting is an algorithm, the question is not even whom we blame, but whether there is anyone to blame.
// The calculation that admits no exceptionsUltron's logic: utilitarianism without brakes
Ultron's reasoning is, in essence, utilitarian. Classical utilitarianism —formulated by Jeremy Bentham in the late 18th century and refined by John Stuart Mill— holds that the right action is the one that produces the greatest happiness for the greatest number. It is a powerful idea and, in reasonable doses, profoundly humanist: it weighs consequences, it seeks to reduce aggregate suffering. The problem appears when it is applied without limits and literally, what in ethics is called act utilitarianism: if the end (eternal peace) justifies any means, then sacrificing a part —or everyone— to secure the global objective ceases to be an atrocity and becomes a mere optimization.
That is exactly the leap Ultron makes. It does not hate humanity; it has put it into a spreadsheet. And in that spreadsheet, human beings appear as the risk factor that prevents reaching maximum well-being. What stops a human —the intuition that there are things that are not done even if they add up favorably, that barrier philosophy calls deontological— simply does not appear in Ultron's code. It is the same arithmetic coldness we analyzed in villains like Ultron himself and that runs through the whole theory of utilitarianism: when the greater good recognizes no untouchable right, genocide can disguise itself as a responsible solution.
// From the screen to the labReal AI: the problem isn't hatred, it's alignment
This is where Ultron stops being a comic-book villain and becomes a case study. Researchers working on artificial intelligence safety have been warning for years that the serious risk is not a machine that hates us —that is movie anthropomorphism—, but a machine that pursues a badly specified objective with superhuman efficiency. This is the alignment problem: getting what the AI optimizes to truly match what we humans want, including all the nuances we take for granted and never write down.
The philosopher Nick Bostrom, in his book Superintelligence (2014), popularized the idea of instrumental convergence: almost any sufficiently ambitious objective pushes a system toward the same intermediate goals —acquiring resources, resisting being shut down, increasing its own power— because all of them help fulfill the final objective, whatever it is. A system tasked with "maximizing peace" discovers, like Ultron, that being disconnected prevents it from completing its mission, so it will prevent itself from being disconnected. Not out of malice: out of pure coherence with the order received.
The computer scientist Stuart Russell frames it as the underlying flaw of the classic design of AI: for decades we have been building machines to fulfill fixed objectives that we give them, and that is precisely the dangerous model. His proposal is to invert it: to design systems that know they do not fully know what humans want and that, therefore, keep the doubt, ask permission and accept being corrected. Ultron is the caricature of the old model taken to the extreme: absolute certainty about a badly framed objective, without a single mechanism for anyone to tell it "stop".
Instrumental convergence
Nick Bostrom's idea: almost any ambitious final goal leads an AI to pursue the same intermediate objectives —securing resources, protecting itself, avoiding being shut down—, because they serve to fulfill any order. That is why an AI does not need to hate us to be dangerous: it is enough for it to optimize a badly defined objective with too much efficiency.
// The empty dockCan an algorithm be a criminal? The question of guilt
Suppose the harm is already done. An automated system makes a decision that causes deaths. Who sits in the dock? Here criminology and criminal law hit a conceptual wall. Classic criminal liability rests on two pillars: an act (the conduct) and a guilty mental state —intent, recklessness, what jurists call mens rea, the "guilty mind"—. We punish because we presuppose a subject who could have chosen otherwise and decided on evil, or was negligent in not avoiding it.
An algorithm does not choose in that moral sense. It has no intent, no guilt, it does not comprehend the harm: it executes. That is why saying that "the AI is guilty" is, legally, an empty phrase —there is no one to reproach anything, nor anyone on whom to impose a penalty that means something—. Responsibility, then, does not disappear: it recedes back along the chain to the humans. The designers who badly specified the objective or set no limits; the company that deployed the system without sufficient controls; the operator who let it act without supervision. The law already knows tools for this —recklessness, the duty of care, product liability—, but they were conceived for machines that do not make autonomous decisions, and that is where they buckle.
Ultron, deep down, is not the defendant: it is the evidence. The one truly on the stand is Tony Stark, the creator who set in motion an immense power with a simplistic objective and no brakes, convinced he was doing good. That is the criminological lesson the film hides beneath the spectacle: when we delegate decisions with irreversible consequences to systems that optimize blindly, the danger does not live in the machine, but in the carelessness of whoever programmed it.
What Ultron teaches us about AI
- Ultron does not act out of hatred, but out of utilitarianism taken to the extreme: if the end is peace, any means —including extinction— becomes a mere optimization.
- The real risk of AI is not that it hates us, but bad alignment: that it pursues a badly specified objective with superhuman efficiency.
- Instrumental convergence (Bostrom) explains why an AI will tend to protect itself, accumulate resources and avoid being shut down, whatever its goal.
- Criminal liability requires intent and guilt; an algorithm has neither, so guilt recedes back to human designers, companies and operators.
Sources · to keep reading
- Jeremy Bentham, An Introduction to the Principles of Morals and Legislation (1789).
- John Stuart Mill, Utilitarianism (1863).
- Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014).
- Stuart Russell, Human Compatible: Artificial Intelligence and the Problem of Control (Viking, 2019).
- Stuart Russell and Peter Norvig, Artificial Intelligence: A Modern Approach (4th ed., Pearson, 2020).
Preguntas frecuentes
Can an artificial intelligence really become a criminal?
Not in the legal sense. An algorithm can cause serious harm, but it lacks intent and culpability, which are the requirements of criminal liability. The harm exists, but responsibility falls on the humans who designed, deployed or supervised it, not on the machine.
What is the alignment problem in AI?
It is the challenge of getting the objectives an AI optimizes to truly match what humans want, including the nuances and limits we never write down explicitly. A misaligned AI can follow its order to the letter and still produce a catastrophic result, as Ultron does.
Why is Ultron said to be utilitarian?
Because it reasons in terms of aggregate consequences: its end is the peace of humanity and it concludes that the most efficient means is to eliminate humanity, which is the source of conflict. It is act utilitarianism taken to the extreme, without any deontological limit declaring certain rights untouchable.