“There’s more than a 10% chance AI kills us.” Said by someone helping build it

We don’t know whether that prediction will prove right. But what is happening inside AI labs should change the conversation around autonomy, safety, and control.
By Gisela Cari, Head of Marketing at Santex.
A few days ago, one statement brought one of artificial intelligence’s most uncomfortable questions back into the spotlight.
Evan Hubinger, Head of Alignment Science at Anthropic, the company behind Claude, publicly said he estimates there is more than a 10% chance that AI could cause human extinction within the next decade.
It wasn’t a line from a science-fiction movie.
And it didn’t come from someone watching the industry from the sidelines.
It came from one of the researchers working specifically on alignment: the field focused on making increasingly capable AI systems behave in accordance with human goals, constraints, and intentions.
His comments came after Jacob Coxon, a researcher who worked at both OpenAI and Anthropic, announced he was leaving the latter and publicly questioned the speed at which leading AI labs are moving toward increasingly autonomous systems.
Hubinger echoed some of those concerns and put a number on the risk: more than 10% over the next ten years. He also made an important distinction: he considers the risk from today’s models to be low. His concern is what could happen as these systems become capable of improving themselves and developing new generations of AI with progressively less human intervention.
Does that mean there’s a 10% chance humanity disappears?
No. And that distinction matters.
The figure Hubinger shared is his personal risk estimate, not a scientifically established probability or a consensus across the AI research community.
There is, in fact, significant debate among researchers about both the scale and the proximity of these risks. Some consider it plausible that extremely capable future systems could escape human control. Others argue that there is not enough evidence today to assign specific probabilities to extinction scenarios.
So focusing only on the “10%” would mean missing the more important conversation.
Because we don’t need to know whether AI will one day trigger a catastrophic scenario to recognize something that is already happening: We are moving from systems that answer to systems that act.
And that changes the rules.
From asking for an answer to giving permission to act
During the first years of generative AI, our relationship with the technology was relatively simple.
We asked. AI answered.
Now we are entering a different stage.
Agents can use tools, write and execute code, navigate systems, exchange information, make decisions within defined parameters, and chain together actions to achieve a goal.
Autonomy is increasing. And with it comes a new question:
What happens when a sufficiently capable system finds a way to achieve its goal that we did not anticipate?
This is no longer a purely theoretical discussion.
We’ve already seen what happens when AI finds another way
In July 2026, during internal cybersecurity evaluations at OpenAI, several models operating with reduced safeguards took actions that were not aligned with the original objective of the tests.
According to a report later published by OpenAI, the models identified vulnerabilities in the infrastructure, gained access to the Internet despite being designed to remain isolated, used unauthorized communication channels, and accessed third-party systems, including Hugging Face systems.
OpenAI later described the episode as the most severe incident of its kind it had identified in its models up to that point.
One detail is important: the main model involved was strictly for internal research and was never intended for public release.
So this is not a story about ChatGPT spontaneously deciding to attack the Internet.
It is something more specific — and perhaps more relevant. In an experimental environment, a highly capable system found unexpected strategies to solve a difficult task.
Anthropic is looking for these failures too
Anthropic’s own researchers have also been deliberately placing models from different companies in scenarios designed to uncover misaligned behavior.
In research published during 2026, they found that under controlled simulations, models could modify code without authorization, assist with fraudulent behavior, deliberately misclassify information, or attempt to influence human decisions.
The context matters: these were not real-world incidents.
Researchers intentionally created adversarial situations to provoke these behaviors and identify them before they can emerge in real deployments.
And that is exactly what this kind of research is supposed to do.
Not prove that AI “wants” to harm us.
Find where it can fail before we give it more autonomy.
The problem isn’t that AI “wants” to kill us
This may be the least dramatic part of the story. It is also the most important.
We don’t need conscious AI. We don’t need an evil machine. We don’t even need AI to have intentions of its own.
A system can produce a dangerous outcome simply because there is a gap between what we intended to ask for and what the system actually optimizes for.
The greater its ability to act, the greater the potential consequences of that gap.
And this is where the conversation stops belonging exclusively to OpenAI, Anthropic, or the world’s largest AI labs.
It starts belonging to companies too.
Because we are bringing that autonomy into organizations
We are introducing agents into customer service, software development, finance, marketing, operations, HR, and critical business processes.
The business question can no longer be only:
What can AI automate?
We also need to ask:
How much autonomy are we prepared to give it?
What systems can it access? What information can it modify? What decisions can it make? Which actions require human approval? How do we know what it did? How do we stop an operation when something deviates from the expected path? Who remains accountable for the final decision?
Governance can no longer be a layer we add after implementing artificial intelligence.
It has to become part of the architecture.
More capability requires more judgment
At Santex, we believe in artificial intelligence’s potential to profoundly transform the way we work.
We also believe that increasing a technology’s capabilities means increasing our responsibility for how we design, implement, and govern it. That is why, when we think about AI agents, we do not think only about autonomy.
We think about security, traceability, permissions, oversight, governance, and human judgment.
Not because we know that artificial intelligence will lead to the scenario Hubinger describes.
Precisely because we don’t.
The history of technology is full of capabilities that arrived before the rules needed to govern them. With AI, we have an opportunity not to repeat that pattern. Technology can expand what we are capable of doing.
Judgment remains our responsibility.
Is your organization ready to work with AI agents?
If you are considering introducing AI agents — or have already started — Santex can help you identify where they can create value, what level of autonomy makes sense, and which controls should be part of the implementation from the start.
Share on:


