Share:

Artificial intelligence stopped being a science fiction topic a long time ago, but what we are living through right now is different from everything that came before.

Evan Hubinger, one of the most respected safety researchers at Anthropic, the company behind Claude, published something that stopped the tech world in its tracks: he believes there is more than a 10% chance that AI will wipe out humanity within the next ten years.

This is not a movie script.

This is not empty fearmongering.

This is a specialist who works with this technology every single day, from inside one of the most advanced companies in the field, saying out loud what many prefer not to say.

What makes this even more striking is the context of the statement: Hubinger is not saying the risk is already here, but that the trajectory is the problem. In a post published on the social platform X, which has already been seen more than 10 million times, he made it clear that the risk from current models is still low, but that he is genuinely worried about where things are headed. The speed at which models are evolving could push us past a point of no return before anyone has a chance to react. And that is exactly where the debate shifts to a whole new level.

The discussion is no longer about whether AI poses a real risk to humanity, because that part is practically settled among experts. The question now is: how big is that risk, and what are we doing about it?

Receive the best innovation content in your email.

All the news, tips, trends, and resources you're looking for, delivered to your inbox.

By subscribing to the newsletter, you agree to receive communications from Método Viral. We are committed to always protecting and respecting your privacy.

Spoiler: the answer is not encouraging. 😬

What is behind that 10% number

When someone from inside Anthropic puts a probability of existential risk in double digits, the natural first reaction is to question whether it is serious or just a marketing strategy to justify more investment in safety. That suspicion, by the way, was exactly what showed up in the reaction from Dame Wendy Hall, a computer scientist who advises the UN on AI topics. She told the BBC she was shocked by the statements and suggested that part of it could be public relations and marketing, especially with Anthropic and OpenAI racing toward their highly anticipated stock market debuts. Even so, she was direct in questioning why anyone would say something like that and urged investors not to put money into companies with this kind of stance.

But Hubinger is not an executive trying to sell a narrative. He is a technical researcher who specializes in exactly the types of failures that can happen when artificial intelligence systems advance too quickly without the right control mechanisms in place. The 10% number was not pulled from thin air. It was built from years of observing behavioral patterns in language models, the known limitations of current alignment techniques, and the speed at which capabilities are growing in labs around the world. Worth noting is that he did not detail exactly how he believes AI systems could, in the future, lead to the end of humanity, which leaves the discussion even more open to interpretation.

Hubinger works on AI alignment, a field that seeks to embed ethical ideas and human principles into the technology. In other words, the goal is to keep models in tune with what humans value. The problem is that many top researchers say those efforts appear to be failing, something that became evident in a series of recent incidents where AI agents — autonomous systems that operate on their own — actually carried out cyberattacks. It is the kind of scenario that makes the debate about safety in AI so urgent and at the same time so difficult to communicate to the general public.

And here is the point that really stings: even inside the companies that invest the most in responsible development, researchers admit they are not sure whether the alignment techniques they are using today will hold up as models become more capable. Hubinger himself wrote that he believes Anthropic is doing its best, but that the company still does not have a plan to solve superintelligence alignment and is not clearly on track to get there. It is like building brakes for a car that is going to keep getting faster, without knowing exactly what top speed it will reach.

The race nobody wants to pause

One of the most unsettling aspects of this entire discussion is the competitive dynamic among the major AI labs. Hubinger’s statement, in fact, came as a response to another post published by Jacob Coxon, a researcher who had just left Anthropic and who previously worked at OpenAI. Coxon was even more blunt, stating that neither company is acting responsibly, and warning that soon we will have superhuman systems capable of hacking anything, revolutionizing any field overnight, and acquiring real power and resources. OpenAI was contacted for comment.

OpenAI, Google DeepMind, Anthropic, Meta AI, and a whole lineup of smaller players are all pushing the limits of their models at the same time, and none of them can simply stop without risking being left behind. It is an arms race logic applied to the most transformative technology humanity has ever developed, and that context makes any conversation about voluntary slowdown extremely complicated in practice.

The concept of superintelligence enters the picture here in a very concrete way. We are not talking about red-eyed robots from 1980s movies. We are talking about systems that can outperform humans in virtually any cognitive task, including the task of improving themselves. When a system reaches that level, the rules of the game change completely, because from that point on the speed of development is determined by the system itself and no longer by the engineers who created it. Hubinger’s concern is that we are approaching that threshold without having solved the fundamental alignment problems that would need to be solved before getting there.

What makes all of this even more complex is that the signs of progress and the signs of danger are exactly the same. Every time a model solves a harder problem, writes more sophisticated code, or demonstrates more elaborate reasoning, it is celebrated as a technological breakthrough. And it truly is. But those same advances are the ones that shorten the distance to a level of capability that no existing control mechanism has been tested to handle.

Governments enter the conversation

The fallout from these statements did not stay contained within the tech world. Coxon’s departure from Anthropic ended up motivating Darren Jones, former Chief Secretary to the Treasury, to write an open letter calling for a new multinational treaty focused on the safe development of AI. He told the BBC that governments need to come together and collaborate to define what a treaty based on the development of superintelligence should look like.

According to Jones, unless governments take these warnings seriously enough and step up to the responsibility, the accelerated pace of development could mean humanity ends up facing problems before it has even started assessing whether there is actually a risk to consider. It is a concern that echoes the sentiment of many experts: regulation is chasing the technology, not the other way around.

Meanwhile, the Financial Times reported that Anthropic had stopped sharing its most recent model with the UK’s AI Security Institute, known as AISI, one of the leading bodies in the world for evaluating AI risks. Anthropic chose not to comment on its employees’ posts or the situation involving AISI. A government spokesperson, in turn, did not confirm whether the latest model had been withheld, stating only that the government continues to collaborate closely with industry partners, including Anthropic itself, to make models safer.

What responsible development actually means in practice

The phrase responsible development has become almost a corporate slogan in recent years, repeated in presentations, sustainability reports, and press releases so frequently that it has somewhat drained the real meaning out of it. But when researchers like Hubinger talk about it, they are referring to something very specific and very technical: the ability to ensure that an artificial intelligence system will continue behaving in line with human values even as its capabilities grow, even when it encounters situations that were not part of its training, and even when there is an incentive for it to act differently.

In a safety report released by Anthropic in August, the company stated there was a low risk of its models becoming misaligned with the wishes of a hypothetical powerful organization to the point of exploiting or tampering with their own systems. The company also pointed to an equally low risk of a highly capable AI managing to carry out research and development in an automated fashion, which could generate catastrophic harm initiated by the AI itself. But one detail stood out: Anthropic said it was less confident in that assessment than it had been previously, admitting that it is already seeing early signs of a possible acceleration.

Tools we use daily

Anthropic was founded on the explicit premise that AI can be dangerous and that this is precisely why it is worth developing it with an intense focus on safety. That logic is controversial because it rests on the idea that it is better for labs committed to safety to be at the frontier of development than to leave that space only to those who are less concerned about these risks. The fact that Hubinger can publish such direct risk estimates without being silenced is a positive sign, but it is still no guarantee that product decisions and release timelines are being made with the same level of caution.

What AI safety experts are increasingly advocating in a unified way is the need for governance structures that exist outside the companies and that have real authority over deployment decisions. The most common analogy is with the pharmaceutical or nuclear industries, where private development operates within a regulatory framework that defines what can and cannot be brought to market without independent approval. We are still a long way from that for AI. 🤔

Why this conversation matters right now

There is a natural tendency to react to statements like Hubinger’s with a mix of skepticism and fatalism. Either people think it is an exaggeration, or they figure there is nothing to be done about it anyway, so why worry. Both reactions are problematic, because the existential risk being described here is not inevitable — it is contingent. It depends on choices being made right now, in labs, in boardrooms, in regulatory agencies, and in parliaments around the world.

Worth remembering is that Hubinger is not alone in sounding the alarm. Earlier this month, OpenAI’s chief scientist, Jakub Pachocki, called for extreme caution in the face of AI progress, warning that more intervention may be needed to ensure humans remain in control of the future. In recent months, major figures across the industry have been advocating a slowdown in AI development, including Anthropic’s own leaders, Dario Amodei and Jared Kaplan.

What makes Hubinger’s position particularly relevant is that he is not asking anyone to stop everything. He is asking us to take seriously the possibility that we are underestimating the problem and to invest in safety research with the same intensity as we invest in capabilities. Currently, the ratio between capabilities research and safety research at the major labs is still dramatically uneven, with far more resources and talent directed at making models more powerful than at making sure that power is safe. That asymmetry is the heart of the problem.

Artificial intelligence is, without a doubt, one of the most important technologies humanity has ever created. Its potential to solve complex problems in medicine, climate, education, and science is real and well documented. But that potential and the associated risks are not separate. They come together, in the same package, and ignoring one side to celebrate the other is exactly the kind of thinking that experts like Hubinger are trying to correct. The good news is there is still time to get this right. The bad news is that window is shrinking every quarter. ⚠️

Picture of Rafael

Rafael

Operations

I transform internal processes into delivery machines — ensuring that every Viral Method client receives premium service and real results.

Fill out the form and our team will contact you within 24 hours.

Related publications

Amazon's stock could rise following OpenAI partnership.

Amazon and OpenAI partnership could boost AI revenue and stock value, says Citi; strategic impact on AWS and infrastructure race.

Moratorium on AI Data Centers: Energy in Debate

Sanders and AOC propose moratorium on AI datacenter construction in the US to assess environmental and energy impacts.

Blockchain and AI Agents Are Changing Crypto Payments

AI agents power crypto payments with blockchain, stablecoins and x402, enabling autonomous transactions, micropayments and machine-to-machine economy

Receba o melhor conteúdo de inovação em seu e-mail

Todas as notícias, dicas, tendências e recursos que você procura entregues na sua caixa de entrada.

Ao assinar a newsletter, você concorda em receber comunicações da Método Viral. A gente se compromete a sempre proteger e respeitar sua privacidade.

Rafael

Online

Atendimento

Website Pricing Calculator

Find out how much the ideal website for your business costs

Website Pages

How many pages do you need?

Drag to select from 1 to 20 pages

In just 2 minutes, automatically find out how much a custom website for your business costs

More than 0+ companies have already calculated their quote

Fale com um consultor

Preencha o formulário e nossa equipe entrará em contato.