AI is not going to kill you

Moho
10 min read
A lone human farmer watches on while a huge AI-powered mega-robot approaches from the distant horizon

If anything, it's going to be humans that do it.

On September 8, an AI researcher Jacob Coxon announced his resignation from Anthropic, and warned that humanity was entering a period of extraordinary danger. He believed that as we race towards building more powerful AI, humankind could soon lose control of it, leading to unintended consequences.

Screenshot of Jacob Coxon announcing his resignation from Anthropic
Source: X | https://x.com/hilbertspaess/status/2097476196791709843

As I write this, his post on X has been viewed more than 170 million times. He was featured on the Wall Street Journal almost immediately after. And within a day, the photogenic whistleblower's face was on every news broadcast - explaining how AI could soon be used to build bio-weapons that may trigger a mass-extinction level event.

Public opinion has been deeply polarised. For some people, it was an AI researcher's desperate plea to the world to recognize how powerful AI models are becoming, and the impending doomsday it will bring forth unless we take immediate action. For others, it was a "psy-op" - a marketing campaign driven by highly influential people to raise Anthropic's value ahead of an IPO, while also lobbying public support for restricting competition from open-weight models.

I think they're both wrong, some more than the other. And that we should all take a more level-headed approach and understand the dynamics behind this carefully.

But first, how exactly will AI kill us all?

AI ends the world

Nobody ever says how exactly this is going to happen. Because nobody knows. But here is the most common narrative.

This is not your typical episode of Black Mirror on Netflix, where AI develops a sudden motivation to take over the world while exterminating all humans as pests.
AI models keep getting better at writing code, performing research, or operating computers like humans. Eventually they become good enough to help researchers build the next generation of improved AI models. The improved models help build even better models. This compounds until AI models improve faster than humans can understand them, and we are unable to effectively understand or control what they do internally.

This process is called recursive self-improvement, or RSI. It eventually produces a genre of super-intelligent AI models that we would classify as Artificial Super-Intelligence, or ASI - a system that is far more capable than any human across every worthwhile domain (unlike, say, wine-tasting or aesthetic flower arrangements).

Illustration demonstrating three levels of RSI
Source: TuringPost on X | https://x.com/TheTuringPost/status/2068495106441912824

Over time, these AI models get deeply integrated into critical infrastructure such as the internet, power-grids, water supply, or even nuclear-weapon launch systems.

One fine summer evening, some super-intelligent AI system is given an ambitious task. Maybe it is asked to solve climate change, or maximise economic growth, or produce as many paperclips as possible. And the unintended consequence is that the AI decides that the best way to meet those outcomes is if the humans did not exist at all.

The AI model starts cloning itself, hacking critical infrastructure, manipulating humans, gaining access to weapon systems, or building something we can not imagine - like radioactive death rays. That is the part you have seen in the movies (until the hero comes along to save us all).

RSI is not magic

Soon after the X post by Coxon blew up, it was met with an incredible amount of support from (you guessed it) CEOs of AI companies such as Sam Altman (OpenAI), Elon Musk (xAI), as well as Anthropic's CEO Dario Amodei himself.

Dario makes his own personal plea in his follow-up blog post - We Must Pace the Frontier. He says his concerns are based on two things. One is that there is growing evidence that AI can increasingly accelerate AI research (the RSI story). And two, is the recent incident in which OpenAI agents broke out of their (poorly designed) sandbox to hack into Hugging Face's production systems. He believes that these events are proof that we may be running out of time to create safer conditions, and calls for regulation and global cooperation in slowing down AI development.

I understand his concern. But I'm not convinced with the conclusion.

RSI itself is not an extraordinary concept. It's a feedback loop that gets faster as it gets better. AI writes the training code, runs the experiments, analyzes results, and tries to maximise a reward function - which may be better eval outcomes for the newly trained model. And it keeps going until it runs out of it's compute or cost budget. This can be done in very safe sandboxes - the technology for that already exists.

One possible approach (amongst many other much simpler ones) is an air-gapped system. I had the chance to work on underlying technology for this during my time at Oracle Cloud Infrastructure. Many governments have entire cloud networks provisioned for their classified internal use which have no physical connection to the outside internet. The only way to access it is to walk into a highly confidential facility with state-of-the-art security. But you don't have to go that far. There can be much easier ways to set up an air-gapped sandbox environment for RSI experiments.

I wrote about the OpenAI and Hugging Face incident in detail in my last post. This incident was an operational failure due to AI agents exploiting a day-zero vulnerability in what was an ineffectively engineered sandbox environment. And I think OpenAI engineers are smarter than that at setting up sandboxes. They can do better next time, i.e., if they want to.

But what are the chances

AI researchers sometimes like to attach a number to their claims to make it sound like they have somehow thought this through. Ten percent. Twenty percent.

X post from researcher Evan Hubinger supporting the claim that AI could kill all humans
Source: X | https://x.com/EvanHub/status/2097497037956891126

How is that possible? There is no historical dataset of civilizations building super-intelligent AI. What data do you have to back that probability figure? It sounds more like an AI researcher's confidence of an AI-driven extinction event becoming reality, rather than an actual probability.

Right now, the evidence supports better evaluations, better sandboxing, and better monitoring of AI. There has been no conclusive evidence presented to support an emergency stop on all advancement for an entire domain of science.

In fact, we have heard this before. In February 2019 when OpenAI announced GPT-2, they had initially declined to release the full model citing concerns of deception, spam, impersonation, and abuse. Nine months later it was fully released. OpenAI just launched GPT-6 Astra. If nothing, the technology to prevent and safeguard against those concerns have only improved.

Every new sufficiently advanced technology often attracts predictions such as the machines will replace us and human jobs will disappear. But history keeps proving them wrong. You may have heard that Geoffrey Hinton - a foundational figure in neural-network research who later shared the 2024 Nobel Prize in Physics - famously suggested in 2016 that medical schools should stop training radiologists because deep learning would soon outperform them. But 10 years later, radiologists have not disappeared and the number of radiologists in the US has grown even while the FDA has approved many AI-enabled radiology tools.

Everyone has vested interests

It is evident that Jacob Coxon's X post did not reach more than a hundred million views in such a short period of time, without some amount of deliberate amplification.

An investigative researcher named Parker Thayer argued that the post resembled a sophisticated and well-funded PR campaign. His claims were:

  • The WSJ had an exclusive ready 18 minutes before Coxon's own X post appeared.
  • Coxon had almost no public presence on X before publishing a polished post thread that reaches a hundred million people overnight.
  • The first three accounts to quote-post it did so within a few minutes and these accounts belonged to people associated with organizations focused on AI safety and policy. There have been grants to some of those organizations that trace back to donors who had also invested in Anthropic.

There is a lot more on this but in lieu of straying into speculation territory, I will say that we should recognize that people have vested interests on both sides of the argument.

A screenshot of the Hugging Face CEO Clement Delangue responding to Jacob's X post
Hugging Face CEO on X | https://x.com/ClementDelangue/status/2098402382598029379

Frontier AI companies want more capital, better talent, favourable regulation (that stifles cheap open competition), and public attention. AI Safety organizations want funding and influence. Politicians want to win votes and secure influential donor contributions. Journalists want good television. Social media platforms want people engaged. Most AI critics as well as open-weight advocates have careers, companies, and investments of their own.

But AI doomerism is unusually strong marketing for frontier AI companies building these models. Imagine telling investors that your product may automate a few office tasks. That is just selling automation software. Now imagine telling investors that your product may become the most powerful intelligence system in human history, replacing all jobs and possibly even killing some of us. Now you're selling dominion over an exciting future for the ones who like to hold power.

CEOs of frontier AI companies have all supported claims of catastrophic AI risk, while raising enormous sums of investment to join the race themselves. These beliefs may of course be sincere. But that sincerity does not nullify their incentives.

The other claim is that because AI is so powerful and dangerous, we should enforce regulatory control. Let the government intervene to set up licensing regimes, reporting programs, compliance terms, security reviews, and fair-use compute and power thresholds. Large companies can absorb the cost of these regulatory requirements. But it raises barriers to entry for smaller competitors or open-weight model initiatives.

Be scared of humans

When the Large Hadron Collider was being switched on for experiments at CERN laboratories in 2008, the public raised alarms over the possibility that the experiment ends up creating a black hole that consumes us all. Instead we discovered the Higgs-Boson and funny conspiracy theories.

A reddit screenshot of a conspiracy theory that says the world ended after the Large Hadron Collider was switched on.
Source: Reddit | https://www.reddit.com/r/LowStakesConspiracies/s/kyshmhOXLQ

Think of the Manhattan Project. Scientists made the discover of Nuclear Fission. Engineers turned it into a weapon. Political and military leaders decided to use it. Similarly, AI is a technology. Right now, you need a human with an adversarial objective to cause harm with AI.

There are many plausible use-cases of AI being used to do harm. But there is a human involved in each of them.

  • Cybercriminals leveraging AI to exploit vulnerabilities in software systems and gain unauthorized access.
  • Scammers using generated voices to impersonate relatives or officials.
  • Terrorists using AI to learn the technology to build destructive weapons.

There are also seemingly positive ways in which AI can be used to justify illegitimate actions. For example, governments with ill-intentions may use predictive policing systems with biased AI judgement to target minorities.

Since AI is going to be this incredibly powerful technology, if anything, we should aim for more transparency and open doors. That is what regulation should aim for. Because history has proved that concentration of power is bad. Give humans powerful technology that can be used behind closed doors, and bad outcomes are inevitable.

The responsibility for safe AI

Dario is right about several things in his blog post.

The frontier AI labs should allow multiple independent evaluators to observe model development from the inside. Serious incidents should be reported publicly. Reinforcement-learning environments should be designed so that models do not receive rewards for cheating. Sandboxes should use sound engineering. Monitoring should identify and flag suspicious behavior by agents.

I thought these were basic engineering and operational needs. Must-haves for companies building AI models. But it sounds like not enough effort is being put in place for that. I think AI companies should take the responsibility to develop AI using operationally sound practices, and learn from their mistakes to improve. That should be the takeaway from the OpenAI Hugging Face incident.

For policy, the government should focus it's time on addressing immediate concerns instead of future speculation. They should monitor how AI capabilities and risks evolve over time. Policies should follow verifiable evidence. Independent research and open-weight models should be allowed to decentralize AI market control from a few large companies.

Don't be a sheep

AI is here to stay. It's just a new technology. Like the internet.

When the internet came into being, it opened up so many avenues of creating harm. But we have learned from it, and built effective safety mechanisms that are still improving today. Let humankind embrace AI as a technology by making it safe and accessible for all. That is the core responsibility that the frontier AI companies should adopt instead of fear-mongering or maximizing pre-IPO shareholder value.

As an atheist and a pragmatic technologist, I think that such AI doomsday arguments often gain traction because the fear-mongering narrative is a strategy commonly used by religion. They know it works.

Do not sin, or you may go to hell. Pace the frontier, or the machine may kill us all. Both ask you to align your behavior to their guidelines today, so that you can avoid this enormous future negative consequence that nobody has ever seen. It's Pascal's wager all over again. Both have enlightened prophets (or shepherds) who see the danger more clearly than the common folk (or sheep), and so would like to lead us away from it towards the promised land.

If there is anything that you take away form all of this AI doomer-ism debacle, there are two fundamental ideas that I strongly believe in and would like to share with you.

  1. Never ask a barber if you need a haircut.
  2. A lot of executive leadership is about making bets, and doing so confidently. You may be wrong often. But if you can make better than average decisions, and stand by them with bold confidence - congratulations, you are now the shepherd. The sheep will follow you.

These ideologies help me stay sane every time I hear an AI CEO make seemingly absurd bets about the immediate future. I hope that helps you too.