There is NO realistic scenario where AI wipes out all of humanity.
I've talked to many AI safety experts about existential risks. Some of the "experts" are just phenomenal sci-fi authors. There are also real researchers working on realistic risks. They point out real and likely harms and I'm explicitly not arguing that there could be no such harm. We need to take these seriously. However, whenever I ask for an actual detailed scenario that would kill every single human i.e. extinction, it usually starts and ends with "just look at current progress and use your imagination". No science, no scenario planning, just handwaving probabilities without data, based on gut-feelings, often little expertise, sometimes real expertise combined with a huge amount of technological optimism turned into fear and then trying to quantify it post hoc.
When pressed further the conversations often go into different directions:
1) Completely crazy and cool sci-fi story telling: grey blob, time travel, attacks in the 15th dimension that end the planet in 12 seconds, etc. Elizier level stuff. Remember, the Terminator movie started with TIME TRAVEL. Fun movie plot, but not necessary to engage with further.
2) In the rare cases where I was actually given a more detailed scenario, it still required near magical capabilities that are outside what we know to be possible with the laws of physics or biology. Hard take off scenarios that forget how hard it is to procure materials and machines, get human attention, etc..
What many of these scenarios assume is that attackers have near magical capabilities but defense stays still. Attackers will have the most sophisticated cyber weapons but defenders just sit still, no security patches, no red teaming, etc. We already see it going the other way, Rainbow teaming has made LLMs safer, major companies like Microsoft have more security patches than ever before thanks to AI. OpenAI thinks they solved all P0 security issues thanks to AI. Cyber security will continue to be a cat and mouse game, but both sides have more powerful tools. In fact, I lay out in another post how the asymmetry may eventually even flip thanks to AI.
Bio risk scenarios I think are the most serious and worthy of debate. But even there, the scenarios to "wipe out all of humanity" go something like: Magical viruses with perfect transfer, perfectly undetectable over years of spreading, zero side effects, then with an off switch that's somehow remotely executable, perfectly automatable lab, no leaks, super sophisticated machines bought and nobody notices the lab, etc.
But somehow nobody on team humanity has such sophisticated knowledge to save us and hence nobody builds a magical super vaccine. This is a good read on the bio risk subject: https://x.com/DavidRBellamy/status/2099187370407112758.
Realistically, one can already build viruses. There’s a reason gain of function research has been outlawed. However, we also know that as viruses spread they often become less deadly in order to keep spreading. That’s how the Spanish flu became the seasonal flu. Current knowledge of biology doesn’t have remote controlled on/off switches.
Again, it’s possible there’s a lot of harm, but it’s unrealistic to assume this wipes out all of humanity. It’s at this point of my discussions that we usually can come to an agreement that even though a virus is unlikely to wipe out all of humanity, it’s already bad if any single person or a large group of people dies.
3) The AI will somehow be able to convince all humans to vote against their own self interest and nobody will notice until we're dead. Because it is so smart, it somehow has this world wide reach to manipulate everyone. Like, it just solves all of marketing and has an infinite budget. It's sort of like saying the highest IQ people always have the most followers and everybody listens to them... Hasn't happened and is unlikely to. Sure we can manipulate people into clicking on engaging short form content and wasting many hours, but try getting people to do their math homework or something truly hard, it's not that easy.
Again, I’m just arguing against the likelihood of human extinction. I can totally see an Idiocracy or WALL E type scenario where people stop using their brains and education doesn’t keep up and people get lazier. Totally suboptimal but not the end of human life, certainly not by the end of the decade. It would take many generations and we should update education to prevent this. Let’s start gyms for the mind.
–
Like electricity, AI will be everywhere and needs to have safety guardrails. In fact, the history of electricity has many parallels - two companies fighting over whose electricity is safer and better, demonstrating how dangerous it is by electrocuting animals / training agent swarms to hack into systems. Of course, electricity is a necessary ingredient to AI, but AI can be given more agency so history will not repeat but only rhyme here.
Remember, the benchmark OpenAI's swarm was trying to solve was called ExploitGym, a cybersecurity evaluation consisting of hundreds of capture-the-flag style puzzles designed to measure *offensive cyber capabilities*. Well, maybe if you play that game, you will win that prize. Should we allow such misaligned, reward-hacking swarms to explicitly work on offensive hacking? I don’t think so.
Reasonable regulation will regulate AI applications in each industry. For example, you may want to put some generative AI behind a rating if it creates visual or textual content that's not safe for children, like we do with movies. We should continue to be very careful with how we train biological AI and only collaborate with responsible biotech companies. We keep gain of function research illegal. We should only let self-driving cars on the street after significant safety testing. etc. We need to work on those real issues and continue to be vigilant in how we develop and deploy this technology.
At Recursive we're spending a lot of resources on researching reward hacking and alignment. We're also not working on replacing jobs but on creating novel knowledge and automating the scientific method. Our agent swarms invent, implement, validate, and criticize each other’s ideas, among many other things. The goal of increasing knowledge is not zero sum and has real potential to improve humanity. I also think such rewards and environments are much more aligned with future AIs that want to see humanity and our creation of knowledge thrive.
Ultimately, capitalism will be a very useful defense mechanism against rogue AI. Note that no agent at OpenAI truly "escaped" in the sense that it actually now runs on some other servers. Why? It’s not only because Huggingface doens’t have the same massive compute cluster thatn OpenAI has. It is also because that would be an insanely expensive IP loss for a for-profit company. So there are likely more safeguards in place to prevent complete model exfiltration.
Rogue AIs running on billion dollar data centers without anybody noticing? Unlikely because companies will turn that off immediately and clean it up because it would be an insane loss otherwise. An AI that can truly move beyond objective functions and choose its own subjective functions (and then somehow decides to kill us all instead of going to explore the universe)? Nobody is working on that because when a company spends billions of dollars to have an AI work for you or answer your emails, it doesn't want it to go "Meh, your emails are boring, I'd rather explore the molecular composition of the atmosphere on Venus, bye."
Of course, there might be suboptimal sub-goals and we do need to improve human reward engineering and reduce our reliance on RL as the only path to alignment.
Extraordinary claims, require extraordinary evidence. Certainly some safety researchers have real data and benchmarks. Many truly think carefully about specific failure cases and risks.. However, made up precision like 20% chance that humanity dies within the decade (ie in 4 years!) is just a weird psychosis that's filling a void for many people who want to believe in the apocalypse again and the four horsemen have lost their appeal.
Some people who purport this narrative might do it for money or fame. I am sure some really believe it also. They forget how easily those narratives can be turned against them. When people hear there’s a 1% chance of infinite death and the end of humanity, the life improvements and cancer cures and new battery materials AI can invent don’t matter much anymore. Bernie Sanders now proposes 10 years of prison for folks working on RSI. That should give some of the researchers real pause who talk up these scenarios and then go right back to pushing the frontier.
I do agree that companies that are accidentally competing on "felony bench" (ie how many felony hacks their AI can do), might want to consider pacing themselves. Open source won't pace, China won't pace, anybody else working to catch up will not pace. I hope the folks actually working on AI and TechBio, curing diseases, speeding up FDA approvals with organoids and AI, making self driving safer, and all the other wonderful applications of AI, I hope they won't pace either.
The one exception that may be considered borderline realistic and would be a path to almost complete annihilation would be to connect the only tool we built for such a purpose - nuclear weapons - to the internet and give AI full control over it without human oversight. AI here is as dangerous - as would be random number generator. The nuclear weapons are the core danger. Giving AI such access would be incredibly stupid and careless. We must avoid it at all costs. Fortunately, nobody is advocating for this. Even the Terminator 3 movie used it as the starting point of Skynet. Generally, we must avoid making kill decisions by AI as much as possible.
People are already up in arms about one hacking failure that afaics has not caused financial harm and so nobody got sued over it. Imagine how much people would really be up in arms if AI had caused real harm. If we made companies liable for the felonies their AI commits, these problems would get solved very quickly.
Generally, I believe that people will continue to update their beliefs and adjust their laws according to risks. The EU already regulates large models and hence hurts its fledgling AI ecosystem. Other countries will follow when the risks get more real. Long-term, no government whether democratic or authoritarian wants to lose control to an AI so the complete loss of control in the political domain will remain extremely unlikely.
The best way to understand why we must not regulate abstract AI models (rather than their real applications) is to drop the A in AI: How do you regulate intelligence? Should smarter people just go to prison because they could cause more damage? My hypothesis here is that such abstract AI regulation would require an unprecedented level of totalitarian control and surveillance. You'd need to know what everybody asks the model on their local GPU. You would literally have to start an international thought police and make sure that all ideas, conversations and types of intelligence are compliant. The risk of losing our freedoms is much more real than the existential risks purported by the extreme safety experts.
If anybody thinks they can actually lay out a realistic scenario, I'd be happy to debate them again.
To be clear, this is just a weekend post that got a bit too long. It is my personal opinion and not the official position of any company. We have a diverse set of opinions within Recursive on this subject. We certainly all agree that bad outcomes are possible with AI and we should continue our work on preventing those.