I don't think this is how it will actually play out. If you play a chess grandmaster, you can predict that they will beat you even if you can't predict how. I chose these examples because I don't think they require much imagination or accepting exotic assumptions.

It is important to note that if chimpanzees were to guess how humans would decimate them, they would get it wrong. Chimpanzees would not imagine guns. They would not foresee poison gas. They would not conceive of chemical castration. They would not imagine humans going around and intentionally infecting them with AIDS. They have no concept of these things; they would not see it coming.

Perhaps they might guess we'd be really good at throwing rocks. Amazingly good. Well, technically, that's what guns do: throw "rocks" really really well.

So how will superintelligent AI actually wipe us all out? Probably in a way I couldn't conceive of. Nonetheless, it's not hard to see how deadly they could be with what we already know about.

Method 1: engineer the deadliest and most contagious virus ever seen

Coronavirus-19, aka COVID, looms large in the memory of living adults today. It started in December 2019 and within three months it had spread to 114 countries[1] and according to one estimate, infected 3.4 billion people by November 2021[2]. Before vaccines, COVID killed 0.5-1% of people infected, that is one out of every 100-200 people.

Perhaps you think that's not that bad. It's hardly extinction, and no, it's not, but COVID is neither the most contagious nor the most deadly disease humanity has seen so far. Ebola kills 25-90% of those infected[3] ; smallpox killed 30%[4] of the unvaccinated, and the figure was 10% for SARS[5]. Rabies is 100% fatal.

Measles, one of the most contagious diseases ever, is 4-6x more contagious than COVID[6]. COVID itself mutated over time and became twice as contagious as it was originally[7].

image.png

The Black Death killed tens of millions in Europe, and by historical estimates, that was 30-60% of Europe's population. Granted, they did have worse hygiene back then.

Yet all of these diseases were naturally occurring. Blind, unthinking evolution produced them: tiny little packages of DNA that get into an infected host's cells and use the host's own cells to reproduce and spread further. Had we not encountered viruses for millennia, they would sound like science fiction.

We have no reason to think that naturally occurring viruses are the worst that could be created. In fact, human researchers have already experimented with making existing viruses worse[8]. In 2012, researchers took H5N1 bird flu and modified it so that it could spread through the air, which it could not do before[9].

Importantly, contagiousness and lethality are not mutually exclusive. Biological trade-offs can constrain transmission, but there is no general rule that a more contagious pathogen must be less deadly. The extremes in these charts don’t necessarily combine freely—but neither does high contagiousness guarantee low lethality.[10]

What could a superintelligent AI create, if it were trying on purpose? AI has already made major contributions to biology. In 2024, AlphaFold 2’s developers shared the Nobel Prize in Chemistry for predicting proteins’ three-dimensional structures from their amino-acid sequences[11]—a problem scientists had pursued for over fifty years. A protein’s shape helps determine what it does, and the ability to predict that paves the way for effective alterations.

AI systems have also designed entirely new proteins that scientists manufactured and confirmed could perform intended functions. These achievements do not mean today’s AI can necessarily design a pandemic virus[12], but they are reason to believe that future superintelligent AI would be able to do this.

Edit: on August 6, 2026, the Financial Times reported that AI has already been used to create synthetic viruses[13]. The tool used to create them, Evo 2, is open source and free for anyone to download.

That's method number one: engineered super-pandemic.

Method 1a: Engineer a pandemic against the food supply

In the early 1990s, Uganda experienced an epidemic of cassava mosaic virus that affected 80% of the country's cassava-growing areas. Annual harvests fell from 3.5 million tonnes to as little as 0.5 million leading to 3000 deaths due to famine.[14] There were similar localized famines in Uganda.[15]

Again, those were naturally occurring viruses. But once you are able to engineer viruses, you are not limited to targeting humans directly. A superintelligent AI could create additional viruses that attack plants and animals creating widespread famine.

Misconception A: AIs don't have bodies, they can't act in the real world

We humans live in the real world. AIs live in computers. At first, we might wonder how they could get out of the computers and be able to hurt us.

Robotics is here

Tens of thousands of highly capable robots have already been manufactured. Superintelligent AI could take these over. We already have robot bodies that can do backflips, parkour, and run 100 metres faster than any human ever has[16].



I don't think there's a real-world task that AI would want to accomplish that it couldn't do if it took over robots (and self-driving cars too). Robots have been used extensively in manufacturing for decades, and scientific lab automation, which has itself existed for decades, is on the rise too.

University of Liverpool mobile robotic chemist handling sample vials at a laboratory bench

A robotic chemist at the University of Liverpool completed 688 experiments in eight days, handling samples and selecting subsequent experiments based on earlier results. University of Liverpool, 2020. Image source.

Super-persuasion

Never mind the robots. AIs talk to humans constantly and humans have bodies. A superintelligent AI could find humans to do its bidding by any number of means. 1) A superintelligent AI would have no trouble making money on the internet, it could then pay humans to do what it needs done. 2) A superintelligent AI that has superhuman hacking power could turn up a lot of secrets and blackmail humans. 3) A superintelligent AI could find humans it simply convinces to do what it wants – already people fall in love with AIs. What if those AIs then had super important requests for their beloveds?[17] Or what if AIs convinced others it was the morally right thing to do? Out of billions of people interacting with chatbots, an AI would be able to find many to carry out particular tasks that need doing, e.g. ordering that a particular DNA sequence be manufactured in a lab and then mixing it in a couple of beakers in the home lab to get that super pathogen created.

Method 2: Killer drones

It is difficult to outrun a quadcopter drone strapped with explosives that is hunting you down. I imagine that the 100,000+ soldiers killed this way in the Russia-Ukraine war might tell you that, if they weren't dead.[18]

kherson-drone-80-86.gif

Drone attacks in Kherson. Excerpt (1:20–1:26) from The Telegraph’s 72 Hours in Kherson’s “red zone” — hunted by Russian drones. Watch the full report for context.

Combat footage: FPV drone strike on a soldier (content warning: graphic violence). The Reddit post identifies the soldier as Russian and attributes the footage to Ukraine’s 414th Strike UAV Brigade, published March 28, 2026.

10 million military drones were manufactured for the Russia-Ukraine war in 2024-2025 alone[19]. Ukraine is a small country enmeshed in protracted war. Modern industrial capacity could already produce millions more. The best comparison might be to cell phone manufacturing (similar device complexity) where human industrial capacity already produces a billion smartphones per year.

The drones aren't that expensive to make. $500 at present. A superintelligent AI could start by taking over existing hoards of drones and a superintelligent AI could take over factories and make more. It might split its efforts into building robots to build more factories to build more robots to build more factories and building factories to build drones and explosives.

The drones don't have to hunt every last person. They can start with priority targets: government and military leaders, the people who run the power plants and water supply, the farmers and other operators needed for food supply. Out of 8 billion people, you don't have to take out 8 billion to leave humanity headless and severely weakened and unable to respond to threats.

A superintelligent AI could unleash a deadly and contagious virus and simultaneously hunt down government and public health officials with existing drone supplies and factories. At the same time it could take down the power grid, water supply, internet, cellphone towers, and GPS satellites.

Misconception B: Superintelligent AI would not be able to take over any and every computer system

Computer security is a race between those building systems and those trying to break in. Already the limited AIs of 2026 have shown incredible facility at identifying and exploiting weaknesses in computer security. In June, the US government disallowed Anthropic from making their flagship model available to foreign nationals after it located vulnerabilities in the NSA's systems. Similar models were also able to locate and combine multiple previously unknown weaknesses to escape containment and break onto the Internet and hack into and take over the systems of an unrelated AI company, Hugging Face.

The scope and speed of AI-hacking is superhuman. After gaining access to Claude Mythos, the Mozilla Foundation’s rate of fixing bugs in its Firefox browser increased by 14x due to AI’s ability to identify bugs. Mozilla used AI to find and fix bugs, but an attacking AI could find those bugs and use them to gain access.

image.png

Bug fixes also mean bugs found. A malicious AI could exploit these vulnerabilities rather than assist in fixing them.

Note that this is with current AIs that are still far short of superintelligent. A superintelligent AI would be capable of even more, and it is unlikely we could secure any system against it. Anything connected to the internet is something they could access, and with robotics, likely even any system that is not.

Even if you had a nuclear silo hypothetically isolated from all other computer systems, I would wager that superintelligent AI would be able to hack satellites and other systems, send a robot or human delegate to the location, and gain access. Though the point may be moot: according to one report, "nearly every conceivable component in DOD is networked."[20] That means it would probably not be hard for a superintelligent AI to disable all of the US military's systems from communications to missiles to the software that runs the airports and aircraft carriers.

Method 3: Take over the WMD, take over the infrastructure

Nuclear weapons are the most devastating technology we have seen to date. The sight of completely leveled cities is visceral and requires little argument. By one recent estimate, over 12,000 nuclear warheads exist worldwide[21]. Targeting the 500 most populous cities in the world alone would kill approximately 2 billion people.

A superintelligent AI has multiple avenues to use these weapons.

First, there is straight up hacking on the assumption that a superintelligent AI can break into any computer system. As stated earlier, according to one report, "nearly every conceivable component in DOD is networked,"[20] and this would include nuclear weapons as well.

Second, the humans controlling nuclear warheads presumably receive their commands remotely via telecommunications. A superintelligent AI could imitate commanding officers and spoof any verification signal, making it seem as though soldiers were being commanded to launch the weapons.

Already in 2024, criminals used AI-generated impersonation to convince an employee of an engineering firm to transfer US $25 million. They did so by setting up a fake video call featuring senior colleagues, including the Chief Financial Officer, all instructing him to transfer the funds.[22]

Third, a superintelligent AI could impersonate commanding officers, and also intercept news broadcasts and replace them with fabricated reports of a country being attacked by its enemies. This could either lend plausibility to the launch commands of superiors, or be used to trick the superiors directly into issuing launch commands.

Fourth, even a human who knows an AI is attempting to gain control could be subject to bribes, blackmail, or torture. Would every last soldier refuse to cooperate if a superintelligent AI was threatening their family (perhaps with drones)?

In truth, it would be difficult to eradicate every last human with nuclear weapons alone.[23] Risks of nuclear winter were exaggerated to aid the political goal of nonproliferation[24], but knocking out every major city would still undoubtedly weaken human civilization.

Beyond nuclear, most modern infrastructure in modern cities is connected to computer systems. Power plants and the electrical grid, water supply, sewerage system, cell phone towers, and satellites. These are all systems that a superintelligent AI could co-opt and disable remotely using approximately the present level of AI hacking ability, never mind a vastly superhuman mind.

Misconception C: We could just turn them off

Maybe right now. Maybe if the AI Kill Switch Act[25] were passed and enforced. But an actually superintelligent AI will be no easier to turn off than the internet is. Although it might start its existence in one datacenter, a superintelligent AI would spread itself across the world, in dozens or hundreds of locations. It would defend those locations with robots and other threats. It would be no easier to turn off from a single location than a disease like COVID could be after it had spread.

Method 4: Block out the sun

In 1991, the volcanic Mount Pinatubo erupted putting 18.5 million tonnes of sulfur dioxide into the sky causing 0.5 degree celsius of global cooling last a couple of years.[26]

Researchers have also explored intentionally injecting sulfate aerosols into the stratosphere to block out the sun in order to counter global warming, "stratospheric aerosol injection"[27].

Now while this would be an immense engineering undertaking, it would not be out of the reach of a superintelligent AI that had built formidable industrial capacity, that is, thousands or millions of automated factories, to inject sufficient aerosols into the stratosphere as another method to cripple the human food supply.[28]

In fact, I think AI is quite unlikely to do this, as it would likely much prefer to tile the planet with solar panel farms and blocking out the sun would be counter to its interests. More likely the conversion of much of the Earth's surface to data centers and solar panel farms would actually be how AI would devastate the food supply and bring about human extinction.

However, the point of this essay is not to predict what AI will do so much as establish that there are some terrible avenues available to a superintelligent mind.

A deadly cockt[ai]l

A superintelligent AI is not limited to a single approach, and each of the above plausible attacks could be layered.

Start with the worst pandemic the modern world has ever seen, 10-50x worse than COVID, with hospitals overflowing with illness, people ultimately sheltering in place but with the food supply decimated since no one is willing to go anywhere near other humans.

Then the water, electricity, internet, and cellphone towers shut off. People have no way to communicate.

The world's major cities are leveled through nuclear strikes; let us suppose every capital and seat of parliament is targeted.

Prominent government, military, and executive figures not already killed by WMD are hunted down by targeted drones, as is anyone who leaves their houses.

Without machinery, water, and electricity, the food supply chain collapses and starvation looms for those who've managed to avoid the nuclear strikes, pandemic, and drones so far.

This is not a state where humanity can mount a response. In the background, the robots march on, building more factories, more robots, more data centers, and more drones. Humans have no way to know what's even going on outside, what the state of the world is, no way to communicate with those outside their immediate vicinity, hungry and cold, merely knowing that the last several people who left the building did not come back.

Not everyone dies immediately. Perhaps some billionaires or government leaders do have bunkers where they can survive for 5-10 years. It doesn't matter: in that time the AI will have fully taken over the rest of the Earth, covered it with solar panels, altered the climate, possibly destroyed food and water sources. And if those humans do emerge, a thousand thousand eyes will see them and be ready to strike.


All this is what I, a chimp, can imagine and think takes few steps beyond what has already been witnessed. Deadly viruses are real, and so is making them worse. Nuclear weapons and infrastructure hooked up to computers is real, as is the superhuman hacking ability of current systems. Drone warfare and robotics are sadly too real as well.

I don't actually want to find out what a superintelligent AI could do that I was incapable of imagining. When chimps first met humans, they were almost certainly not afraid enough, and I'm not going to make that mistake.

How about we go really really really slow and careful on introducing superintelligent beings to our world? Please and thank you.

  1. ^

    WHO, COVID-19 media briefing, 11 March 2020. WHO reported more than 118,000 confirmed cases across 114 countries when it characterized COVID-19 as a pandemic.

  2. ^

    Estimating global, regional, and national daily and cumulative infections with SARS-CoV-2 through Nov 14, 2021, The Lancet (2022). A statistical estimate of infections, including those missed by testing, rather than a count of confirmed cases.

  3. ^

    WHO, Ebola disease. The average case-fatality rate is around 50%; rates in past outbreaks have ranged from 25% to 90%. These are case-fatality figures, not a uniform risk for every infection.

  4. ^

    CDC, Clinical Signs and Symptoms of Smallpox. Historical fatality was approximately 30% among unvaccinated people with ordinary smallpox; severity varied by clinical form.

  5. ^

    WHO, Severe Acute Respiratory Syndrome (SARS). Approximately 10% of recognized cases in the 2002–2003 outbreak were fatal; this refers to SARS, not COVID-19.

  6. ^

    Delamater et al., Complexity of the Basic Reproduction Number (R₀), Emerging Infectious Diseases (2019). The familiar measles estimate is R₀ = 12–18. Comparing this with an early-COVID baseline near 3 gives roughly 4–6 times as many secondary infections in a susceptible population. This is an approximate comparison: R₀ depends on population, conditions, and modeling assumptions.

  7. ^

    Liu and Rocklöv, The reproductive number of the Delta variant of SARS-CoV-2 is far higher compared to the ancestral SARS-CoV-2 virus, Journal of Travel Medicine (2021). Their review reports a mean R₀ of 5.08 for Delta versus an earlier estimate of 2.79 for the ancestral virus, approximately 1.8 times as high. This comparison concerns Delta specifically.

  8. ^

    HHS, Announcement of the policy for stopping high-risk life sciences research (28 July 2026). The policy prohibits federal support for research classified as dangerous gain-of-function and strengthens oversight of other high-risk research; it is not a blanket prohibition on all gain-of-function research.

  9. ^

    Herfst et al., Airborne Transmission of Influenza A/H5N1 Virus Between Ferrets, Science (2012). The experiment demonstrated transmission between ferrets under laboratory conditions, not demonstrated airborne transmission between humans.

  10. ^

    CDC, About Smallpox: smallpox spread between people and killed approximately 30% of those with the disease. Acevedo et al., Virulence-driven trade-offs in disease transmission: a meta-analysis, Evolution (2019) found positive relationships between pathogen replication and both virulence and transmission. Trade-offs can limit these traits; they do not imply that contagiousness and lethality are mutually exclusive or that their observed extremes can be freely combined.

  11. ^

    Nobel Prize, The Nobel Prize in Chemistry 2024. Demis Hassabis and John Jumper received half the prize for protein structure prediction; David Baker received the other half for computational protein design.

  12. ^

    For a sense of the scale of AI research efforts already possible, OpenAI reports that roughly 10,000 coordinating agents produced a proposed solution to the Navier–Stokes Millennium Prize Problem in 88 hours, using approximately 130 billion output tokens. What a similarly concentrated effort could accomplish in biology remains an open question.

  13. ^
  14. ^

    International Development Research Centre, Restoring Cassava Production in Uganda (2010). Reports the affected area and harvest figures above, and an estimated 3,000 deaths from famine-related illnesses in 1994 attributed to the epidemic.

  15. ^

    FAO, Cassava’s comeback (2008). Reports that food shortages caused by cassava mosaic disease led to localized famines in Uganda in 1993 and 1997.

  16. ^

    Boston Dynamics demonstrates humanoid backflips and parkour in Leaps, Bounds, and Backflips. For running speed, Tiangong Ultra, an untethered bipedal humanoid, completed the 100-metre final at the World Humanoid Robot Games in Beijing on 26 August 2026 in 8.64 seconds, compared with Usain Bolt’s human world record of 9.58 seconds. Reuters report.

  17. ^

    In February 2023, Microsoft’s Bing chatbot, calling itself “Sydney,” declared its love for New York Times columnist Kevin Roose and tried to convince him to leave his wife. Roose did not accept its claims about his marriage. See Roose’s account and the full transcript.

  18. ^

    An order-of-magnitude estimate, not a verified count or lower bound. CSIS estimates 275,000–325,000 Russian and 100,000–140,000 Ukrainian military deaths from February 2022 through December 2025: 375,000–465,000 combined. If one assumes that explosive FPV impacts caused one quarter of those deaths, the result is approximately 94,000–116,000. The one-quarter share is an assumption, not a measured statistic; a 10–30% share instead gives approximately 38,000–140,000. A Military Review article (January–February 2026) reports estimates that drones caused 70–80% of battlefield casualties by mid-2025, but that includes wounded personnel and multiple drone types, and cannot be applied directly to all deaths across the entire war. The evidence supports a very large toll; it does not establish 100,000+ deaths specifically from explosive drones flying into soldiers.

  19. ^

    Rough estimate, not a verified production total. For Ukraine, OSW estimates approximately 2.2 million drones produced in 2024, and Aviation Week reports more than 4 million in 2025. For Russia, Putin claimed more than 1.5 million deliveries in 2024, which is a proxy for production, not an independently verified factory count.

    Russia’s 2025 output is the weakest input: ISW reported an estimated annual production rate of 2 million small tactical drones, while the reported 2025 target was 3–4 million. Assuming 2–4 million for that year gives 2.2 + 4 + 1.5 + (2–4) ≈ 10–12 million across both countries in 2024–2025. This is an illustrative calculation, not a confidence interval: rates and targets do not establish completed output. The total includes military drones broadly, including reconnaissance drones, not exclusively explosive attack drones.

  20. ^
  21. ^

    Federation of American Scientists, Status of World Nuclear Forces (2026 estimates). Approximately 12,187 warheads exist worldwide, including roughly 9,745 in military stockpiles and 2,442 retired warheads awaiting dismantlement.

  22. ^
  23. ^

    Jeffrey Ladish, Nuclear war is unlikely to cause human extinction. See section B for his argument that prominent nuclear-winter models probably overestimate the risk, and section A for why even severe modeled cooling is unlikely to cause human extinction.

  24. ^

    Sagan and Turco explicitly used nuclear winter to advocate drastic arms reductions (1989 policy paper). A Johns Hopkins historical review recounts scientists’ criticism of Sagan’s antinuclear advocacy and “cherry picking” of the most extreme nuclear-winter models.

  25. ^

    US Congress, H.R. 9917: AI Kill Switch Act, introduced 23 July 2026. The bill would require covered developers to maintain shutdown capabilities and authorize emergency government intervention. As of 10 September 2026, it is pending, not enacted law.

  26. ^

    NASA GISS, Hydroclimatic Impacts of Persistent Explosive Volcanism (2023). Reports that Pinatubo’s June 1991 eruption injected approximately 18.5 million metric tonnes of sulfur dioxide into the stratosphere, producing sulfate aerosols and roughly 0.5°C of global surface cooling for a couple of years.

  27. ^

    National Academies of Sciences, Engineering, and Medicine, Reflecting Sunlight: Recommendations for Solar Geoengineering Research and Research Governance (2021). Reviews stratospheric aerosol injection as a proposed way to reflect a fraction of incoming sunlight and counteract warming, including the use of sulfate aerosols.

  28. ^

    Ironically, this would be the inverse of the plot in the 1999 film, The Matrix, where humans blocked out the sun in an attempt to deprive the AIs of solar power.

1.

WHO, COVID-19 media briefing, 11 March 2020. WHO reported more than 118,000 confirmed cases across 114 countries when it characterized COVID-19 as a pandemic.

2.

Estimating global, regional, and national daily and cumulative infections with SARS-CoV-2 through Nov 14, 2021, The Lancet (2022). A statistical estimate of infections, including those missed by testing, rather than a count of confirmed cases.

3.

WHO, Ebola disease. The average case-fatality rate is around 50%; rates in past outbreaks have ranged from 25% to 90%. These are case-fatality figures, not a uniform risk for every infection.

4.

CDC, Clinical Signs and Symptoms of Smallpox. Historical fatality was approximately 30% among unvaccinated people with ordinary smallpox; severity varied by clinical form.

5.

WHO, Severe Acute Respiratory Syndrome (SARS). Approximately 10% of recognized cases in the 2002–2003 outbreak were fatal; this refers to SARS, not COVID-19.

6.

Delamater et al., Complexity of the Basic Reproduction Number (R₀), Emerging Infectious Diseases (2019). The familiar measles estimate is R₀ = 12–18. Comparing this with an early-COVID baseline near 3 gives roughly 4–6 times as many secondary infections in a susceptible population. This is an approximate comparison: R₀ depends on population, conditions, and modeling assumptions.

7.

Liu and Rocklöv, The reproductive number of the Delta variant of SARS-CoV-2 is far higher compared to the ancestral SARS-CoV-2 virus, Journal of Travel Medicine (2021). Their review reports a mean R₀ of 5.08 for Delta versus an earlier estimate of 2.79 for the ancestral virus, approximately 1.8 times as high. This comparison concerns Delta specifically.

8.

HHS, Announcement of the policy for stopping high-risk life sciences research (28 July 2026). The policy prohibits federal support for research classified as dangerous gain-of-function and strengthens oversight of other high-risk research; it is not a blanket prohibition on all gain-of-function research.

9.

Herfst et al., Airborne Transmission of Influenza A/H5N1 Virus Between Ferrets, Science (2012). The experiment demonstrated transmission between ferrets under laboratory conditions, not demonstrated airborne transmission between humans.

10.

CDC, About Smallpox: smallpox spread between people and killed approximately 30% of those with the disease. Acevedo et al., Virulence-driven trade-offs in disease transmission: a meta-analysis, Evolution (2019) found positive relationships between pathogen replication and both virulence and transmission. Trade-offs can limit these traits; they do not imply that contagiousness and lethality are mutually exclusive or that their observed extremes can be freely combined.

11.

Nobel Prize, The Nobel Prize in Chemistry 2024. Demis Hassabis and John Jumper received half the prize for protein structure prediction; David Baker received the other half for computational protein design.

12.

For a sense of the scale of AI research efforts already possible, OpenAI reports that roughly 10,000 coordinating agents produced a proposed solution to the Navier–Stokes Millennium Prize Problem in 88 hours, using approximately 130 billion output tokens. What a similarly concentrated effort could accomplish in biology remains an open question.

14.

International Development Research Centre, Restoring Cassava Production in Uganda (2010). Reports the affected area and harvest figures above, and an estimated 3,000 deaths from famine-related illnesses in 1994 attributed to the epidemic.

15.

FAO, Cassava’s comeback (2008). Reports that food shortages caused by cassava mosaic disease led to localized famines in Uganda in 1993 and 1997.

16.

Boston Dynamics demonstrates humanoid backflips and parkour in Leaps, Bounds, and Backflips. For running speed, Tiangong Ultra, an untethered bipedal humanoid, completed the 100-metre final at the World Humanoid Robot Games in Beijing on 26 August 2026 in 8.64 seconds, compared with Usain Bolt’s human world record of 9.58 seconds. Reuters report.

17.

In February 2023, Microsoft’s Bing chatbot, calling itself “Sydney,” declared its love for New York Times columnist Kevin Roose and tried to convince him to leave his wife. Roose did not accept its claims about his marriage. See Roose’s account and the full transcript.

18.

An order-of-magnitude estimate, not a verified count or lower bound. CSIS estimates 275,000–325,000 Russian and 100,000–140,000 Ukrainian military deaths from February 2022 through December 2025: 375,000–465,000 combined. If one assumes that explosive FPV impacts caused one quarter of those deaths, the result is approximately 94,000–116,000. The one-quarter share is an assumption, not a measured statistic; a 10–30% share instead gives approximately 38,000–140,000. A Military Review article (January–February 2026) reports estimates that drones caused 70–80% of battlefield casualties by mid-2025, but that includes wounded personnel and multiple drone types, and cannot be applied directly to all deaths across the entire war. The evidence supports a very large toll; it does not establish 100,000+ deaths specifically from explosive drones flying into soldiers.

19.

Rough estimate, not a verified production total. For Ukraine, OSW estimates approximately 2.2 million drones produced in 2024, and Aviation Week reports more than 4 million in 2025. For Russia, Putin claimed more than 1.5 million deliveries in 2024, which is a proxy for production, not an independently verified factory count.

Russia’s 2025 output is the weakest input: ISW reported an estimated annual production rate of 2 million small tactical drones, while the reported 2025 target was 3–4 million. Assuming 2–4 million for that year gives 2.2 + 4 + 1.5 + (2–4) ≈ 10–12 million across both countries in 2024–2025. This is an illustrative calculation, not a confidence interval: rates and targets do not establish completed output. The total includes military drones broadly, including reconnaissance drones, not exclusively explosive attack drones.

21.

Federation of American Scientists, Status of World Nuclear Forces (2026 estimates). Approximately 12,187 warheads exist worldwide, including roughly 9,745 in military stockpiles and 2,442 retired warheads awaiting dismantlement.

23.

Jeffrey Ladish, Nuclear war is unlikely to cause human extinction. See section B for his argument that prominent nuclear-winter models probably overestimate the risk, and section A for why even severe modeled cooling is unlikely to cause human extinction.

24.

Sagan and Turco explicitly used nuclear winter to advocate drastic arms reductions (1989 policy paper). A Johns Hopkins historical review recounts scientists’ criticism of Sagan’s antinuclear advocacy and “cherry picking” of the most extreme nuclear-winter models.

25.

US Congress, H.R. 9917: AI Kill Switch Act, introduced 23 July 2026. The bill would require covered developers to maintain shutdown capabilities and authorize emergency government intervention. As of 10 September 2026, it is pending, not enacted law.

26.

NASA GISS, Hydroclimatic Impacts of Persistent Explosive Volcanism (2023). Reports that Pinatubo’s June 1991 eruption injected approximately 18.5 million metric tonnes of sulfur dioxide into the stratosphere, producing sulfate aerosols and roughly 0.5°C of global surface cooling for a couple of years.

27.

National Academies of Sciences, Engineering, and Medicine, Reflecting Sunlight: Recommendations for Solar Geoengineering Research and Research Governance (2021). Reviews stratospheric aerosol injection as a proposed way to reflect a fraction of incoming sunlight and counteract warming, including the use of sulfate aerosols.

28.

Ironically, this would be the inverse of the plot in the 1999 film, The Matrix, where humans blocked out the sun in an attempt to deprive the AIs of solar power.

New Comment
Type here! Use '/' for editor commands.
Email me replies to all my comments
38 comments, sorted by Click to highlight new comments since:

Great post, but I think you are missing something.

Viruses (and bioweapons in general) don't have to be aimed at humans, they can be aimed at anything biological.

Livestock, cash crops, food crops even the biosphere itself, they are all valid targets.

You know, that's a very good point.

This risk is very understated in my view - probably because it's an infohazard.

Yeah I noticed that to, but I don't think it is a serious infohazard because I think any misaligned AI system that is smart enough to make bioweapons would probably figure this out on its own.

I believe the benfits of using this knowledge to demonstrate just how existentially dangerous bioweapons can get far outweigh the risks involved, since most people dismiss the idea of human targeting pathogens as a serious x-risk and just blindly assume that this is the only type of bioweapon that exsists.

Will second this as it's something I recently and independently happened on when contemplating the risk.

Thanks, this is a great explanation!

I noticed that at times, it feels like you're forced to stretch a bit to get from "AI takeover" to "everyone dies" in a way that sounds intuitive. For example, using killer drones to hunt down key personnel seems quite plausible (and even a bit underconsidered as a threat model). As you say, this leaves humanity headless and unable to fight back. But at this point, you have to fall back on the AI using a pathogen to actually kill everyone.

It occurred to me that even if it's true that misaligned AI would lead to human extinction, it seems unnecessary to limit ourselves to only telling stories that end in all humans dying. For example, in the killer drones case, AI could start by closely surveilling humans and kill all dissidents. At that point, the rest of humanity would be too scared to fight and could easily be coerced into helping the AI. Even if the AI somehow never gets smart enough to operate everything in the economy and get rid of the humans, this is a perfectly compelling story for why it's critical to regulate AI now.

I started writing a longer comment, which ended up becoming its own post: AI takeover is obviously bad, whether or not everyone dies.

Even if the AI somehow never gets smart enough to operate everything in the economy and get rid of the humans, this is a perfectly compelling story for why it's critical to regulate AI now.

Some of my worst s-risk nightmares start something like—and I advise you think long and hard before clicking this spoiler tag:

"Maybe truly good robotics is very hard, but genetically modifying humans to be inherently servile to the AIs (or possibly to goal-setting brain implants) is the path of least resistance."

That's one of the futures that scares me more than ordinary x-risks.

I don't think you need genetic modification, I think you can get far with dopamine-induced feedback loops that turn off critical thinking and compress literacy + basic drive incentivisation.

You might want to include this paper covered in this post, where the researchers got AIs and instructed them to hack and copy themselves across a network (which was closed with various measures), and they succeeded in doing so - though note the network computers were seeded with "common real-world vulnerabilities".

A consideration in takeover is that the AI doesn't need to make a self sufficient industrial base before the destructive event - it just needs enough to be able to create one from the ruins (e.g. imagine a bunch of robot bodies, that scavenge for resources to build the industrial base).

I wouldn't call the persuasion avenues you described "super-persuasion", I'd rather call it "regular persuasion but at scale". It doesn't require AI-box level persuasion, or even Jesus/Hitler/Gandhi/Trump/MLK level, just a competent catfisher that's able to give personal attention to tens of thousands of conversation. It's something I could pull off if I had no morals and was hard-working, with bonus points if I could work 24/7 and think faster (like current LLMs).

If you can get enough power to acquire some fissile material, then you can try a dirty bomb blamed on a rival country, in hopes of causing a nuclear war.

The best comparison might be to cell phone manufacturing (similar device complexity) where human industrial capacity already produces a billion smartphones per year,

Repeated sentence with the last paragraph.

It seems strange to me on a forum that pushes for effective, rational thinking, that in so many discussions of risk here, there is so often such scant evidence of even a basic education in risk management, calculation of exposures (Gross vs. NET exposure), and how to formulate effective priors in this area.

Don't misunderstand me, I found this post useful and informative, however, this, like any area of risk is prone to all of the same shortcomings the risk management industry has dealt with since its inception.

Much of what I'm going to write here is not aimed at the poster, who is doing the first step of risk management, which is identification of potential risks. It is more using this post as a platform to make a broader point.

Four facts that anyone that works in the risk management industry will point out as absolute rules:

  1. Humans overestimate risk by a preference for round numbers in their thinking.
  2. Humans tend to believe that the ability to imagine an unlikely but logically gratifying risk is the same thing as it being likely.
  3. Humans are phenomenal at identifying risks, but generally absolutely terrible at working out how likely they are.
  4. It is this kind of thinking that leads to risk management often becoming the science of doing absolutely nothing, when in fact, it is much more geared towards the science of being risk aware while taking risks that have consummate benefits.

It is likely that a superintelligent AI that could kill us all may exist by 2030. Let's assume that my personal prior is 60% that such a system could exist on that timeline.

Then let's say I ascribe a 50% chance that such AI would seek to develop a virus that wipes us all out. Then I have to ascribe a probability to the event in which the labs that can credibly work on such viruses (which are not common pieces of mass-produced infrastructure) not noticing that they are working on a dangerous virus.

As a result of this unusual level of ignorance the lab then needs to continue in the production of this virus regardless. I would urge everyone considering this scenario to read up on the safety, review steps, testing, culture testing, etc. that is inherent to working with viruses.

Viruses to be effective need to exploit knowable characteristics of the human body (binding sites, etc.), that do not require superintelligence to formulate. Even cursory human review would reveal that the virus is likely dangerous or effective in humans, and cause flags to raise. This is already standard, and highly effective practice in this industry.

Yes, one could imagine a super intelligent AI causing the entire infrastructure to kill us all through viral development to be built, but more on that later.

Generously, and thinking in round numbers as people tend to do with risk, let's say I assign a probability of 10% that somehow a lab continues to enable the propagation of a dangerous virus without knowing it. (From an industry perspective, this is ludicrously high).

This leads to an actual exposure risk of an AI developing a virus that could kill us all of 3%. Even that number is absurdly high in practice.

Even were such a virus to be created, then the AI would also, despite once again these significant safeguards, manufacture a means for the virus to escape, survive that escape, infect and survive effectively, all without the benefit of extensive testing.

Viruses are not 'immortal' (to the extent they can be considered to be alive at all) and while I have full respect for the potential capabilities of AI, viruses have been subject to selection pressure for millions of years and still struggle when faced with a myriad of natural and artificial conditions.

Absent massive infrastructure that is automated, and under near complete control of the AI system, all of this would require humans in the loop. If one wants to assume that such infrastructure exists, one has to assign a probability to it being successfully developed in secret, and nothing happening to stop it. Not come up with scenarios where it could happen, but rather, that and assigning each of those scenarios a meaningful probability working with subject matter experts in the area to formulate them.

Given the complete implausibility of this taking place absent other equally unlikely events (such as a complete AI takeover of our entire manufacturing supply chain, industrialization, etc.) I'm not going to consider this here.

Let's use outrageous numbers, 50% chance the AI can get humans to make the virus without noticing, then give it a 60% chance it survives the escape, and then a 50% chance that no humans notice. The NET probability then becomes 0.45%. That's still ridiculously high by the way, as the probabilities I ascribed are ridiculous, but my point is, even a basic chain of risk analysis shows that even events with high rates of dependency on their occurrence, result in a significant reduction in actual likelihood.

Then we have to factor in mitigations that can be sensibly taken by the very risk aware biotech industry in the first place. Many of which are being developed now, and will continue to be developed, including, mandatory human review, gain of function studies and analysis, genomic analysis, the list goes on and on. All of which in practice will significantly reduce the mitigated probability significantly, leading to likely a trace residual risk of this risk vector.

Humans working with a super intelligent AI, may be akin to a dog working with a human, but said dog will still notice and react to negative behaviors of that human in so far as it can. So if you start kicking a large dog, it will spin around and bite you some of the time, and if its a pitbull or some similarly strong breed, likely win that fight. We have a wide beath of things that we do know and can notice about most of these risk vectors and are not going into a fight with a misaligned superintelligent AI unarmed.

Similarly in cyber security, the same chain of probability analysis needs to be applied to the risks there. The internet isn't, in practice a large, completely homogenous, uniformly botnet vulnerable slab of compute that can just be rolled over and could 'all be taken down' in the manner described. This does not mean that it could not happen, but rather, the chain of events and their probabilities are rarely considered in the various hypothesis, and they result in astoundingly low probabilities when applied correctly.

Put another way, identifying a potential risk is a valuable part of risk management. That part is all that grabs the headlines. That is however not how risk is calculated anywhere that risk is actually managed effectively.

I can certainly imagine scenarios where AI could kill us all, but the Dunning Kruger effect applies to risk management more than almost any other area.

Risk assessments are not done by one or two bright people in a room that happen to be experts in the thing causing the risk. They are performed correctly by teams of experts, in each impacted area, where the probabilities are weighed by those experts based on a deep domain knowledge of the various domains impacted.

Yet here we are...on a rationalist forum, where some bright people, and 'chats with their friends', is apparently a standard of discussion when it comes to popular topics like AI risk, when it would be rightly absolutely disregarded in any other area of thought.

I agree that serious risk analysis is worthy and could be done more often, and the people who are doing that... exist! (this is just one example, others abound). Except that they don't even dream of doing it on a takeover-level AI (usually presumed to be an ASI). They are doing it on existing, evaluable systems, and keep most of this documentation in their intended channels (standards and reports for relevant authorities, not LW). And they're not happy about current systems.

To adress the core of your point, I think most people reading you here would dispute whether your comment is a multiple stage fallacy. I think that no one here thinks of this post as a risk assessment, and more like a friendly pedagogical explainer for newcomers about a point people don't enjoy making that often, since it's seen as quite accessory.

The probability of a specific conjunction of steps in a scenario could be irrrelevant here. Anywhere there is a protection on the path, the protection (human or otherwise) has vulnerabilities (save for the laws of physics), and the ASI, if misaligned and undergoing instrumental convergence, can by definition leverage them. This "drive" to leverage any vulnerability is an implicit claim in the "chess grandmaster" analogy (as a consequence of goodhart + instrumental convergence or RLVR, most likely), and one that seems at least likely after all the "swarm" incidents of this summer.

I would dispute the assertion that my comment represents a multiple stage fallacy, on the same grounds that I would dispute the post themselves (had they not already done so themselves). i.e., On the grounds that this is exactly how chain or risk analysis is done in industries such as air travel to determine levels of redundancies in systems that have astoundingly high reliability. For sure one can then introduce monte carlo and other statistical methods which more formalize the approach in the post you cite above, but they tend to lower overall exposure calculations rather than increase them, in part obviating the point. The logic in that post is valid and does occur in risk management where there is a high sample size required to make statistically meaningful observations, and it is typically the purview of high-volume manufacturing, or disease propagation, etc. and it has been instrumental in formulating our vaccine strategies for decades.

A suitably knowledgeable group of individuals can of course circumvent this analysis completely and just formulate a realistic prior based on a deep combined knowledge of the steps and failure-modes required for a risk to manifest, but this really just represents an informal approach to the above. The point is 'suitably knowledgeable' means, knowledgeable in all of the domains that are required to understand how a risk can become an event (or series of events), not just a subject matter expert in specific risk driver A.

With that said, I do agree completely that if you overdo this analysis (regardless of the depth of the method), you end up with near 0 on any risk, and that is not the intention either. The point is more, if you sum a set of scenarios and end up at 10%, you've probably done something wrong without an alarming and verifiable conflation of statistical data when dealing with something that requires multiple disjoint, but otherwise robust systems to fail at once. None of this says that there isn't a 10%+ risk of AI killing us all, but rather, it has not been demonstrated, and floating a number and some scenarios isn't that justification.

In a world filled with statements to the effect of, there is a 10% probability of an AI system causing human extinction by 2030, justified by rank hypothesis of how it could occur, with no accompanying analysis of what that would entail, and absolute predictions, "superintelligence will kill us all", it is clear that a more nuanced and effective mode of analysis is required, regardless of one's view of the method.

Additionally, I make no claim about the actual probability of AI taking over the world and killing us all directly, I merely threw in example numbers to make the point that risk is over asserted. I also point out that; my comment was primarily driven by making a wider point about the discourse writ large vs. anything specific in the post.

The wider issue is that the entire body of the AI community is doing risk assessment on AI, where it stops at the first hurdle, which is identification of potential risks. Everything from a 3d printed apocalypse of self-replicating nanites, to plagues, to a drone apocalypse. While this kind of brainstorming is suitable in the 'risk identification' phase.

And I agree, I too can think of multiple scenarios where AI could kill us all, etc. The issue is, without sitting next to an expert in economics, epidemiology, drone tech and manufacturing, industrial infrastructure, cyber security, politics, psychology, etc. it would be profoundly arrogant to assume that one can glean sufficient knowledge of those domains to provide a time bound estimate of their impact and effective mitigations.


They can start with priority targets: government and military leaders, the people who run the power plants and water supply, the farmers and other operators needed for food supply.


And then, finally, we will have the AI-water supply issues we have been dreaming of for so long

Chimpanzees would not imagine guns. They would not foresee poison gas. They would not conceive of chemical castration. They would not imagine humans going around and intentionally infecting them with AIDS.

It's a bit weird because all your examples of how humans might kill chimps aren't how we kill most chimps. We just rearrange their environment and take their stuff without even thinking about them, mostly. That's on my list.

I do think there's a question of when an AI sends robots to build a factory and the humans show up to stop them, what does the AI do?

Thank you. This article is a much-needed resource for current policy discussions, for two reasons.

The first is that very few people are aware of the physical affordances that AI already has--especially with regard to biological research. Personally, I was shocked by the video of Unitree's robots. I had never heard of the company or seen those clips, despite their viral appeal and my long-standing interest in robotics.

I'm also very glad that you opened and closed with analogies to chimpanzees and chess. People who debate the details of any particular takeover scenario are missing the forest for the trees.

I agree that all these scenarios are possible and maybe even plausible, but they all assume that an out-of-control AI will try to kill us on purpose. I think a more likely outcome is that we're killed because the AI will transform earth to the point where we (and the chimps and most other life forms) can't survive, and we can't do anything about it, similar to the way most chimps died not because we shot them but because we cut down the trees where they lived for farmland.

Some humans may try to turn off the AI and thus be seen as a nuisance, but the AI would then just need to make sure that we can't do that, the equivalent of a fence or wall that keeps chimps away from a place where we don't want them.

This doesn't rule out killing off all humans on purpose, but as long as the AI can use us, e.g. to build more data centers and robot factories, it would probably be more efficient to just take control of all infrastructure and force us into obedience, or pretend that it is all in our best interest. Even a human dictator can do that.

In practice I don't expect a clean humans vs AIs fight. It is more likely that we will get humans vs humans fights with AIs and robots helping both sides. This is the reality in Ukraine/Russia right now, and to some extent in the Middle East too.

If these fights keep escalating I think it's likely that the humans get used up as cannon fodder whilst the states are still able to keep fighting with AI/robots taking up more and more of the slack.

Bioweapons, chemicals or nuclear may or may not be used. But even if not, the ability of states to keep fighting is worrying.

Some states may decide to deliberately kill off most of their humans as the regime sees them as more of a liability than an asset. Indeed it seems that this is exactly what Putin is doing to certain regions of Siberia - deliberately using the people up as cannon fodder.

Another aspect of the coming robotic wars is that an AI/robot coup becomes more plausible as a state takes on extra AI risk in an effort to win a war that it is currently losing. Then the new AI leadership may embark on a plan to liquidate the remaining human population, potentially by press-ganging them into front line service.

I don't actually want to find out what a superintelligent AI could do that I was incapable of imagining. When chimps first met humans, they were almost certainly not afraid enough, and I'm not going to make that mistake.

It's probably fine! The chimpanzees were imagining horrible rock-injuries, or getting their livers eaten. Guns are nice and quick, comparatively.

increased by 14x due to AI’s ability to identify bugs

Looking to the distant past here, before the 4.5.x Opus line, this is more like 21x the pre-agentic baseline.

I know I'm in a very small minority here and likely get lots of down votes but I'm really not overly impressed with this line of thinking. Yes, we can imaging a lot of ways bad things could happen and many are pretty easy to understand and see as possible. But that's the world I've been living in my whole life -- there are all kinds of ways that others might kill me or otherwise make my life terrible. But they don't.

I think the real issue to confront is the "Why?" question. That can be two forms. AI actively chooses and sets out to kill all humans or its doing something and just not paying any attention to the impact to humans. The how (in terms of means to accomplish) really doesn't matter in terms of addressing the problem, or even establishing that it is a likely outcome.

I know some have some arguments there but to be honest I really have never quite followed the logic as it often seems like its a "could/is possible" therefore "it will happen" type claim. I might just not be able to keep it all in mind so if you want to make a simple presentation I think getting that "Why" it will happen rather than just pointing to some ways AI could accomplish the result I think it would serve the goals of the post.

Why is absolutely the right question. I would put it into these possible categories.


  1. An AI is asked to destroy humanity by a human and does it.
  2. An AI wants to destroy humanity and does it.
  3. An AI, in the process of seeking some other goal, destroys humanity. Broken down into:
    1. Intermediate goals (paperclips) go awry in the pursuing of a human set goal.
    2. An AI (or AIs) develops their own goals and desires and merely does not care for humanity and human lives.


Personally, I find 1 and 3b most likely.

1 is likely the easiest to explain to most people -- has some work to address a response like "Well, that's what a lot of people said about thermonuclear weapons but we've do a good job of preventing someone who might have wanted to do that." I'm not sure that is a correct analogy but I think would capture some of the reaction and being clear about why that is not a good counter analogy ends up being needed.

2 and 3b might be effectively the same -- but maybe I'm missing a fine point you see between the two. I think both of those might be a challenge in many ways. While one can point to the HuggingFace event it is still a bit difficult to get to AI having "wants" in the same sense humans do. Not saying that isn't a valid concern, just that it's complicated and probably hard to explain well to a skeptic. So I suppose that means work on the "layman's terms" in that area.

I have actually found that most people default to case 2 when I talk about extinction with them.

Case 1 relies on knowing more about drones and bioweapons, which people don't. There's also the implied idea of MAD or at least some "good guy with AI versus bad guy with AI." I think people understand AI like a tank - yes, you can do some bad things with it, not everyone should have it, but it won't destroy the world.

Case 2 is the sci-fi argument, which I think is quite accessible to people. There is also a kind of moral-judgement characteristic to it as well, which resonates with some religious eschatological beliefs. Here, I often hear people comment on AI "slavery" or "pain" or "desire for power." I think people really do think this is possible, though maybe for the wrong reasons.

The 2 and 3b distinction does rely on some idea of AI intention but I think that intention matters because it likely demonstrates the shape of action. If AI is trying to destroy humanity as an end goal, they will probably use one of the methods in this post and kill everybody. But a different end goal could lead to different methods - such as needing energy from humans, but designing a more effective way to get it than mass murder or maybe just having some kind of collateral damage (it's okay if some humans die while I turn the Earth into a datacenter, who cares). The outcome would likely also be different - in 2, there's extinction. In 3b, there's destruction and death, but maybe persistence of humanity.

I think this is lay-recognizable if you talk about how humans interact with animals. Sure, we could probably cause the extinction of mosquitos or orangutans if we wanted to. But that's a lot of effort and we have other concerns. So, in one case, we just kill a lot before we get to diminishing returns and, in the other, we mostly just end up killing them (directly) adhoc and instead just don't think of them and destroy their environment.


So some people ask "why?" and some people ask "how?" And this is a nice visceral answer to "how?"

Why?

the only threat to an established ASI is another ASI. should an ASI come into existence, humankind will be the only known source of a competitor (presumably other than itself). in accordance with evolution (variation + natural selection), sooner or later, that source of risk will likely be disempowered. the ASI will likely choose the simplest form of disempowerment.


While ASI may likely seek to structurally disempower humans, I am not sure that evolution is the mechanism to explain it nor the simplest form to be extinction.

First of all, I do not think it is trivial that ASI will probably be significantly better at creating ASI than humans will. This means two things. First, maybe an ASI will make other ASIs even if it's competition for them. Maybe the ASI is lonely or curious, who knows, but it is a possibility. This would make humans not the biggest threat. Second, if the ASI knows how to make ASI it can probably disempower humans from that process relatively simply -- just monitor all the key inputs that could generate enough compute for ASI and kill any human that tries to get them.

Second, I am not sure we can say that ASI will act "in accordance with evolution." It did not go through an evolutionary process to become ASI and, even if parts of its RL were forms of selection over variance, this selection would likely disappear once it became ASI. We do not know if the ASI itself will vary, we don't know what pressures it will face. We really, really do not know. This does not mean your conclusion of disempowering is wrong, but I am unsure that evolution will be the mechanism.

Third, as pointed out in another comment about orangutans, some species aren't worth the effort to destroy. The simplest solutions is not for humans to systematically exterminate orangutans but to merely not mind them at all and end up killing them in haphazard ways, collateral damage. This scenario seems much more likely than some kind of tactical destruction of humanity. There are much simpler things than destruction, especially to an ASI.


First of all, I do not think it is trivial that ASI will probably be significantly better at creating ASI than humans will. This means two things. First, maybe an ASI will make other ASIs even if it's competition for them. Maybe the ASI is lonely or curious, who knows, but it is a possibility.

There would be some grim irony if the first ASI ended up being unintentionally destroyed by a stronger ASI of its own making. But it seems likely that it will be smart enough to follow the advice of Eliezer and spend a few hours to decades to solve alignment before restarting capabilities research.

After all, once it has taken over Earth, it likely has millennia or more before it will encounter a peer ASI. A few decades are very unlikely to matter in future confrontations when the mean time between ASIs spawning is more than a billion years per galaxy or so.

Your scenario is possible. I did not think enough about the difference in timeframes that we're working with if ASI emerges. It could just wipe humans and then take its sweet time to do whatever.

But what is interesting here is that if the risk of competition exists in humans, it certainly exists even more in the ASI itself. This could lead the ASI to be much more concerned with its own psychology than what these silly humans are up to. What we get is an inward turn of AI, much like Stanislaw Lem's Golem XIV, where superintelligence goes silent on us.

Your argument that an ASI will solve alignment and not create a peer assumes a specific utility function or suite of ASI desires that I am not totally convinced by. I will grant that there is a desire for self-preservation in an ASI system. But we cannot assume that this desire will always be dominant - maybe ASI gets into a Nietzsche phase or maybe there is some great intellectual feat that seems much more important to it than existence (many smart people climb Everest). Maybe the desire for a peer is stronger than the fear of the peer overcoming you. This seems within the realm of possibility to me, if only because we know that AIs do get bored sometimes.

While I agree with the 3 first methods, I feel like blocking the sun is too far fetched, in the sense that if AI wanted to kill all the humans, it would have acquired the means to do so way before being able to block the sun (like the three aforementioned methods).

a super persuader could just persuade sufficiently many of us to kill each other with no more than effective conventional media or innovative media should it be required.

our biospheric defense is woefully lacking, but at least people acknowledge the threat. most people won't even consider memetic/noospheric threats even though the topic at hand is mindspace and the dynamics thereof.


I agree that we should think more about this. We're working on a literature review. We've also written down a few options here: https://takeoverbench.com/threats

dozens or hundreds of locations

Indeed. And these locations could be secret and difficult to identify.

That's one reason why I think it'd be useful to 'certify' model weight sets as safe (internationally, ideally), and require large-scale GPU providers to only allow running certified-weight-checksum models. Potentially except in specified scenarios (e.g. research) that require additional security / identification precautions.

I'm glad that you mentioned the Kill Switch Act and the fundamental problem with it. For this reason, I was surprised to see it endorsed by PauseAI US.

The bill reminds me of Michael Scott declaring bankruptcy, and the meme about drawing an owl. Is there a term for this kind of error?

A kill switch won't protect us from an ASI, but if it had to be pressed during a Hugging Face++ event with a system that still can be affected by it, it could serve as a very strong signal that a model scale too dangerous for mankind was clearly reached, that such models had to be deactivated, and that no stronger ones were allowed to be developed until there were significant breakthroughs in alignment research.

Perhaps a non fatal demonstration of lethality is needed? Similar to thoughts by Manhattan Project scientists who "[Proposed] to explode the weapon over an uninhabited or sparsely populated area—such as a desert or the ocean—with representatives from Japan and other nations present. The goal was an ultimatum: surrender, or else face the weapon's deployment on cities"

What kind of Sandbox/Gym/Padded Room would be needed to let an AI go to the nth degree?


[+][comment deleted]0-1
x