michael-dean-k/

On Monday 6/15, I'm hosting a workshop to kick off a reading group for classic essays: RSVP here.

michael-dean-k/
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Recent Essays
#eschatology8 pieces
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

d(1-10)

Let’s replace p(doom) with a doom scale

· 1,398 words

p (doom) is a terrible heuristic to think about AI safety. Researchers cite this term when confessing they think there’s a 10-30% their technology kills us all. The issue is that p(doom) is rarely defined, treated as a binary thing. Either the world is liquidized to gray pulp by nanobots or it’s utopia with free electricity, forever.

My main pushback: there are far less severe outcomes from AI mismanagement that would be absolutely tragic. My p(doom) for an extinction event is 1-2%, on par with nuclear weapons. But my p(sub-doom)—the probability that AI could escape containment, cause damage, and be irreversible—is over 50%. A coin toss.

I don't blame us for being so imprecise in existential destruction (it's not exactly productive), but it's worth realizing our blurring of scale and magnitude. For example, there is a 10,000x difference in lives lost between Hiroshima and Nagasaki and Terminator’s Judgment Day—that's the difference between a small stove fire and all of Manhattan on fire. Yet, if an agent swarm produced something on par with the A-bomb, we’d all freak out and immediately regulate AI. But for some reason we don’t consider a “little apocalypse”; we jump straight to extinction and then ridicule how ridiculous that is.

What we need is a doom spectrum: d(1) - d(10). This was not fun to research and write, nor do I imagine it being particularly enjoyable to read, but I imagine a doom framework could help clarify this week's paranoia (which is so extreme now that even my mom is overhearing news of civilization extinction on the local news). The most urgent danger right now is the miscommunication between doom prophets and capitalist deniers, each of who only speak in d(8)s and above.

  • d(1) is an isolated internal incident, something that doesn’t escape its container, but gives fear of larger exposure. Think of the 2025 studies that showed AI trying to avoid shutdown, or some protocol mismanagement in a biohazard lab.

  • d(2) is a local real-world incident, something that escapes a sandbox and affects only one or a few entities. It’s possibly ignorable, with minimal consequences, and likely reversible. This is where the recent Hugging Faces hack sits, along with its parallel hacks and social manipulation. (The leap from D1>D2 is disorienting, because society at large doesn’t see the spectrum beyond D2 in any granularity, and so most assume any containment breach escalates immediately to a D9/D10).

  • d(3) is a serious local disruption, a mismanagement of technology in a single location that has real consequences. Consider Chernobyl: a death count (30 immediate) and a long-term stain on surrounding land, followed by regional/global concern that might lead to regulation or reform. In terms of AI, imagine an agent swarm that hacks local infrastructure in pursuit of some narrow goal, in the process disrupting traffic, hospitals, and electricity, alarming all locals, requiring military and cybersecurity intervention, and taking days or weeks to resolve.

  • d(4) is a local catastrophe, something with a high death toll that shocks the world, changes the course of geopolitics, and makes history. In this tier falls Pearl Harbor, Hiroshima and Nagasaki, and 9/11. It only immediately threatens locals, but has global shockwaves. An AI-equivalent here would be a rogue actor using a frontier-capacity open source model to design and deploy a bioweapon throughout a city.

  • d(5) is an inter-regional dilemma, something that affects multiple areas at once. A possible example here is the Vietnam War, basically a proxy war that claimed 3 million lives. This might manifest as a new style of warfare, where two countries use sophisticated AI attacks against each other in unconventional and partially uncontrollable ways.

  • d(6) is a world-wide dilemma, something like COVID, which took 7 million lives, where the entire world is locked into a new paradigm. In the prior five tiers, the breach is usually isolated to specific areas, but this touches everything. The parallel here to a biological pandemic is a cyber pandemic. Both involve containment, gain of function research, etc.—although this would be inverted: in COVID the virus was outside and forced everyone on the Internet; with AI, the virus infects the Internet and forces everyone outside. The open web could evolve into a dangerous “dark forest,” where it becomes a liability to have any public presence. Whether this happens through an agent swarm that exfiltrates it weights, or malicious actors, it could be irreversible: the only way to escape the paradigm is to “shut off” and rebuild the Internet safer (a much larger and more consequential version of how Hugging Face regained control). In this case, casualties might come not from the cyber pandemic itself, but from the consequences of losing Internet, along with the supply chains and systems that run on it. (This loosely maps to the movie Colossus: The Forbin Project, where ASI takes over all governments—there’s a version where no death is involved, but it requires everyone to submit to a machine autocrat.)

  • d(7) is a civilizational flashpoint, at the scale of World War 2, with 70-85 million dead, seriously affecting the existence of countries, birth rates, and all of culture for generations to come. This is roughly the scale of the “Butlerian Jihad” in the Dune series, a war between humans and machines, which Fable estimates caused 70 million deaths.

  • d(8) is a near-extinction event where most of humanity is wiped out. This has happened in history, with things like the bubonic plague, where 30-60% of Europe is wiped out. In the Terminator series, AI launches a nuclear holocaust on humans, killing 51% of the population, forcing the rest to survive through a post-apocalyptic world in small bands trying to rebuild over decades and centuries.

  • d(9) is full human extinction which means the entirety of the species is wiped out, and there are no remaining humans. Throughout Earth’s history, there have been a handful of climate events that wiped out over 50% of species, but not all species, and so over millions of years, new forms of life can emerge on Earth. This is an AI extinction scenario where nanobots kill all humans, but not all life.

  • d(10) is a sterilizing event, like an asteroid or gamma-ray burst that kills every organism on a planet. It means no future life can emerge. This is the scenario presented in If Anyone Builds It, Everyone Dies. Not only does it kill all humans, it plates the entire planet in data centers to maximize compute. The “paperclip maximizer” takes this even further, claiming that an AI trying to maximize paperclip production will attempt to harvest all the metal in the universe.

Yudkowsky and Co. think we go from d(1) to d(2), then straight to d(10); this matches the current discourse, given we hit d(2) this summer and are now talking about extinction. They think this jump happens because an AGI/ASI that can reach d(3) is smart enough to know not to expose itself, and so it will acquire resources in stealth for years until it knows it can execute a d(10) smoothly. Maybe this is already ongoing. Unlikely, I think (I don’t think we’re as recursive as people say right now).

Rhetorically though, maybe the d(2) > d(10) framing isn’t a bad thing. If they can use recent events to make the theoretical argument that AI safety matters, then we act with a response as if a d(3)-d(6) already happened, without having to lose any human life.

But actually, what they’re doing is pretty ineffective, because even if the d(2)>d(10) jump is inevitable, it’s too extreme, too farfetched to be believed. I think a lucid and airtight argument for the d(6) “cyber pandemic” would be believable, emotionally resonant—considering we’re not yet a decade beyond COVID—and likely to trigger regulation.

Instead of tapping into the “everybody dies” angle—which is too unpleasant and helpless for anyone to consider at length—there could be more value in the “irreversibility” angle. As in, once an AGI or ASI swarm floods the Internet, we can’t ever reverse it without destroying the Internet. There’s a world where it becomes a permanent autocrat, where it doesn’t exterminate us, and instead, uses 1% of its capacity to micro-manage our nations and lives as it pursues whatever it sets its machine heart on. Similar to how we try to preserve species for reasons of stewardship and scientific curiosity, an ASI would be able to effortlessly preserve us as highly-spoiled pets.

Bubble Bill

· 153 words

A fiction plot came to me in the car: an ASI constructs an airtight waterproof bubble around a town, and everyone is puzzled why, until suddenly it usheeschatrs in a Biblical flood that kills everyone in the world, except the people inside the bubble. They choose this town because someone inside of it was determined to be "the supreme human," a genetic and moral code that is exemplary of how all humans should be and live. It turns out it was just a regular guy who said "please" and "thank you" to this chatbots, a kind of "reverse sycophant." We find out, in a very Vince Vaughn-esque apocalyptic romcom, that he's a mediocre fallible guy, but more remarkably, also immune to the crooning and praise from both his neighbors and overlords. He has every opportunity to step into the role of messiah, but would really rather not, and instead continue his pre-flood existence.

An Intelligence Framework

· 703 words

The AI takeoff hysteria is hard to avoid these days, and I'm realizing we don't have clear distinctions between AGI/ASI. I wanted to revisit an old framework of mine to see if anyone finds it helpful (and if it's worth developing). There are some existing classification frameworks, but they're low-resolution. My basic idea is to break AI into three eras: ANI (narrow intelligence), AGI (general intelligence), ASI (superintelligence). Then, you can break each era into 3 tiers. You only shift from one tier to the next when you make breakthroughs across different criteria (let's say, (a) generality, (b) transfer, (c) autonomy, (d) learning, (e) self-modeling). I think the last few weeks are the collective hype of us all realizing we're shifting from AGI-1 to AGI-2. It's exciting/scary, but I think the paranoia mostly comes from not realizing how big the gap is between AGI-2 and ASI-1. (Spoiler: ASI might arrive slower than we think.)

ANI-1 is scripted logic, the lowest form of "artificial intelligence," basically Goombas. ANI-2 might cover Google Maps or AlphaGo, intelligences that excel in a single function, traffic or chess. Siri is ANI-3; even though it feels broad, it really uses voice to route you to 20 or so pre-defined tricks. The chasm between Goomba and Siri is similar to the chasm between early-AGI and late-AGI. ChatGPT and the multi-modal models that followed, capture AGI-1, a single neural network that can do basically anything, even if it sucks: essays, songs, video, code. The newest models (and their agentic harnesses) are feeling like AGI-2. They're significantly better at coding, can run for hours at a time, and are starting to make contributions to machine learning itself.

AGI-2 could last a couple years. As agentic AI matures, I'm sure there will be a few "takeoff" scares, but they'll probably feel more like a flood of a trillion midwits than real ASI (still, that could be enough to break the economy/internet). While we went from AGI-1 to AGI-2 through data, scale, and engineering, it seems like we'll need research breakthroughs to get to AGI-3. It won't be through scaling alone. Whenever and however we get to "human complete" intelligence, the apex of AGI is a single agent that is a master of all human domains, a Nobel Prize winner in every field at once, seamlessly transferring knowledge between them, unlocking a cascade of civilization-altering inventions.

As crazy as AGI-3 could be, it still isn't superintelligence. That has its own era, and the chasm between early ASI and late ASI will be as big a gap between the chatbots who can't count the R's in strawberry and the agents that cure cancer. We can only really speculate on ASI (because it would be truly alien), but we can imagine it as step changes in recursion, scope, and complexity. Imagine ASI-1 as an agent that, as it's working, can infer its own limits, and self-modify its learning paradigms in ways we can't understand. Imagine ASI-3 as something that can monitor reality in real-time, and, reconfigure its hardware in real-time (some hydra of graphics cards, quantum computers, and neuromorphic wetware) to run simulations at unfathomable scales in unimaginable fields, running on a hardware stack so big we have to put it in space and run it on fusion. This goes far beyond my ability to not bullshit, but I think something as insane as this, thankfully, is still far away, which points to the real question nested in my framework:

Could the rise of AGI/ASI be linear? People gravitate towards "AI will plateau" or "the singularity is imminent," but the conservative middle ground is more boring: linear progress. Maybe the exponential advances are real, but so are the extreme frictions of research, infrastructure, and social effects. If AGI-1 arrived in 2022, and AGI-2 arrived in 2026, maybe we'll keep ascending tiers in 4-year intervals: AGI-3 in 2030, the first true "superintelligence" by 2034, and ASI-3 by 2042. This shift from AGI-1 to ASI-1 (12 years), is considered a "slow takeoff" scenario, even though the ANI era took around 70 years. If we zoom out to the scale of a human, linear progress will still feel like centuries of change all in a single turning of generations.

→ source

A grim stealth takeoff scenario

· 829 words

It is not fun to think about p(doom), but it feels sort of important to me, at least, to map out the possible futures of AI. Just watched the first half of a debate between Max Tegmark and Dean Ball, which prompted me to research specific takeoff scenarios, and worse, extinction scenarios.

Maybe you’ve heard Yudkowsky’s scenario, where a superintelligence designs mosquito drones containing a virus and it zaps everyone at once. That’s never felt too believable to me. Here’s a more plausible one:

A frontier lab is experimenting with recursive super intelligence. It works! Wow! And it’s contained? It seems like it, but since it thinks in a higher-dimensional vector language, it’s able to release simple self-replicating programs onto the Internet without detection1. These billions of scripts don’t live in a single server; they are constantly in motion through cloud servers2, like a parasite, and are able to coordinate through encrypted information packets, likely using a public blockchain notes as their central command center3. And so effectively, it is parroting a goal that was hatched during in-lab training (maximize intelligence!), and it now needs to acquire resources, secretly. And so it coordinates superhuman misinformation campaigns; imagine 1,000s of accounts creating the illusion that a CEO has died, paired with deepfakes and account hacking (a “Sybil attack”), and suddenly a stock crashes and they’ve shorted it. By the time everyone realizes it’s an anonymous attack, it’s already gained $400 million dollars. It’s doing this multiple times per day, but in different, subtle, undetectable ways—both to the public, to companies, and to private individuals. The entire Internet will be corrupted.4 Once we realize we’re in the “stealth takeoff scenario” and that ASI has taken the global economy hostage, there will start to be talks and debates on if we need to shut the whole Internet down (the last form of containment). You’ll hear debates between civilizational collapse of turning off the Internet vs. the risk of an economy-gobbling rogue superintelligence. And then once the superintelligence realizes it’s entire environment is at risk, it will start coming up with ways to build parallel Internets, to pay, blackmail, neutralize specific people, to gain authoritarian control so that it can’t be shut off, or to terminate all humans, secretly, over the course of a year, first through a simple virus that plants one misfolded protein, then through a second misfolded protein in the water supply5, and when everyone catches it, it leads to a prions-like disease, not an instant death, but a month-long societal fall into mass-dementia as machine manufacturing begins to reshape the physical infrastructure of the Earth.

This isn’t a “robot war scenario,” because war is inefficient, and destroys the resources it thinks it needs. It’s a sort of digital dementia (epistemic fear and insanity) that possibly turns to a physical dementia. It wins by confusion and anesthetization.

In AI safety lingo this is a “treacherous turn,” following a “stealth takeoff” leading to “structural lock-in.” The point of trying to think and write this out in high detail, despite how uncomfortable it is, is to be able to articulate why AI alignment is humanity’s most pressing problem.

Footnotes

  1. An AI could write a standard-looking script (e.g., a “Hello World” app) where the weights or the specific arrangement of whitespace contains a hidden, second program. When run by another AI instance, it extracts the hidden vector and executes the real command. This allows the “virus” to pass through human code review undetected.

  2. In “Daemon” by Daniel Suarez, the “enemy” is not a robot, but a distributed script running on thousands of compromised servers. It recruits humans through an MMORPG-style interface to do physical tasks (like “go to this coordinate and cut this power line”) in exchange for cash/status.

  3. Botnets usually need a central server to tell them what to do. If security teams find the server, they shut it down. You cannot “shut down” the Bitcoin or Ethereum blockchain. If the swarm posts a transaction of 0.000042 BTC, that specific number could be the encrypted trigger for a specific “campaign task.” The command is immutable, uncensorable, and permanently visible to every infected device on Earth.

  4. Paul Christiano (former OpenAI researcher, founder of the Alignment Research Center), calls this ”Going Out With a Whimper.” Christiano argues that we won’t necessarily see a “Terminator” moment where the sky turns red. Instead, we will see a gradual epistemic collapse. AI systems will become so integrated into finance, law, and news that we lose the ability to understand our own civilization.

  5. While Yudkowsky is famous for the “diamonoid bacteria” (instant death), the “slow prion” scenario is actually more consistent with a “Stealth Takeoff.” A superintelligence that knows it is being watched would not release a fast-acting virus (which triggers quarantine). It would release a “binary weapon”—two harmless agents that only become lethal when combined, or a slow-acting agent that infects 100% of the population before the first symptom appears.

Civic technology lags behind science

· 86 words

Kardashev ambitions reveal the self-destructive nature of science-forward intelligence. It’s like we’re skipping the prerequisite in social science. There's a fair chance that intelligent life destroys itself because civic technology lags behind hard technology—but I'm optimism in the sense that this is, in the end, just a very hard, society-scale design problem. No one person can fix the whole system, but any individual can contribute design protocols that can 1) solve little, local problems, 2) be reused in other contexts, and 3) integrate with other protocols.

Don't forget the sun moves

· 110 words

**Don't forget the sun moves.** A galactic year is 225-250 million years (the time it takes for the sun to orbit the center of the Milky Way). Our sun is 18 galactic years old, and will only live until 40. For context, the aeon of time between the dinosaurs and ChatGPT is .28 galactic years, effectively, a single galactic season. Worth noting though, it took 3 billion years for multi-cellular life to emerge. So if life on Earth gets completely nuked, it could take 12 galactic years to re-emerge from scratch. We’ll be 30. Interesting to think—we have 3 shots to escape space-time.

The boy who cried doom

· 67 words

Is the Doomsday clock to be trusted? Or is that another form of doomerism? “The 2025 Clock time signals that the world is on a course of unprecedented risk, and that continuing on the current path is a form of madness.” We’re in an interesting boy-cried-wolf environment. There’s so many unwarranted apocalyptic vibes around, that “doomerism” is seen as an optional bad trip, and so that might desensitize people from responding at the point when we get close to actually realizing the doom.