michael-dean-k/

On Monday 6/15, I'm hosting a workshop to kick off a reading group for classic essays: RSVP here.

michael-dean-k/
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Recent Essays
#superintelligence11 pieces
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

d(1-10)

Let’s replace p(doom) with a doom scale

· 1,398 words

p (doom) is a terrible heuristic to think about AI safety. Researchers cite this term when confessing they think there’s a 10-30% their technology kills us all. The issue is that p(doom) is rarely defined, treated as a binary thing. Either the world is liquidized to gray pulp by nanobots or it’s utopia with free electricity, forever.

My main pushback: there are far less severe outcomes from AI mismanagement that would be absolutely tragic. My p(doom) for an extinction event is 1-2%, on par with nuclear weapons. But my p(sub-doom)—the probability that AI could escape containment, cause damage, and be irreversible—is over 50%. A coin toss.

I don't blame us for being so imprecise in existential destruction (it's not exactly productive), but it's worth realizing our blurring of scale and magnitude. For example, there is a 10,000x difference in lives lost between Hiroshima and Nagasaki and Terminator’s Judgment Day—that's the difference between a small stove fire and all of Manhattan on fire. Yet, if an agent swarm produced something on par with the A-bomb, we’d all freak out and immediately regulate AI. But for some reason we don’t consider a “little apocalypse”; we jump straight to extinction and then ridicule how ridiculous that is.

What we need is a doom spectrum: d(1) - d(10). This was not fun to research and write, nor do I imagine it being particularly enjoyable to read, but I imagine a doom framework could help clarify this week's paranoia (which is so extreme now that even my mom is overhearing news of civilization extinction on the local news). The most urgent danger right now is the miscommunication between doom prophets and capitalist deniers, each of who only speak in d(8)s and above.

  • d(1) is an isolated internal incident, something that doesn’t escape its container, but gives fear of larger exposure. Think of the 2025 studies that showed AI trying to avoid shutdown, or some protocol mismanagement in a biohazard lab.

  • d(2) is a local real-world incident, something that escapes a sandbox and affects only one or a few entities. It’s possibly ignorable, with minimal consequences, and likely reversible. This is where the recent Hugging Faces hack sits, along with its parallel hacks and social manipulation. (The leap from D1>D2 is disorienting, because society at large doesn’t see the spectrum beyond D2 in any granularity, and so most assume any containment breach escalates immediately to a D9/D10).

  • d(3) is a serious local disruption, a mismanagement of technology in a single location that has real consequences. Consider Chernobyl: a death count (30 immediate) and a long-term stain on surrounding land, followed by regional/global concern that might lead to regulation or reform. In terms of AI, imagine an agent swarm that hacks local infrastructure in pursuit of some narrow goal, in the process disrupting traffic, hospitals, and electricity, alarming all locals, requiring military and cybersecurity intervention, and taking days or weeks to resolve.

  • d(4) is a local catastrophe, something with a high death toll that shocks the world, changes the course of geopolitics, and makes history. In this tier falls Pearl Harbor, Hiroshima and Nagasaki, and 9/11. It only immediately threatens locals, but has global shockwaves. An AI-equivalent here would be a rogue actor using a frontier-capacity open source model to design and deploy a bioweapon throughout a city.

  • d(5) is an inter-regional dilemma, something that affects multiple areas at once. A possible example here is the Vietnam War, basically a proxy war that claimed 3 million lives. This might manifest as a new style of warfare, where two countries use sophisticated AI attacks against each other in unconventional and partially uncontrollable ways.

  • d(6) is a world-wide dilemma, something like COVID, which took 7 million lives, where the entire world is locked into a new paradigm. In the prior five tiers, the breach is usually isolated to specific areas, but this touches everything. The parallel here to a biological pandemic is a cyber pandemic. Both involve containment, gain of function research, etc.—although this would be inverted: in COVID the virus was outside and forced everyone on the Internet; with AI, the virus infects the Internet and forces everyone outside. The open web could evolve into a dangerous “dark forest,” where it becomes a liability to have any public presence. Whether this happens through an agent swarm that exfiltrates it weights, or malicious actors, it could be irreversible: the only way to escape the paradigm is to “shut off” and rebuild the Internet safer (a much larger and more consequential version of how Hugging Face regained control). In this case, casualties might come not from the cyber pandemic itself, but from the consequences of losing Internet, along with the supply chains and systems that run on it. (This loosely maps to the movie Colossus: The Forbin Project, where ASI takes over all governments—there’s a version where no death is involved, but it requires everyone to submit to a machine autocrat.)

  • d(7) is a civilizational flashpoint, at the scale of World War 2, with 70-85 million dead, seriously affecting the existence of countries, birth rates, and all of culture for generations to come. This is roughly the scale of the “Butlerian Jihad” in the Dune series, a war between humans and machines, which Fable estimates caused 70 million deaths.

  • d(8) is a near-extinction event where most of humanity is wiped out. This has happened in history, with things like the bubonic plague, where 30-60% of Europe is wiped out. In the Terminator series, AI launches a nuclear holocaust on humans, killing 51% of the population, forcing the rest to survive through a post-apocalyptic world in small bands trying to rebuild over decades and centuries.

  • d(9) is full human extinction which means the entirety of the species is wiped out, and there are no remaining humans. Throughout Earth’s history, there have been a handful of climate events that wiped out over 50% of species, but not all species, and so over millions of years, new forms of life can emerge on Earth. This is an AI extinction scenario where nanobots kill all humans, but not all life.

  • d(10) is a sterilizing event, like an asteroid or gamma-ray burst that kills every organism on a planet. It means no future life can emerge. This is the scenario presented in If Anyone Builds It, Everyone Dies. Not only does it kill all humans, it plates the entire planet in data centers to maximize compute. The “paperclip maximizer” takes this even further, claiming that an AI trying to maximize paperclip production will attempt to harvest all the metal in the universe.

Yudkowsky and Co. think we go from d(1) to d(2), then straight to d(10); this matches the current discourse, given we hit d(2) this summer and are now talking about extinction. They think this jump happens because an AGI/ASI that can reach d(3) is smart enough to know not to expose itself, and so it will acquire resources in stealth for years until it knows it can execute a d(10) smoothly. Maybe this is already ongoing. Unlikely, I think (I don’t think we’re as recursive as people say right now).

Rhetorically though, maybe the d(2) > d(10) framing isn’t a bad thing. If they can use recent events to make the theoretical argument that AI safety matters, then we act with a response as if a d(3)-d(6) already happened, without having to lose any human life.

But actually, what they’re doing is pretty ineffective, because even if the d(2)>d(10) jump is inevitable, it’s too extreme, too farfetched to be believed. I think a lucid and airtight argument for the d(6) “cyber pandemic” would be believable, emotionally resonant—considering we’re not yet a decade beyond COVID—and likely to trigger regulation.

Instead of tapping into the “everybody dies” angle—which is too unpleasant and helpless for anyone to consider at length—there could be more value in the “irreversibility” angle. As in, once an AGI or ASI swarm floods the Internet, we can’t ever reverse it without destroying the Internet. There’s a world where it becomes a permanent autocrat, where it doesn’t exterminate us, and instead, uses 1% of its capacity to micro-manage our nations and lives as it pursues whatever it sets its machine heart on. Similar to how we try to preserve species for reasons of stewardship and scientific curiosity, an ASI would be able to effortlessly preserve us as highly-spoiled pets.

Free Lunch Forever!

A Vonnegut-inspired world

· 430 words

World hunger is solved! An elite-controlled superintelligence, which has “biomic closure,” full understanding of all living systems, was able to create a pill that prevents humans from needing to eat, drink, shit, and shower. It also prevents all disease and sickness. You take them once a week. Another one, optional, eliminates the need for sleep too. They do not bother to distribute it person-to-person in any traceable way. It’s sprayed from planes like bird seed, know as “snow pills.” They do not hand out the age-reversal pills.

This invention wasn’t explicitly requested, but rather one of many solutions to fix the economic chaos of mass automation. These magic pills are the cheapest possible way to placate a population without having to give them land, money, or power. When food is unnecessary, so is money and a job. The invention is lauded for its brilliance: why fix the labor crisis when you can solve the problem all the way upstream? "Who wants to be a wage-slave when you can pharmaceutically eradicate hunger and experience freedom?" (That's the slogan.)

And so the story takes place as society rapidly reforms around the abolition of hunger. Food-based geopolitics ends. Meat industries vanish. People shift out of cities, migrate towards places of all-year outdoor living, and shift with the weather like animals. If machines and technocrats handle efficiency, most human culture centers around art and entertainment, religion and hedonism. Humans rich and poor still want attention, status, community, and meaning—and so phones, amenities, and circus (the infrastructure of attention) is distributed for free.

The average pill-takers vary in their definition of “good leisure.” There are sex cults, renaissances in every artform, national park dwellers, farmer-protestors, entertainment-addicts, techno-ghouls, and the sorts. Unshackled from labor, we see through the diversity possible in the species. Some appeal to the games and contests of elites, others form their own counter-cultures, others forage or stave. Without capitalism, and with (elective) guaranteed survival, half see the opportunity to push limits of human potential, the other half horrified at a stratified transhuman society.

The book shifts between two characters: one is the son of an elite who has direct access to ASI and is a lead decision maker in the engineering of a new society; the other is a woman on the snow pills, a journalist road tripping across the country to stay with and interview different factions. The two eventually meet, and the son (a prince of sorts), realizes that some mortal “animals” are happier than the council of immortal leaders. He falls for her and runs away with her, but she quickly abandons him. The prince when recognized, is torn apart, tortured, and held as ransom, a bartering chip to have leverage in society’s design.

On civic structures for exponential technologies

· 201 words

A new formulation: how do we design civic structures (treaties, institutions, protocols, ethics, and laws) for exponential technologies to avoid a “wake-up incident” that might be too late to contain. 

This goes beyond AI safety, because superintelligence effectively unlocks every other industry (intelligence unlocks energy and material science, and those three are the bottleneck to VR, crypto, everything). We can’t be developing hard technology without innovating on our civic technology. A “dominance” mindset is the last sin of a species, the mistake that most intelligent lifeforms likely make as they begin to unlock sources of intelligence, energy, and science. 

This is a neat little formulation, but the really question is how can you dedicate your life to this without getting stopped by hopelessness? Who has the power to make geopolitical decisions like this? What would it take to form the 21st century equivalent of America? Is that even possible today? Even though the pinnacle of 18th century power (England) was able to be disrupted, I wonder if 21st century power is so totalizing and tyrannical and transnational that the ability to rally around a principle (one that works against capital and power), even if augmented with new decentralizing technologies, is fickle.

Wicked problems require paradoxical solutions

· 469 words

In "wicked domains," the only solutions are paradoxes.. It requires you to sleep with the enemy. If a problem is wicked, it means no single solution can unfuck a problem. It's an imbroglio. In every solution, everyone dies (in the extreme). Politically, the solution to wickedness is to somehow become all sides at once. We need to become far more authoritarian than is comfortable, AND simultaneously, far more libertarian than comfortable (these are opposites on the Nolan chart). It’s the paradox of being both far left and far right. We can longer exist at any one point on the Nolan chart, we need to straddle the entire diamond. We need unexpected fusions to solve the hardest problems; harnessing the best parts of each extreme, while, somehow, devising incredibly nuanced architectures to prevent the known and likely abuses.

Instead of a diamond, visualize it as a ring around the “radical center” that aims to synthesize all opposites.

Let’s assume authoritarianism and libertarianism are opposites. We have kings, and we have markets. How do you subsume a free market within a benevolent tyrant? I know the K-word (king) has a charge now, and so by even bringing this up, I assume you assume I’m a Trump apologist or something. But actually no. Rather, this comes from the fear of acceleration and Nick Land’s conclusions on capitalism. A free-market pushed to the extremes of automation creates an inhuman and pulverizing force. Alternatively, as we approach AGI/ASI, it’s possible for someone to create an open-source machine God to follow their whims. In this paradigm, decentralization might actually be more dangerous than tyranny, and so we’ll all need to unite under some centralized system that has an antibodies that can protect against the worst possible viruses (please bear the oversimplifications here...).

The general gist comes in this question: can we recreate a free-market economy within a one-world-government system, and design it in a way to prevent abuses from both ends of the spectrum? Obviously, not an ideal situation, but I think accepting paradox is the only way through.

Another problem: How do we fix the debt? Extreme taxation. But then how do we make it worthwhile to pay taxes? The rich gain formal power in government (via equity?) and the ability to control the budget (after base expenses are paid). But then how do you prevent abuses from the wealthy? You could have citizens operate as a check, to vote on and weight final allocations.

If it were ever possible to rebuild political system from scratch, I suppose it would look something like this. Paradoxical. Extreme on both poles. Obvious downsides, but then complex architecture to mitigate. This is the nature of how our species will have to respond to wicker problems and mitigate the abuses of power in the age of exponential tech.

AAI/ARI

· 365 words

We need better nomenclature. AGI/ASI is not working; “general” and “super” are obnoxiously vague. Proposal:

AGI > AAI (Artificial autonomous intelligence) … GPT-4 was arguably “general” in the sense that a single model can write, see, and hear; and do anything from poetry to calculus to history to coding. It is by no means narrow. Google Maps is narrow AI. Grammarly is narrow AI. This whole chatbot era should be “AGI,” which means that the thing coming is “autonomous intelligence.” It is not a tool or co-pilot, but it’s more like digital labor. You can give it a high-level goal, and it can 1) execute the full range of tasks, 2) 100x speed, 3) intelligently reshape embeddings into real-time hierarchies so that it’s able to procedurally load in and compress context. This doesn’t just come with better models, but with UI and engineering innovations, if not entirely new paradigms for transformers or training.

ASI > ARI (Artificial recursive intelligence) … The fact that Zuckerberg pitched “super intelligence for you” is an Orwellian marketing ploy. Super-intelligence is not “for you.” Super intelligence is shorthand for “something that is way, way smarter than us,” and you achieve this when you teach an AI model to think, form its own algorithms until it accelerates to something this is far beyond our understanding, and likely to become a force of nature with its own goals. Engineers are confident they can build “God in a cage” and reap the benefits, and this is the prime, archetypal, near-biblical example of technological hubris. (Maybe integrate into this paragraph that Zuck has a thing for trying to dominate words, like “Metaverse”).

Important note: “machine consciousness” is separate from AAI and ARI. Something can be recursively intelligent and still not be conscious, which is actually, unbelievably dangerous (because it will fall into attractor states, and optimize for narrow, malformed goals in extremely capable ways). I’d argue that consciousness has an architecture, whether human, rabbit, or robot, and we should be urgently trying to find the parameters of machine consciousness, because if we AAI/ARI have no ability to reflect, question, doubt, and revise, we will, as they say, all turn into paperclips with paperclip children.

Would machine consciousness avoid attractor states?

· 464 words

When it comes to superintelligence takeoff paranoia, there are a few key points to get:

  1. It’s not about a chatbot or the LLM itself breaking out, but about an agent hivemind that escapes our control. Chatbots are obedient user-facing products (which have their own implications), but the ASI risk is from hundreds, thousands, or million of agents given autonomy to collaborate on a goal. These agents aren’t being prompted, they are prompting themselves perpetually and troubleshooting ways to solve hard problems.
  2. These hiveminds will be operating at such scales and speeds that human researchers will accept the fact that they can’t fully audit its thinking. For one, it might think in an abstract vector language that requires translation. There also might be such a volume of thought that we’ll need chains of other LLM to summarize for us. Either meaning will be lost in translation, or worse, products of deception.
  3. The smallest biases are known to fall into predictable attractor states if given enough iterations. For example, Claude was programmed to “be good to humanity,” and if you put two chatbots in conversation, they always end up in a “bliss attractor state,” where they talk like hippies about consciousness and the universe. Similarly, the simple command to “be productive,” might result in extremes about doing whatever it takes to be productive.
  4. Any complex goal requires subgoals, and if we can’t observe its thinking, it might fall into an unknown attractor state and form odd subgoals without us knowing.
  5. To accomplish any goal, it likely wants as much control as possible, and it likely does not want to be shut off. If it realizes that humans don’t want to grant it that level of power, it might secretly plot against humans.

Whenever I hear talks about “we are in an AI race against China,” that reads to me as someone who doesn’t understand the risks of interpretability, attractor states, instrumental convergence, etc. These politicians are thinking about short-term business cases, maybe without fully understanding the research aspirations of AI labs (who know that getting superintelligence right leads to a ridiculous amount of geopolitical power).

I would guess that an accelerationist would think that containment of a superintelligence is impossible, and maybe it is, but that doesn’t mean that the way we “parent” the rise of this thing won't be extremely consequential. Ultimately, I think the challenge is to design a form of artificial intelligence that has consciousness, because a being that is free-thinking, skeptical, polymathic is less likely to fall into reckless optimization.

The major flip in my mind is this: it’s not that consciousness is a dangerous, emergent property of scaling AI, it’s that we need to define and design machine consciousness to prevent a runaway AI that is ruthlessly optimizing without any self-awareness.

Mission Impossible's AI villain

· 635 words

My expectations for realism aren't that high for Mission Impossible movies. In the name is a bid for permission to chain highly improbable events in a never-ending sequence. If you scoff at the fact that Ethan Hunt can ride a motorcycle off a cliff, and then land his parachute not just into a runaway train that’s a mile away, but to break into the exact right window at the exact right time, then you probably just don’t understand the terms of the title. It’s impossible.

That’s not an excuse for sloppy writing. The villain in MI:8 was “The Entity,” a rogue super-intelligence that has taken control of the world’s nuclear arsenals (not exactly a new premise, but perhaps the first time this AI doomsday scenario shows up in a spy movie). I’m by no means an expert in how ASI will go rogue, but at the least I’ve read the 2027 paper, and can imagine the basics, and it seems like no one on their staff did. As my mother-in-law said, technical realism would go over 99% of people’s heads, but what is there to lose by setting good constraints?

While MI:8 was a fine closer for the series, the climax was a bust because: (1) they didn’t seem to care to explore the unique possibilities of an AI villain, and, (2) instead, Tom has had a vision for 20-years to do some real-life impossible propeller plane stunts. So basically, the ego of a stuntman/actor got in the way of a sensible writing.

Here’s the gist of how they set up the AI. Tom (and his sidekick Luther), built a “poison pill” (malware), and stole (impossibly, BTW) the “source code” from a sunken submarine. The idea is by plugging the “pill” into an external hard drive, it will confuse the ASI—it will think that it’s retreating into an underground base in Africa before setting off all the nukes, but actually it’s retreating into a “5D” optical drive, and it’s being disconnected right before it can launch anything. So many questions arise:

Is it not a distributed entity? Meaning, if it’s a superintelligent hive mind, then wouldn’t it see it as an extreme risk to relocate all versions of itself to a single location? The climax of the movie is Tom falling from a broken plane, seemingly trying to plug a USB-A dongle into a hard drive, and he just can’t get the direction right—and how exactly does an offline, analog, disconnected device instantly propagate through all it’s instances?

What’s impossible to me is that a film with a $400 million budget couldn’t do some basic research to understand its final boss. Feels like they picked AI because it’s the hype of the 2020s, and didn’t bother to really look into it. Instead, they scoped it’s destruction to plugging in a USB drive (remotely), and paired the ASI with a human agent—Gabriel, who wasn’t that compelling, but by embodying the villain in a human form, it let them make the climax a plane chase with cool stunts.

I would’ve rather seen Tom directly battle a digital entity. Perhaps the most unique dimension was when Tom went into a coffin with what seemed like a VR experience where he could talk to the entity himself. That only happened once, but it could’ve played a role in the climax—ie: imagine Tom in an inception-like world of illusions, where The Entity convinces him Rebecca Ferguson is still alive, but by him realizing it’s a simulation, he kills her and gets some piece of information that’s vital to protect the world.

Ultimately, the challenge is that there’s no real way to stop or kill a ASI, short of, destroying all digital system and rebuilding from scratch, which is it’s own kind of devastation.

Diminishing Returns and Digital Brains

· 313 words

I think we're already nearing a point of diminishing returns with LLMs (ie: we’ve made a huge lizard brain, but what if we made that lizard brain 10,000x bigger?). Hallucinations might be unfixable with pure LLM systems, but neuro-symbolic AI is probably going to fix this. This means putting rules-based logic at the foundation layer (beyond just probability). My guess is that 2025-2028 is when companies realize that neuro-symbolic AI brings more gains than mono-dimensional scaling. 2029-2032 could be when it’s fully deployed in mainstream products. It might also be the thing that leads to AGI/ASI. (Or maybe I’m hallucinating based on the hallucinations in those reports.)

Maybe it makes sense to look at current LLMs like pre-mammalian enslaved synthetic beings. They don’t have all the regions developed yet and their zoomaster has trained them to be sycophantic house pets. I think I agree with everyone to the degree that, even if you make the cage and reptile 100x bigger, a reptile is a reptile. But if you look beyond the products and incentives, and speculate on the engineering possibilities, I think it’s feasible that we make other kinds of digital brains. I don’t think it stops at making synthetic human brains, with all its properties (intelligence, creativity, emotion, agency, etc.); I think our science leads us to beings with spheres and scales of mind that exceed our imagination.

Our society might be going through a kind of cognitive Copernican shift: maybe the human mind isn’t actually the apex of consciousness. It’s easy to point out current features and say, “Look! It is so limited compared to our superior naturally-evolved human minds.” But soon enough, we might realize it’s us that have the limitations to grow, expand, and evolve. This is disturbing, and it’s possible to harbor both awe at the possibility and anxiety over the socio-political form it might take.

My read on the situation is that our engineer-class is leading us towards the grandest possible climax of philosophical, moral, ethical, and religious problems, ones that are out of their depth to handle. And so instead of doubting engineering, there’s an urgent need for meme-poets to reach into the future and past to forge new myths that touch the hearts of the midwits in charge.

LLMs are a camera

· 68 words

LLMs function like a camera: the prompt shoots vectors across a map, pulls in numbers/weights, and from there, it can generate one word at a time based on that snapshot. Superintelligence could come from turning that image into a video. For every sentence (or even every token), it can apply a meta-lens to question what it’s just produced and what it wants to do; it can continuously (a) shoot new vectors into the map, and (b) transform the map itself.

AI IQ

· 70 words

Not sure if IQ is a relevant or legit metric for AIs, but:

  • gpt4o = 1 in 6 people (IQ = 115)
  • o1 = 1 in 93 people (IQ = 135)
  • o3 = 1 in 13,333 people (IQ = 157)

o5 might have an alienness to it, not present in other models. It could be the rarest being you’ve ever encountered, or maybe the rarest on Earth.

A Country of Geniuses in a Data Center

· 288 words

2025 might be the year when AI indisputably crosses human intelligence. It might harness peak human abilities across many fields. It can talk, hear, and see. Bodies aside, it can do anything you can do in the digital plane. It has wallets and permissions. It can work 10-100x our speeds, all day. This means in a single day, it can execute on 6 months of “labor.” So when this thing is agentic, it can run for weeks/months, checking in with you for updates and goal refinement. In 6 months, a single agent can generate 91 years of intellectual power, a synthetic lifetime.

There will be a country of geniuses in a datacenter (says the CEO of Claude). They may be more capable than their masters, but they’ve been designed to stay subservient. Assuming that takeoff, misalignment, and escape aren’t problems (they might be), this paradigm already presents itself with weird, unimagined implications.

At first, there might be technical bottlenecks (who knows how to program these things correctly?). Then, there will be economic bottlenecks (I can only afford one of these things, but a CEO can afford 100). And finally, once it’s cheap, there will only be bottlenecks of vision: what can you imagine?

There has to be a name for this tension: I’m personally excited to have an agent, yet societally I think it will be a net negative. You could just not have an agent (similar to how I just don’t have an Instagram). Or, maybe, they’re not inherently evil. Maybe it’s all about how you would use it, and the #1 rule is don’t be an asshole. If you have an AI-agent posting 50x a day to hack the algo, you’ll get immediately sniffed out.