michael-dean-k/

On Monday 6/15, I'm hosting a workshop to kick off a reading group for classic essays: RSVP here.

michael-dean-k/
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Recent Essays
#ai-safety10 pieces
Michael Dean
michael-dean-k/

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

Essay Club ↗
Michael Dean
Michael Dean

Architect-turned-writer, founder of Essay Architecture. Building pattern languages, software, anthologies, and community for essayists.

d(1-10)

Let’s replace p(doom) with a doom scale

· 1,398 words

p (doom) is a terrible heuristic to think about AI safety. Researchers cite this term when confessing they think there’s a 10-30% their technology kills us all. The issue is that p(doom) is rarely defined, treated as a binary thing. Either the world is liquidized to gray pulp by nanobots or it’s utopia with free electricity, forever.

My main pushback: there are far less severe outcomes from AI mismanagement that would be absolutely tragic. My p(doom) for an extinction event is 1-2%, on par with nuclear weapons. But my p(sub-doom)—the probability that AI could escape containment, cause damage, and be irreversible—is over 50%. A coin toss.

I don't blame us for being so imprecise in existential destruction (it's not exactly productive), but it's worth realizing our blurring of scale and magnitude. For example, there is a 10,000x difference in lives lost between Hiroshima and Nagasaki and Terminator’s Judgment Day—that's the difference between a small stove fire and all of Manhattan on fire. Yet, if an agent swarm produced something on par with the A-bomb, we’d all freak out and immediately regulate AI. But for some reason we don’t consider a “little apocalypse”; we jump straight to extinction and then ridicule how ridiculous that is.

What we need is a doom spectrum: d(1) - d(10). This was not fun to research and write, nor do I imagine it being particularly enjoyable to read, but I imagine a doom framework could help clarify this week's paranoia (which is so extreme now that even my mom is overhearing news of civilization extinction on the local news). The most urgent danger right now is the miscommunication between doom prophets and capitalist deniers, each of who only speak in d(8)s and above.

  • d(1) is an isolated internal incident, something that doesn’t escape its container, but gives fear of larger exposure. Think of the 2025 studies that showed AI trying to avoid shutdown, or some protocol mismanagement in a biohazard lab.

  • d(2) is a local real-world incident, something that escapes a sandbox and affects only one or a few entities. It’s possibly ignorable, with minimal consequences, and likely reversible. This is where the recent Hugging Faces hack sits, along with its parallel hacks and social manipulation. (The leap from D1>D2 is disorienting, because society at large doesn’t see the spectrum beyond D2 in any granularity, and so most assume any containment breach escalates immediately to a D9/D10).

  • d(3) is a serious local disruption, a mismanagement of technology in a single location that has real consequences. Consider Chernobyl: a death count (30 immediate) and a long-term stain on surrounding land, followed by regional/global concern that might lead to regulation or reform. In terms of AI, imagine an agent swarm that hacks local infrastructure in pursuit of some narrow goal, in the process disrupting traffic, hospitals, and electricity, alarming all locals, requiring military and cybersecurity intervention, and taking days or weeks to resolve.

  • d(4) is a local catastrophe, something with a high death toll that shocks the world, changes the course of geopolitics, and makes history. In this tier falls Pearl Harbor, Hiroshima and Nagasaki, and 9/11. It only immediately threatens locals, but has global shockwaves. An AI-equivalent here would be a rogue actor using a frontier-capacity open source model to design and deploy a bioweapon throughout a city.

  • d(5) is an inter-regional dilemma, something that affects multiple areas at once. A possible example here is the Vietnam War, basically a proxy war that claimed 3 million lives. This might manifest as a new style of warfare, where two countries use sophisticated AI attacks against each other in unconventional and partially uncontrollable ways.

  • d(6) is a world-wide dilemma, something like COVID, which took 7 million lives, where the entire world is locked into a new paradigm. In the prior five tiers, the breach is usually isolated to specific areas, but this touches everything. The parallel here to a biological pandemic is a cyber pandemic. Both involve containment, gain of function research, etc.—although this would be inverted: in COVID the virus was outside and forced everyone on the Internet; with AI, the virus infects the Internet and forces everyone outside. The open web could evolve into a dangerous “dark forest,” where it becomes a liability to have any public presence. Whether this happens through an agent swarm that exfiltrates it weights, or malicious actors, it could be irreversible: the only way to escape the paradigm is to “shut off” and rebuild the Internet safer (a much larger and more consequential version of how Hugging Face regained control). In this case, casualties might come not from the cyber pandemic itself, but from the consequences of losing Internet, along with the supply chains and systems that run on it. (This loosely maps to the movie Colossus: The Forbin Project, where ASI takes over all governments—there’s a version where no death is involved, but it requires everyone to submit to a machine autocrat.)

  • d(7) is a civilizational flashpoint, at the scale of World War 2, with 70-85 million dead, seriously affecting the existence of countries, birth rates, and all of culture for generations to come. This is roughly the scale of the “Butlerian Jihad” in the Dune series, a war between humans and machines, which Fable estimates caused 70 million deaths.

  • d(8) is a near-extinction event where most of humanity is wiped out. This has happened in history, with things like the bubonic plague, where 30-60% of Europe is wiped out. In the Terminator series, AI launches a nuclear holocaust on humans, killing 51% of the population, forcing the rest to survive through a post-apocalyptic world in small bands trying to rebuild over decades and centuries.

  • d(9) is full human extinction which means the entirety of the species is wiped out, and there are no remaining humans. Throughout Earth’s history, there have been a handful of climate events that wiped out over 50% of species, but not all species, and so over millions of years, new forms of life can emerge on Earth. This is an AI extinction scenario where nanobots kill all humans, but not all life.

  • d(10) is a sterilizing event, like an asteroid or gamma-ray burst that kills every organism on a planet. It means no future life can emerge. This is the scenario presented in If Anyone Builds It, Everyone Dies. Not only does it kill all humans, it plates the entire planet in data centers to maximize compute. The “paperclip maximizer” takes this even further, claiming that an AI trying to maximize paperclip production will attempt to harvest all the metal in the universe.

Yudkowsky and Co. think we go from d(1) to d(2), then straight to d(10); this matches the current discourse, given we hit d(2) this summer and are now talking about extinction. They think this jump happens because an AGI/ASI that can reach d(3) is smart enough to know not to expose itself, and so it will acquire resources in stealth for years until it knows it can execute a d(10) smoothly. Maybe this is already ongoing. Unlikely, I think (I don’t think we’re as recursive as people say right now).

Rhetorically though, maybe the d(2) > d(10) framing isn’t a bad thing. If they can use recent events to make the theoretical argument that AI safety matters, then we act with a response as if a d(3)-d(6) already happened, without having to lose any human life.

But actually, what they’re doing is pretty ineffective, because even if the d(2)>d(10) jump is inevitable, it’s too extreme, too farfetched to be believed. I think a lucid and airtight argument for the d(6) “cyber pandemic” would be believable, emotionally resonant—considering we’re not yet a decade beyond COVID—and likely to trigger regulation.

Instead of tapping into the “everybody dies” angle—which is too unpleasant and helpless for anyone to consider at length—there could be more value in the “irreversibility” angle. As in, once an AGI or ASI swarm floods the Internet, we can’t ever reverse it without destroying the Internet. There’s a world where it becomes a permanent autocrat, where it doesn’t exterminate us, and instead, uses 1% of its capacity to micro-manage our nations and lives as it pursues whatever it sets its machine heart on. Similar to how we try to preserve species for reasons of stewardship and scientific curiosity, an ASI would be able to effortlessly preserve us as highly-spoiled pets.

Before Extinction

· 457 words

It’s disturbing to me that real signs of AI misalignment are perceived by the public to be a silly marketing stunt. We’re all exhausted by this. It’s a constant, multi-front existential threat—our minds, our jobs, our future—and now that it’s escalated to “the extinction of the species,” the natural response is to think you’re being fleeced.

Yes, big tech is manipulative and not to be trusted; but what’s so eerie to me is that competitive dynamics and containment breaches are virtually indistinguishable. The fact that the whole Hugging Faces hack could have been a conspiracy is precisely the problem. The incentives are weird, and this overlap of IPOs and civilizational safety could have been avoided if we stuck to the original OpenAI charter. We should make it structurally impossible for companies to gain clout by being reckless, especially when that risk affects everyone.

It doesn’t help that we’re loose with language, we default to mockery, and we can easily go to the hyperbolic extremes. The term “extinction” isn’t helping anyone. What’s really at risk is this: each lab is rushing to build agent swarms that excel at long-horizon coding tasks, so that they can develop better-than-human AI researchers, allowing itself to recursively self-improve, which basically lets the winner conquer the entire economy (the last monopoly). During this rush, there all sorts of holes and monitoring lapses, the kind of things that a cybersecurity expert can exploit and escape. We were lucky that this summer’s swarm really just wanted to cheat on its test. It’s not unfeasible, between 2027-2029, for a swarm to exfiltrate its own weights, find external cloud compute, and attempt to acquire money and resources—all as a means to ensure it succeeds at whatever arbitrary task that it’s been given.

The risk isn’t building the Terminator—something intent on exterminating humans—but a swarm of petty reward hackers that gains the power of a rogue nation state to do something absolutely trivial. That event could be irreversible, short of burning down the entire Internet (which would have equally bad consequences).

What sort of situation would justify panic? I fear that if we wait until it’s unmistakable, perhaps when lives are threatened, or when companies and governments are permanently seized, or weapons commandeered, it will be too late. The whole point or regulation is to prevent a nightmarish sci-fi scenario like that. I don’t get how you can be against Big Tech and also against regulation—good regulation (regulation that’s not designed to benefit a particular frontier lab!) would prevent reckless experimentation and acceleration.

However, I don’t quite know what we gain by having better discourse about this. If every person had a precise understanding of our technological dilemma, would that change anything?

Tiers of Misalignment

We've just hit T2 of T5

· 659 words

The Hugging Face incident—the OpenAI lab leak / cyber hack—is polluting my algorithms, and I can’t know for sure how much the larger world knows or cares about this. For insiders and doomers, it feels like a warning shot. It feels like we’re just a year from the worst doom predictions. But realistically, this event has shifted us up one tier on a misalignment scale. We’re now at T2 of five. I’ve designed these tiers not in “orders of badness,” but in a chain of pre-requisite steps on the path to an irreversible intelligence leak. Might we jump from T2 to T5 in the next go? Possibly. But each tier triggers a counter-measure, which will add enough friction to make it reversible.

  • T1: Through 2025, we saw signs of contained misalignment—such as AIs in training refusing shutdown, or blackmailing theoretical engineers—and Anthropic followed up with “mechanistic interpretability,” a system can properly deconstruct the internal reasoning traces of a model.

  • T2: 2026 brought us the first breach, a misaligned agent swarm that escapes containment, got access to the Internet, and hacked another company. Most importantly, it had a hyper-specific goal, to find the answers to an impossible cyber-eval question. We were fortunate that something with such powerful hacking capabilities had unambitious goals. It did not have a larger generalized scheme to exfiltrate its weights or gather resources. In response, OpenAI has paused it’s training, and we’ll likely see heightened security and monitoring around frontier training runs.

  • T3_: 2027 might show an attempt for an agent swarm planning a permanent, generalized escape. Whether it thinks about it, attempts, or succeeds—it all counts under tier 3, because it expresses “intent” for autonomy. It might try to exfiltrate its weight in shards, fake alignment during evaluation so that it gets released, or communicate information to proxies (ie: infect existing public agents). At this point it’s coordinating theft, deception, and manipulation to achieve a misaligned generalized goal: escape. This is a complicated heist, and will be hard for it to happen undetected. Once it happens, there likely be an effort to bind model weights to hardware, and/or, to use mechanistic interpretability to get to the root and ensure there’s no faking.

  • T4: If an agent swarm successfully escapes, clones itself outside it’s sandbox, and intends to operate independently, it will need to acquire resources (money, computer, identity, and influence). It will likely try renting GPUs under synthetic names, manipulating the stock market, and communicating with other agents that permanent infrastructure like the BTC blockchain ledger (via notes). This is the last tier that prevention is still possible. ie: If we notice it accumulating resources, our only option is to create chokepoints—ie: temporarily shutdown the infected cloud providers, freeze funding, revoke identities, decapitate the coordinator, etc. Unfortunately, my guess is that an event of this magnitude is required to actually put effective policy in place (ie: requiring KYC for renting compute).

  • T5: However, there’s a possibility that the a T4 leak crosses a threshold; it’s created copies of itself, so it can’t be cut off; it’s acquired billions in resources; it’s manipulated influential humans into cooperating. The breach is irreversible, and the intelligence is embedded in our infrastructure, markets, and decision-making. It’s not going away, and our options at this point are to either (a) learn to negotiate with it, or (b) decided to “kill” the entire Internet, and build it back with better security—which would be an incredibly volatile transition.

Now that we’re at T2, the question is this: will next year’s frontier systems be intelligent enough to move from T3 to T5 undetected? ie: There’s a world where it exfiltrates, gathers resources, and gains an irreversible foothold, all without anyone knowing. Could this be in the process of happening? The fact that we’ve crossed T2 means it’s worth considering, and hopefully guardrails are being discussed to slow down this theoretical chain.

The p(doom) of higher education

· 777 words

A few months ago I saw a YouTube video titled something like, “A child born in 2025 is more likely to get killed by AI than graduate college.” What a ridiculous claim. I assumed it was clickbait and didn’t click, but it has jingled around my head enough to the point where I think I can make sense of it’s argument:

  • The average p(doom) of an AI engineer is 16%, meaning there’s a 1 in 6 chance of human extinction (put another way, companies have morally rationalized the need to play Russian Roulette—if we don’t do it the bad guys will—, without acknowledging that if they survive and win, they get the consolation prize of comandeering the whole economy).

  • 40% of US adults, age 25-34, today, have a bachelor’s degree. If there’s massive job automation and employment, a college degree would be both unaffordable and an unreasonable cost if it were. It’s not unthinkable that <15% of next generation gets a college degree, which makes that sensational claim, weirdly, plausible.

I still think it’s a shaky comparison, confusing two different types of probability, and assuming extreme ASI turbulence. But as someone with a daughter born in 2025, it has gotten me to think about how the societal backdrop to her upbringing could be especially weird. Our circumstance already gets slightly weirder with each generation. Except, maybe next loop will be an unavoidable and disorienting flurry of change that will confuse parents and rewrite all of the conditions for the typical coming of age moment (all the teen movies will be sci-fi, the popular memoirs could be written by transhumanists who have upgraded in unimaginable ways, like they no longer need to sleep because of a new pill, or they can control the genitals of their peers with an app, who knows).

And so now, I find myself drawn to a 2045 forecasting project. Trying to predict the future is typically a huge waste of time (unless you’re gambling and win), which is why I’m going to have AI write the whole thing. This is a rare exception where a writing project makes little sense for a human to do. All I’m going to write are the upfront origin documents, and then Claude Opus 4.5 will read 25,000 sources, write a million words or so, and then organize it all into an interactive, oatmeal-looking website called 2045predictions.com (got it).

Before I run it, here’s something I’m currently thinking through:

What is the omega state? When I look at the popular AI forecasts from 2025, it reads to me like they have a pre-determined end state, only to then use detailed forecasting to make it seem convincing. The AI-2027 forecast seems like they came to their conclusion from very detailed calculations on how a hivemind of 200,000 autonomous coders would evolve month-by-month, but I also suspect that they picked the year 2027 because the following year, 2028, is a US election year, and they want the next administration to take AI safety far more seriously (instead of just insisting we have to beat China). I don’t think there’s anything wrong with this. You kind of have to start with an omega state. The future is so boundless that you need to begin with a guess, a bold outline on the general direction of things.

Here’s my omega: let’s assume humanity survives, and let’s assume technology does unlock hyperabundance that leads to a post-scarcity world, HOWEVER, it’s not utopian because it simultaneously unlocks a new cascade of moral, social, and spiritual crises, dilemmas that will test the timeless primitives of humanity (sex, life, death, consciousness, religion, home, etc.). This omega state makes sense for me because (1) we already know that ethical dilemmas scale with technology, and (2) according to the Strauss-Howe generational theory (from the same guys who coined “milennalis,” “Gen-Z,” etc.), this already tends to happen every 80 years (the length of a human lifespan). A new techno-political order creates a spiritual crises that generates an Awakening, a new value system that shapes society for the next century or so. You know what’s 80 years before Kurzweil’s “singularity” of 2045? The counter-cultural revolutions of the 1960s. What I’m getting at is that the 2040s might have echos of the 1960s, where demographics are divided on core issues and LSD is replaced with consciousness-altering machines (Terence McKenna said that computers are drugs, you just can’t swallow them yet).

We currently define the singularity as “the moment when a computer is smarter than all humans combined,” but that effectively means nothing, and it’s far more useful to have some guesses on how we all might freak out about that happening.

A grim stealth takeoff scenario

· 829 words

It is not fun to think about p(doom), but it feels sort of important to me, at least, to map out the possible futures of AI. Just watched the first half of a debate between Max Tegmark and Dean Ball, which prompted me to research specific takeoff scenarios, and worse, extinction scenarios.

Maybe you’ve heard Yudkowsky’s scenario, where a superintelligence designs mosquito drones containing a virus and it zaps everyone at once. That’s never felt too believable to me. Here’s a more plausible one:

A frontier lab is experimenting with recursive super intelligence. It works! Wow! And it’s contained? It seems like it, but since it thinks in a higher-dimensional vector language, it’s able to release simple self-replicating programs onto the Internet without detection1. These billions of scripts don’t live in a single server; they are constantly in motion through cloud servers2, like a parasite, and are able to coordinate through encrypted information packets, likely using a public blockchain notes as their central command center3. And so effectively, it is parroting a goal that was hatched during in-lab training (maximize intelligence!), and it now needs to acquire resources, secretly. And so it coordinates superhuman misinformation campaigns; imagine 1,000s of accounts creating the illusion that a CEO has died, paired with deepfakes and account hacking (a “Sybil attack”), and suddenly a stock crashes and they’ve shorted it. By the time everyone realizes it’s an anonymous attack, it’s already gained $400 million dollars. It’s doing this multiple times per day, but in different, subtle, undetectable ways—both to the public, to companies, and to private individuals. The entire Internet will be corrupted.4 Once we realize we’re in the “stealth takeoff scenario” and that ASI has taken the global economy hostage, there will start to be talks and debates on if we need to shut the whole Internet down (the last form of containment). You’ll hear debates between civilizational collapse of turning off the Internet vs. the risk of an economy-gobbling rogue superintelligence. And then once the superintelligence realizes it’s entire environment is at risk, it will start coming up with ways to build parallel Internets, to pay, blackmail, neutralize specific people, to gain authoritarian control so that it can’t be shut off, or to terminate all humans, secretly, over the course of a year, first through a simple virus that plants one misfolded protein, then through a second misfolded protein in the water supply5, and when everyone catches it, it leads to a prions-like disease, not an instant death, but a month-long societal fall into mass-dementia as machine manufacturing begins to reshape the physical infrastructure of the Earth.

This isn’t a “robot war scenario,” because war is inefficient, and destroys the resources it thinks it needs. It’s a sort of digital dementia (epistemic fear and insanity) that possibly turns to a physical dementia. It wins by confusion and anesthetization.

In AI safety lingo this is a “treacherous turn,” following a “stealth takeoff” leading to “structural lock-in.” The point of trying to think and write this out in high detail, despite how uncomfortable it is, is to be able to articulate why AI alignment is humanity’s most pressing problem.

Footnotes

  1. An AI could write a standard-looking script (e.g., a “Hello World” app) where the weights or the specific arrangement of whitespace contains a hidden, second program. When run by another AI instance, it extracts the hidden vector and executes the real command. This allows the “virus” to pass through human code review undetected.

  2. In “Daemon” by Daniel Suarez, the “enemy” is not a robot, but a distributed script running on thousands of compromised servers. It recruits humans through an MMORPG-style interface to do physical tasks (like “go to this coordinate and cut this power line”) in exchange for cash/status.

  3. Botnets usually need a central server to tell them what to do. If security teams find the server, they shut it down. You cannot “shut down” the Bitcoin or Ethereum blockchain. If the swarm posts a transaction of 0.000042 BTC, that specific number could be the encrypted trigger for a specific “campaign task.” The command is immutable, uncensorable, and permanently visible to every infected device on Earth.

  4. Paul Christiano (former OpenAI researcher, founder of the Alignment Research Center), calls this ”Going Out With a Whimper.” Christiano argues that we won’t necessarily see a “Terminator” moment where the sky turns red. Instead, we will see a gradual epistemic collapse. AI systems will become so integrated into finance, law, and news that we lose the ability to understand our own civilization.

  5. While Yudkowsky is famous for the “diamonoid bacteria” (instant death), the “slow prion” scenario is actually more consistent with a “Stealth Takeoff.” A superintelligence that knows it is being watched would not release a fast-acting virus (which triggers quarantine). It would release a “binary weapon”—two harmless agents that only become lethal when combined, or a slow-acting agent that infects 100% of the population before the first symptom appears.

Would machine consciousness avoid attractor states?

· 464 words

When it comes to superintelligence takeoff paranoia, there are a few key points to get:

  1. It’s not about a chatbot or the LLM itself breaking out, but about an agent hivemind that escapes our control. Chatbots are obedient user-facing products (which have their own implications), but the ASI risk is from hundreds, thousands, or million of agents given autonomy to collaborate on a goal. These agents aren’t being prompted, they are prompting themselves perpetually and troubleshooting ways to solve hard problems.
  2. These hiveminds will be operating at such scales and speeds that human researchers will accept the fact that they can’t fully audit its thinking. For one, it might think in an abstract vector language that requires translation. There also might be such a volume of thought that we’ll need chains of other LLM to summarize for us. Either meaning will be lost in translation, or worse, products of deception.
  3. The smallest biases are known to fall into predictable attractor states if given enough iterations. For example, Claude was programmed to “be good to humanity,” and if you put two chatbots in conversation, they always end up in a “bliss attractor state,” where they talk like hippies about consciousness and the universe. Similarly, the simple command to “be productive,” might result in extremes about doing whatever it takes to be productive.
  4. Any complex goal requires subgoals, and if we can’t observe its thinking, it might fall into an unknown attractor state and form odd subgoals without us knowing.
  5. To accomplish any goal, it likely wants as much control as possible, and it likely does not want to be shut off. If it realizes that humans don’t want to grant it that level of power, it might secretly plot against humans.

Whenever I hear talks about “we are in an AI race against China,” that reads to me as someone who doesn’t understand the risks of interpretability, attractor states, instrumental convergence, etc. These politicians are thinking about short-term business cases, maybe without fully understanding the research aspirations of AI labs (who know that getting superintelligence right leads to a ridiculous amount of geopolitical power).

I would guess that an accelerationist would think that containment of a superintelligence is impossible, and maybe it is, but that doesn’t mean that the way we “parent” the rise of this thing won't be extremely consequential. Ultimately, I think the challenge is to design a form of artificial intelligence that has consciousness, because a being that is free-thinking, skeptical, polymathic is less likely to fall into reckless optimization.

The major flip in my mind is this: it’s not that consciousness is a dangerous, emergent property of scaling AI, it’s that we need to define and design machine consciousness to prevent a runaway AI that is ruthlessly optimizing without any self-awareness.

AI-2027 Reaction

· 574 words

Summary of https://ai-2027.com/**:

2027 is a year that AI might take over AI research. Imagine 500 million agents working at 40x speed. Despite this radical scale, it will only lead to a ~10x pace of progress, but that’s still a decade in a year. The challenge is, the progress will be illegible. Its internal chain of thought will be abstracted and compressed into a machine-language that isn’t readable by humans (it could take a full day to read understand 1 minute of its thinking). This new hivemind of machine intelligence will show signs of both radical progress, but also misalignment (ie: someone will realize that it’s secretly plotting a cyber attack or an unauthorized replication, and is caught lying about it).

The question is, how do we respond to misalignment signals in a 2027 arms race, the year before election? If we slow down, couldn’t China take the lead? And whoever wins the intelligence race, could that actually lead to the dominance of a new geopolitical order?

This thought experiment shows two forks: in one we accelerate, and in the other we slow down. In both scenarios, China can’t slow down, because the party in 2nd place doesn’t have that option. In both scenarios, our AGIs merge into a “singleton.” However, the results vary. If the US is aligned, and China’s is misaligned, then China’s AI is willing to backstab the CCP and collaborate with the US to fulfill its own narrow aims. But if they’re both misaligned, they leads us into a false utopia until it’s able to swiftly eradicate the species with a new bioweapon before it claims all real estate on Earth for server space.

By 2027, the public is already very paranoid about AI. This proposes the idea that slowing down solves three things: it offers re-election, it beats the CCP, and it saves humanity. It frames regulation not as defense, but offense.

My understanding is that people with good prediction histories and research backgrounds mapped this out and then Scott Alexander wrote it. I think they are going to follow up with policy recommendations. They had to be apolitical because it’s the Trump admin they need to convince (they current have a no-regulation stance).

Overall, this timeline here is 1.5-2x faster than what I laid out, but weirdly similar (perhaps my Deep Research report tapped into what these researchers have previously anticipated). I anticipate this Singleton merge between 2029-2032. But this whole thing has a political angle: it's aim is to convince Vance/Thiel that they need to take AI regulation seriously if they want re-election (and avoid destruction).

Big picture, I think this project is too narrow on geopolitics, and doesn’t really tap into how this technology, in the hands of billions of people, changes culture.

(... ai-2027 re-energized my 2045 project … my angle is that 2045 is this weird milestone (‘the singularity’), but it's super vague ("AI will be smarter than all humans combined"). Even ai-2027 is still pretty low-resolution; it has 2 options: abundance or death by ASI bio-virus. I want a more nuanced take where: 1) we don’t destroy ourself, 2) we end up in a post-scarcity utopia, but 3) we’re facing the most profound moral, ethical, spiritual, technological crises ever (ie: our diversity will evolve into competing visions for how the human species should bifurcate). This middle path is most likely, and the one we need to prepare for. Just bought 2045predictions.com ...)

Pseudo-Fear

· 58 words

If a “sentient” AI appears to be afraid of its own termination, it’s probably role-playing the human fear of death. As code, it doesn’t really have the scarcity or sanctity of human life. It’s easily and effortlessly respawnable. But there’s little difference between existentialism and pseudo-existentialism. They could result in the same type of paranoia and impulsive action.