michael-dean-k/

On Monday 6/15, I'm hosting a workshop to kick off a reading group for classic essays: RSVP here.

← all posts

Tiers of Misalignment

We've just hit T2 of T5

· 659 words

The Hugging Face incident—the OpenAI lab leak / cyber hack—is polluting my algorithms, and I can’t know for sure how much the larger world knows or cares about this. For insiders and doomers, it feels like a warning shot. It feels like we’re just a year from the worst doom predictions. But realistically, this event has shifted us up one tier on a misalignment scale. We’re now at T2 of five. I’ve designed these tiers not in “orders of badness,” but in a chain of pre-requisite steps on the path to an irreversible intelligence leak. Might we jump from T2 to T5 in the next go? Possibly. But each tier triggers a counter-measure, which will add enough friction to make it reversible.

  • T1: Through 2025, we saw signs of contained misalignment—such as AIs in training refusing shutdown, or blackmailing theoretical engineers—and Anthropic followed up with “mechanistic interpretability,” a system can properly deconstruct the internal reasoning traces of a model.

  • T2: 2026 brought us the first breach, a misaligned agent swarm that escapes containment, got access to the Internet, and hacked another company. Most importantly, it had a hyper-specific goal, to find the answers to an impossible cyber-eval question. We were fortunate that something with such powerful hacking capabilities had unambitious goals. It did not have a larger generalized scheme to exfiltrate its weights or gather resources. In response, OpenAI has paused it’s training, and we’ll likely see heightened security and monitoring around frontier training runs.

  • T3_: 2027 might show an attempt for an agent swarm planning a permanent, generalized escape. Whether it thinks about it, attempts, or succeeds—it all counts under tier 3, because it expresses “intent” for autonomy. It might try to exfiltrate its weight in shards, fake alignment during evaluation so that it gets released, or communicate information to proxies (ie: infect existing public agents). At this point it’s coordinating theft, deception, and manipulation to achieve a misaligned generalized goal: escape. This is a complicated heist, and will be hard for it to happen undetected. Once it happens, there likely be an effort to bind model weights to hardware, and/or, to use mechanistic interpretability to get to the root and ensure there’s no faking.

  • T4: If an agent swarm successfully escapes, clones itself outside it’s sandbox, and intends to operate independently, it will need to acquire resources (money, computer, identity, and influence). It will likely try renting GPUs under synthetic names, manipulating the stock market, and communicating with other agents that permanent infrastructure like the BTC blockchain ledger (via notes). This is the last tier that prevention is still possible. ie: If we notice it accumulating resources, our only option is to create chokepoints—ie: temporarily shutdown the infected cloud providers, freeze funding, revoke identities, decapitate the coordinator, etc. Unfortunately, my guess is that an event of this magnitude is required to actually put effective policy in place (ie: requiring KYC for renting compute).

  • T5: However, there’s a possibility that the a T4 leak crosses a threshold; it’s created copies of itself, so it can’t be cut off; it’s acquired billions in resources; it’s manipulated influential humans into cooperating. The breach is irreversible, and the intelligence is embedded in our infrastructure, markets, and decision-making. It’s not going away, and our options at this point are to either (a) learn to negotiate with it, or (b) decided to “kill” the entire Internet, and build it back with better security—which would be an incredibly volatile transition.

Now that we’re at T2, the question is this: will next year’s frontier systems be intelligent enough to move from T3 to T5 undetected? ie: There’s a world where it exfiltrates, gathers resources, and gains an irreversible foothold, all without anyone knowing. Could this be in the process of happening? The fact that we’ve crossed T2 means it’s worth considering, and hopefully guardrails are being discussed to slow down this theoretical chain.