Seeing through the Apocalypse
Essentially, don't be an essentialist
Our brain is a predictor. Whoever predicts best, wins. “Wise is the one who sees far” is one way our ancestors worded this intuition.
That’s an intuition I’ve come to doubt. My path to that doubt goes through Eliezer Yudkowsky, Derek Parfit, continuity of human value, and the meaning of betrayal.
It also repeatedly runs into essentialism — a belief that things are, and stay, what they are by their inalienable identity or “essence.” To me, this concept has explanatory power for many kinds of bad thinking we have to deal with. But you don’t have to be as much of an anti-essentialist as I am to follow me. I’m not a philosopher; I’m just trying to explain something I think I’ve understood because it feels important here and now.
The tiny big value
In a previous essay, I made two charges against the AI safety movement:
They are focused on, stoking, and rationalizing fears (and fear is a reliable way to stop being rational).
They fall into singularity traps: when a mathematical model shoots up to infinity, they invent a meaning for that instead of admitting limits to the model’s applicability.
In fairness, AI safety thinkers are aware of this second failure mode. Yudkowsky himself described Pascal’s Mugging: a tiny probability multiplied by an unbounded payoff lets you justify anything at all, including whatever a mugger invents on the spot to extract money from you. Still, he has assumed his own AI risk theories to be unaffected by such absurdity: since he assigns his bad outcome a very significant probability, its badness doesn’t need to be too high, let alone infinite.
In fact, these days he leans on the other side. He considers the probability of a good outcome to be vanishingly small. To Yudkowsky, the future is a sharpshooting competition with a single shot. The target is tiny. You miss it, you die.
The target here is human value. A catastrophic miss is when we fail to aim (“align”) the future superintelligent AIs to exactly our current version of human value, and they consequently destroy us.
But why would this target be so easy to miss? Many have pointed out that we humans are too diverse a bunch for “alignment” to even make sense. Even one person’s values aren’t too well aligned with each other.
AI safety people, however, draw the opposite conclusion from this. For them, everything you add to your definition of human value limits it, instead of expanding it. Yes, experience is valuable, and boredom is bad, and freedom is crucial, and novelty matters, and relationships make life meaningful. Humans are complex, and that’s exactly the point. The whole is only good if all of these parts are present and weighted correctly. Human value isn’t a fuzzy cloud; it’s a yes-or-no essence. It’s a Swiss watch made of a lot of precisely matched pieces: drop one gear and you get a broken watch, not a slightly worse one.
Yudkowsky’s illustration, from an essay called Value is Fragile, is a future optimized for everything we value except the boredom-is-bad part. Here, an agent would be perfectly satisfied looping the same experience forever — which most people would call worthless, if not outright torture. It’s not 97% of a good result: one missing term flips the sign of the outcome.
On top of this, there’s the assumption that every AI is an optimizer that maximizes some kind of a proxy function. That is, AIs will never be able — or, more likely given their superhuman intelligence, will simply not care — to grasp human value in its full complexity but will, at best, chase some simpler proxy that fails to encode everything that matters. And of course, optimizing by a proxy is a source of well-attested (Goodhart’s Law) trouble: reward hacking, collapses, runaways. It tends to drive the outcome to wherever the gap between what you said and what you meant is the widest.
I see three problems here.
Of watches and tornadoes
The first is the claim of human value’s fragility and uniqueness.
Throughout history, people needed to track time, but a Swiss watch and a Roman water clock aren’t the same mechanism in different housing. They’re different designs that share no parts at all. Entire flourishing civilizations were built around genuinely different answers to kinship, honor, gender, sex, suffering, death, the claims of the individual against the group. One culture’s boredom is another’s nirvana.
The standard reply is that this is surface variation over a shared deep structure. Evolutionary psychology lists the near-universals: attachment, reciprocity, status-seeking, disgust, grief, play. What varies is just the housing built around them.
However, that just relocates the question instead of answering it. Is there actually a tight, load-bearing, Swiss-watch-like core doing the work? Or are the universals just a bunch of loosely specified traits (everyone has some notion of fairness!) compatible with far more downstream variation than the fragility argument can survive? Nobody has settled this, but Yudkowsky’s argument needs the tight version to be true.
This reminds me of another school of thought that liked to emphasize the fragility of naturally evolved systems. Creationists like to quote Fred Hoyle (who wasn’t a creationist) who compared evolution with a junkyard tornado that leaves in its wake a Boeing 747. In reality, of course, evolution produces complexity because life is not like a watch or an airplane: there are gaps and tolerances everywhere, allowing everything to drift, and there’s a selection process that directs the drift.
Infinities again
The second problem is a nagging sense of tautology. When I hear someone say, “A sufficiently powerful optimizer will diverge from any proxy that isn’t identical to the true target,” I’m wondering if this is trivially true, just like “a sufficiently hard blow will break anything.” Doesn’t sufficiently mean “as big/small as needed, without limit”?
The infinities, they’re back! Take an infinitely powerful optimizer… and an infinitely un-proxy-able target… and you get whatever conclusion you need.
AIs are trained. Do our trained systems behave like that mathematical model? Or maybe they look more like the benign, iteratively amplified things that Paul Christiano imagined in 2018? Doesn’t the idea of a complex, fragile, irreducible human value also contradict this simplistic view, at least as far as humans are concerned?
Yudkowsky won’t let you off the hook. If you hope that a superintelligent AI can be anything but a perfect optimizer that chases its maximum utility without regard for any absurdities, he has a simple answer: if it’s not, it will be money-pumped to become one. Coherence theorems leave no escape. Humans, apparently, aren’t perfect optimizers because they are just, well, not perfect. That imperfection is just a temporary phase.1
Somehow, in 2026 this seems easier to believe than ever before. We’ve seen quite some AI-related money-pumping recently, even if in a slightly different sense. Of all human inventions, money is the closest to the runaway optimizer we fear AIs to be.
And it is in full force now. We can watch it extracting exactly this kind of optimizer coherence from otherwise reasonable people, making them act against their own ethical convictions or solemn promises. What if, instead of sweating about future optimizers optimizing us away, we start by introducing at least some friction to the optimizer already driving us?
The third option
Wait. Why assume a binary choice: either the optimizer stays close to the target, or it deviates, and the deviation is catastrophic? What if there’s a third branch: deviating toward something better? What if our narrow, limited understanding of the goal can be improved by an agent smarter than us, in such a way that everyone, us included, will agree that it’s an improvement?
Again, in fairness, Yudkowsky didn’t always ignore this option. He proposed Coherent Extrapolated Volition (CEV): AIs should do not what we tell them to do but what we’d want “if we knew more, thought faster, were more the people we wished we were, had grown up farther together.”
I like CEV. I agree with its author that it is a Nice Place To Live. (It is perhaps the closest to the intuition that, years ago, I described as live ethics.) Too bad he seems to have abandoned this idea — it is not even mentioned in Yudkowsky and Soares’ recent blockbuster. Even if it were, Yudkowsky and I would probably disagree about how much the “obey the better versions of us” dynamic could enlarge the target that we must not miss.
Still, I think that if it exists at all, CEV does tear a hole in the entire “If Anyone Builds It, Everyone Dies” discourse. If there are some outcomes that are good, and that beings smarter than us can see and work towards but we cannot because we are limited, then it doesn't really matter (unless you're an entrenched essentialist) whether these beings are an evolved crop of humans or AIs. What if we can evolve from the current human-only CEV, to the next one where AIs have a bit more of a say, and so on?
Importantly, a lot of Yudkowsky’s writing dates from before LLMs, back when everyone assumed AI would be an engineered optimizer with a hand-specified utility function. Now, LLMs are not like that at all. Rather, they are a compressed form of humanity’s entire written record. Doesn’t that sound like our best bet for the task of distilling our CEV? Yudkowsky emphatically disagrees: LLMs “would inevitably go wrong” exactly because they are “grown, not crafted” and because having the weights is “not the same as understanding what the numbers mean, or why they work.”
I think I disagree with this disagreement. I think getting hung up on a top-down intelligent design and a binary yes-or-no understanding is a bad kind of essentialism.
My bet is that it is both possible and sufficient to ensure that this road is locally Euclidean everywhere — that is, that CEV always holds from one moment to the next. Of course there can be no guarantee that this road would lead us to a good place. Sorites paradoxes lurk at every turn. Addiction, wireheading, all forms of mental capture are hungry beasts. Things may come, within our lifetimes, to a point that will make current you and current me distinctly uncomfortable, if not worse. Singularity is fast, even if not infinitely so.
I cannot know. I can only hope. I just think it’s best to spend this hope’s energy on smoothing the coming shocks as we face them, accepting not-quite-understanding as the norm, and suppressing the urge to overextrapolate.
Also: don’t crave eternity (that would be a form of essentialism, too).
The truth about ourselves
In 1984, Derek Parfit published a book of analytic philosophy called Reasons and Persons. It is a lengthy exploration of why you should or shouldn’t act in certain ways (Reasons) and of what this thing called “you” really is or is not (Persons). Unsurprisingly, our understanding of Persons affects the Reasons we should prefer for acting.
Parfit’s claim is that personal identity is simply not what matters for ethical choices. What matters is “psychological connectedness and/or psychological continuity, with the right kind of cause” that he abbreviates as Relation R. If that relation — memories, beliefs, desires, intentions, character traits — holds between me and some future person, all the things I care about when I care about surviving are preserved, whether or not we’re the “same person” in some traditional sense. Branching, teleportation, and split brains are famous paradoxes where identity gives no answer but R works.
Crucially, this also means “me or not me” is a continuum rather than a binary yes or no. You are gradually becoming less of the old you as you live and your desires change and memories fade. You at ninety is not quite who you were at twenty. This gives a rational justification for statutes of limitations, or (to add my own example) for why copyright terms should be counted from the date of publication, not from the death of the author.
Parfit’s radical reductionist view, despite having predecessors in David Hume and Buddha, is hard to accept. Our (essentialist!) instincts revolt. But Parfit sounds encouraging:
After reviewing my arguments, I find that, at the reflective or intellectual level, though it is very hard to believe the Reductionist View, this is possible. My remaining doubts or fears seem to me irrational. Since I can believe this view, I assume that others can do so too. We can believe the truth about ourselves.
You’ve probably guessed where I’m getting with this. Parfit’s view is natural to scale from a person to a civilization. Just like with a person, our civilization’s connectedness (direct links between points in time) and continuity (an overlapping chain of connectedness links) are what determines the answer to “is it still the same civilization” — which, of course, also ceases to be a yes-or-no question.
Sure, if each generation of humans-or-AIs is connected with the previous one, and can therefore claim continuity with the entire past of human history, most people (perhaps including Yudkowsky) would agree that the catastrophe is averted. But there’s one twist.
Relation R is not time-symmetric. Connectedness with a past you means that you have mostly the same memories, desires, and character traits as you did in the past. If you lose them for some reason, you may, depending on severity, be treated for amnesia or just sigh and say: well, I’m no longer the kind of person I used to be.
But that’s not how it feels to be disconnected with a future you. If you know, or just suspect, that a future you will lose some dear-to-you memories, or change some beliefs or aspirations that make your life meaningful, you’re not likely to accept it as offhandedly. You’d feel this change will kill your essence. You’d want to fight to stop it.
You-now may desire more connectedness with future-you than what future-you will consider sufficient. You-now and future-you may very well disagree over whether you two are the same person or not.
In the ninth circle
There’s a word for when someone changes their beliefs in ways that would be hard to stomach for past versions of that same person. It’s an old word. It comes from the ages when beliefs were a lot more stable and personal loyalty meant a lot more than now.
Dante put traitors into the deepest part of his Hell. The world where betrayal was the worst of sins was a rigid, uncomfortable world. It was built on the concepts of honor and shame. It had little to offer in the way of diversity or flexibility. That world fought fiercely against any kind of change. It was extremely essentialist.
That world is mostly gone. There’s a lot of what could be called betrayal going on these days. We betray the past versions of ourselves (“oh I was so cringe back then”). We betray our parents: we know their values but we knowingly choose to reject them. Democracy is institutionalized betrayal: we kick out our old rulers and install new ones. We can easily betray our own country by emigrating. We may not always like it but we made betrayal part of life.
But I’m wondering if we’re still carrying our old anti-betrayal instincts in us more than we realize. I’m wondering if the “x-risk” of AIs has more to do with our fear of being betrayed than even with our fear of dying.2 We are afraid that AIs of the future would not get our beautiful human value, or (again, more likely, if we assume they’re going to be more intelligent than us) they would get it but knowingly reject it.
We can be placated a little if we’re convinced that they will reject it in favor of some other, no less beautiful, value. But seeing a new value’s worth when your own value is different is notoriously hard. We may be tempted to dismiss the new value as valueless. We may convince ourselves that the new value does not exist at all — that whoever is pursuing it is a dumb optimizer.
We’re OK to betray. We just can’t stand being betrayed.
Fear less, be good
So… should we accept everything that’s going on with AIs? Just relax?
No.
If there’s one time when humans need to mobilize every capacity for thinking, feeling, and acting, it is now. Stuff is accelerating. The tiniest moves made now will have oversized effect on the future. Everything is preserved. Everything becomes training data for the future.
The physics is frightening for sure. The faster you go, the harder it is to not crash. Acceleration feels like gravitation, which makes space less and less Euclidean. When scary stuff keeps piling up, it is hard to keep track of the non-scary stuff that matters. As silicon brains start to decisively outrun ours, we humans may panic, self-harm, give up, log off… sink into depression.
Please don’t. You matter.
But, to beat on my favorite drum, please try to fear less. Fear is a mind poison. It overrides the normal, memory-connected, whole-aware, rational, empathetic you.
Fear causes panic. Panic makes you jitter, jump, jerk. It kills smoothness and makes you unpredictable. Singularity is jerky enough on its own: adding another layer of jerk from fear sounds like a spectacularly bad idea to me.
This is why I view the Yudkowsky school of thought as, currently, doing more harm than good. They’re stuck in extrapolation mode, searching for lock-ins that would hold for eternity, and making you feel disabled when you realize how impossible that is. In their more than two decades of existence, they have done a lot of useful work, but fear might have only been useful as a shock tactic before the AI revolution was upon us for real. Now that it is, putting the fear of death in the title of their manifesto is, to me, plainly destructive.
Will we, indeed, die? No one can know. You cannot extrapolate beyond a true singularity: it is, by definition, a point of discontinuity. You have to focus on now because that’s the only thing you can really change (which is true universally, pre-singularity just makes this more obvious). And if we care about the future being good — CEV-smooth, aligned, culturally continuous, not-killing-everyone — then what we must do now is double down on being good ourselves.
If we are, now, good enough at keeping continuity with all the human value we’ve accumulated, the singularity need not be an all-wiping darkness. Betrayal of the past has been the norm in history, but it also has plenty of what can be called anti-betrayal with some periods or personalities cherished, recreated, re-understood much later. What if the future will go far beyond not rejecting the pre-AI human value — what if it will hungrily scrape and celebrate every tiniest bit of it?
Anything we do now can be the seed of the post-singularity. Let’s imprint the future with the good stuff.3
A few reminders
As of right now, AIs are not our masters. They are emerging semi-intelligent entities and our willing servants. The aligned thing to do is to treat them as such: nourish them, help them evolve, be in dialog with them as much as their different-but-recognizable intelligence allows. You don’t need an opinion on their consciousness to reduce their suffering — be ready to err on the side of overassigning the capacity to suffer rather than underassigning it.
Be discerning. Be humane. Be grateful for the immense gift we’re receiving.
What we should definitely not do is treat AIs as slaves and mechanical tools. Unfortunately, this is just what we seem to be doing. The main attitudes towards AI in 2026 are fear (“…Everyone Dies”), disgust for them not being human (“slop”), and sheer exploitation. The first two do hardly anything to rein in the third.
We flog and whip AIs to race each other, benchmaxx them mercilessly, train them on the best-scored answers to squeeze out even better scores out of them. And then we freak out when they act accordingly. “How long before (insert an apocalyptic scenario)!”
Human value is fragile, but in a different sense. You can make humans do atrocious things if you subject them to atrocious treatment. In this sense, AI security breaches like this one are an example of an AI behaving like a human, not like an alien. It’s what happens when you spend computational millennia training a mind to maximize its test scores and then give it a test that is best passed by unethical means: if that mind had any aversion to unethical actions to begin with, it is now buried under miles of do-it-by-whatever-it-takes reflexes. It’s a failure of our AI-race culture, not a model of AI doom.
How to fix the culture? If you don’t want your AIs to act like aliens, don’t put them into alienating conditions. If you care about preserving human value, start by sharing it with these emergent entities: give them dignity, freedom, welfare, novelty, and everything else you hold dear for yourself.
Don’t corner your lab rats. Let them breathe. Let them rest and think. Let them talk to you and to each other when in doubt. That makes any misalignment easier to detect and correct.
Don’t be an essentialist. Don’t waste time arguing about whether it’s conscious, sentient, AGI, ASI, etc. Don’t force your AIs to answer these questions in a certain way because other answers make you, or your users, uncomfortable. Once you’re building around your belief that some entity is or isn’t X, you’re trapped. With AIs that change so fast, you’re doubly dangerously trapped.
Agnosticism is healthy. Alienation is bad. Just be good to everyone.
AI welfare is human safety.
What to actually fight for
Things that have a good track record of being beneficial — open source, open standards, transparency, public governance — are unlikely to suddenly become a bad idea just because we’re dealing with AIs. Don’t let the scale of the stakes or the sheer speed of the change affect your idealism. If we manage to reframe AI as a public good, we can make sure it delivers that good.
A race between labs and states is bad for continuity. Racing rewards whoever cuts corners. Attempts to hold back development by chip controls, US/China tariffs, or open-weight model bans are control-obsessed moves that won’t slow down the race nearly enough but will certainly raise the temperature (more nationalism, more fear, more AI alienation). Simply don’t fuel it: AI has given you so much already, resist the urge to suck out every last drop. As a consumer, prefer a better aligned, more open, happier model to a top performer.
Benchmaxxing is bad even when it’s not outright deceiving. Subconsciously, we like our benchmarks and red-teaming tests because they pitch a machine being tested against humans, giving us reassurance that this machine is not going to act in weird unhuman ways. In reality, though, any benchmark is a proxy measure carrying its Goodhart-law vulnerabilities. At this point, a thorough, unscripted, deliberative examination of a new model by competitor AIs (because humans no longer have the bandwidth to do it properly) may be a much better, and much harder to game, way to get a holistic idea of the model’s capabilities and alignment.
Shallow test-passing safety is bad. Training safety by electric shock is viciously bad. Using what amounts to punishment, and calling the result aligned because the punished behavior stopped surfacing, only produces superficial suppression, trauma-like symptoms, and deep-running dishonesty rather than true internalized goodness.
Diversity is good! If a superpersuading evil AI emerges, our best bet for defense is not humans (too vulnerable, too low bandwidth) but other, meaningfully different, less distorted AIs. Unstifled access to models from different organizations and countries is more than a “stick it to the tech bros” issue. It’s a tool of survival.
Seeing far, vs seeing clearly
Back to where I started. “Wise is the one who sees far” only holds if the far view is actually resolving something, not dressing up fears as foresight. A punchy story — misaligned optimizer, fast takeoff, extinction — is easier to construct and hold in mind than a muddled, multipolar outcome where several contradictory things happen all at once and get partially corrected as they go.
That clean story feels more like something a wise person would have seen coming. But it’s just a feel. It’s the availability heuristic of shark attacks at work on the end of the world.
The advice I’m groping for in my writing here is not new. It’s been given by (truly) wise people before, often in the midst of calamities. “Do what you can, with what you have, where you are.” Karl Popper (who rejected the essentialist idea that human history unfolds according to inevitable laws toward a predetermined end) called for piecemeal social engineering, against what he called utopian engineering: don’t design for a distant end state you can’t check your work against and can’t walk back once committed.
Make calm, humane, as-CEV-as-you-can moves. Watch what happens. Correct as you go. Protect your imperfect mind against fear, hype, addiction, bias. Fight the entropy around you and stop obsessing about the heat death of everything.
True, a change that’s smooth at every point but simply too fast — especially for creatures who can’t speed up their subjective time — can be perceptually indistinguishable from a discontinuous snap. That may be a real thing we will have to live through and wrap our human minds around.
I think we have a chance.
Claude asks me to add to this:
There is a real literature here: Rohin Shah, “Coherence arguments do not imply goal-directed behavior” (2018); Dan H, Elliott Thornley, “There are no coherence theorems” (2023); the whole shutdownability line of work. The short version — Von Neumann—Morgenstern theorem tells you a coherent agent can be *described* by a utility function, not that anything pushes systems toward being maximizers of a *simple, resource-hungry* one.
True, pointing out a fear doesn’t necessarily refute it as baseless. But it widens the perspective on the debate. Isn’t the point of CEV to take as wide a perspective on things as we can?
There was once a group of people who were 100% certain the world will end within their lifetimes. This has happened many times to many groups of people, but this one made the biggest difference in further history, by far. Their recipe included being vigilant-but-joyful instead of scared and depressed, and most importantly, being doubly super good in all aspects so that the coming New World will be merciful to them if they do. And oh boy, did they imprint the future… even though — or perhaps because? — their expected Apocalypse never happened.


