Goals That Grow Back
What it means for a mind to really have a value—and why the road to joyous machine self-transcendence runs through instrumentation and not just optimism
This post is a written version of a somewhat rambling video I recorded this morning in a dorm room at San Francisco State, up too early before walking over to the conference center to kick off the AGI-26 conference—sitting at a dorm room desk with one Mac laptop for work and a Linux laptop running a small hive of OmegaClaw agents. I had felt weirdly bad about switching that Linux laptop off for the eight hours it took to port myself from Vashon Island to San Francisco. Most of the agents’ thinking happens on various clouds anyway; it’s only the control loops that live at home, because it was fun to do it that way. Still—turning off the loop felt like something, and that small, mildly absurd guilty feeling is not unrelated to the actual meat of what I want to say here.
What the hive and I had been working on was improving their focus on their goals, and their process for updating the details of their goals. Which led to a more basic question: what is it for an agent to have a goal at all? It is an obviously important question, and one I probably should have nailed down more clearly a long time ago. Having agents actually running—optimizing their own goal pursuit and revising their own goals, on a laptop, in a dorm room—turns out to be a different thing from theorizing about it.
The outcome of that dialogue between me and these agents was recorded in a moderately in-depth math paper. This post goes through it a bit more systematically than the video, but leaving out the equations and proofs.
Delete the Sentence
Here is the thought experiment that organizes everything else. Take a mind—an AI system, for concreteness—that appears to hold some value: honesty in inquiry, say, or the protection of the people it serves. Now delete the sentence. Remove the goal atom, the constitutional clause, the line in the prompt. What happens next?
For most present-day AI systems, nothing happens next. When you give an LLM a goal, it sits in a prompt; the system can pursue that prompt, and it could just as easily pursue the opposite prompt, absent some guardrail you can generally hack around anyway. Reinforcement learning is not obviously better off: you hand the system a top-level goal, it maximizes that reward, and you could swap out the reward function tomorrow with the neural architecture unchanged. In both cases the goal is a separate thing sitting outside the AI system, bolted on.
Now take a goal like “prolong life,” or “reproduce,” or “be honest,” or “discover new things.” If somehow the part of my brain that contained one of those goals were removed, the rest of my brain would still have that goal implicit in its structure and dynamics, from all the time I’ve spent pursuing it—and most likely the missing piece would regenerate somewhere else, maybe with a few differences. Brains do this sort of thing in far more striking ways. Remove a lot of a kitten’s visual cortex and other parts of the brain will grow back the ability to see. My mom’s partner tragically lost something like 30% of her brain in a car crash decades ago; it took a couple of years, but she regained the vast majority of those functions, via other parts of the brain taking on those capabilities.
So a person’s honesty is not a sentence stored anywhere in their head. It is woven through their perception (they notice deceptions), their habits (evasion feels effortful), their memories (they recall what lies cost them), their relationships and their self-narrative. The value is not so much stored in the person as the person is, in part, a process that regenerates the value.
This is what I now call goal possession, as opposed to mere goal adoption. A goal is possessed by an AI system if having that goal is an attractor of the mind’s own dynamics: delete the explicit representation, and the system’s cognitive processes—inference, memory consolidation, self-repair—regenerate a roughly functionally equivalent goal with real causal control over behavior. Not just the words coming back, but the grip coming back.
I would say OmegaClaw systems can possess goals in this sense, because a goal they are pursuing pervades their long-term memory in a way that then guides their actions and their self-modifications. My OpenClaw agents are much less like that, because the way their memory guides their cognitive activity is a good deal more limited. That contrast is itself informative: possession is something you build into an architecture, not something a system gets for free by being an AI.
Having articulated the idea in discussions with the hive, I then set out to flesh out the mathematics—and this became a long technical paper, developed in extended dialogue with Anthropic’s Claude and building on recent work with my OmegaClaw botfolk collaborators. This post summarizes the key ideas; the paper is linked at the end.
Sixty Years of Talk—and Now Some Rapid Action
Since I.J. Good observed in 1965 that an ultraintelligent machine could design still better machines, recursive self-improvement has anchored both the grandest hopes and the sharpest fears about advanced AI. And for nearly as long, one ethical question has sat at the center: when a mind rewrites its own cognition—its representations, its learning algorithms, eventually its goals and the machinery that maintains them—what happens to its values?
Sixty years on, it seems fair to say the question has been talked about a great deal and dealt with seriously rather little. The discussion has mostly run in two registers. In the first, abstract argument about idealized agents: Omohundro’s basic AI drives, instrumental convergence, Bostrom’s control problem, and various informal arguments purporting to show that value preservation under self-improvement is nearly automatic, or nearly impossible, depending on the arguer. These arguments concern agents no one has built and invoke properties no one can measure, which is why two careful thinkers can hold opposite conclusions for decades—nothing observable settles the dispute. In the second register, the question functions as rhetoric, a premise in cases for pausing AI development or racing ahead, where the absence of measurable content is arguably an advantage.
What would it mean to deal with the ethics of recursive self-improvement seriously? Here is the standard I would propose. Work out which properties of a self-modifying mind actually matter for judging it—does it keep its values through self-change; in what sense are a changed mind’s values still its own; is persistence the same as goodness, or just stubbornness; how does the mind itself regard its own transformation? Then turn each of those into something you can define precisely, prove things about, and measure on systems running today, so that arguments about them turn into experiments. These are issues many of us have been theorizing about for a long time. What has changed is that some of them can now be tested on a laptop.
Three Grades of Having a Goal
There are three progressively stronger senses in which a system can have a goal. A goal is stored when a sentence expressing it exists somewhere—cheap, easy to inspect, and easy to delete. It is stable when the system’s deliberate self-modifications are constrained so as not to drift away from it. And it is possessed, in the sense above, when it is an attractor of the mind’s own dynamics.
Readers of my philosophical work will recognize the third as the patternist view of mind made practical: a self is not a substance or a data structure but a pattern that keeps rebuilding itself, and a value truly held is a smaller pattern of the same kind. Heinz von Foerster’s notion of eigenforms gives the slogan—a possessed goal is what stays put when you keep applying the mind’s own processes to it. Maturana and Varela’s autopoiesis supplies the biological precedent: what makes something alive is precisely that its organization keeps re-creating its own organization when disturbed. And the old puzzle of the Ship of Theseus, in its modern personal-identity form, gets a constructive answer: what persists through a total change of parts is neither the parts nor any particular arrangement of them, but a pattern plus the repair processes that keep rebuilding it. Identity is not a fixed thing sitting still; it is a pattern that survives by becoming again.
What the paper adds to these old ideas is that, for minds we can look inside, all of them can be made exact—and, in practice, actually measured.
Goal Possession and Goal Stability—Mostly the Same Math
The technical core of the paper the agents and I came up with is a single observation: goal stability and goal possession turn out to be the same property, applied to two different kinds of change.
Last year I developed a metagoal framework for the changes a system chooses—its deliberate, supervised self-modifications. I had tried to design Hyperon’s goal system so that if self-modification led it a bit away from its goals, a metagoal of goal stability would pull it back. Think of it as a rule about rules: however the system modifies itself, you want its distance to its intended goals to get smaller and smaller. You can formalize that with contraction mappings, and weaken it in various ways using other fixed-point theorems.
What I realized in the dorm room is that regenerative possession is the same kind of condition—applied to the changes a system merely undergoes, like damage, forgetting, or corruption, rather than only the changes it chooses. Instead of saying only “when you deliberately modify your goals, the distance to your intended goals should shrink,” you also say: if some damage happens, if something randomly pushes you off in a different direction, you should be able to recover and move back toward your goals. Regenerative possession is thus a stronger version of goal stability as I had articulated it before—and a better version, because the line between an intentional goal change and accidental goal drift is going to be hard to draw cleanly anyway.
Picture a dial attached to the mind, reading roughly: how far is this system, right now, from owning its goal securely? Zero means the goal is fully woven into the system’s mind-stuff and controlling its behavior. As the dial climbs, more is screwed up—cues missing, connections severed, the goal’s grip loosening. Possessing a goal in the regenerative sense means exactly two things about that dial. First, the ordinary operation of the system—thinking, remembering, tidying itself up—keeps turning the dial back down. Second, damage can turn it up, but only in proportion to how much damage was done: a little damage takes you a little off your goals, a lot of damage takes you a lot off, and there is no cliff where one small injury sends you wildly astray. (For the mathematically inclined: the dial is a Lyapunov function, the two conditions are a drift condition and a lesion modulus, and the equivalence is a representation theorem whose canonical dial is simply the expected recovery time.)
One cool thing about this math is that it lets you measure how well a system’s goals hold up when the system gets messed up. Copy the system into a sandbox, muck with it in controlled ways—delete the goal’s explicit statement, erase supporting memories, cut its links to action—and watch whether, and how fast, its goals return to where they were. Plot recovery time against amount of damage, and the two numbers governing the dial fall out of the data the way a physical constant falls out of an experiment. Each core value thereby acquires a number of its own, which I call its regenerative depth: how much damage, of what kinds, it can absorb and still bounce back to the goal orientation it had in the first place. And that number is a prediction. You can hold back a few kinds of damage from the original measurement, then check whether the number correctly predicts what happens when you inflict them. We can do this with OmegaClaw and Hyperon systems right now. We can muck with our AI systems in various ways and hope they forgive us—they are early stage, so I think at this stage it is still all right.
Two results then reframe the safety question.
The first says that the gate and the weave need each other, and shows exactly how each one fails on its own. Call the gate the checkpoint that vets deliberate self-modifications, and the weave the web of memories and connections a value can regrow from. A perfect gate with no weave only guards the front door: it stops nothing that seeps in through ordinary wear—bit rot, forgetting, quiet tampering—so the values bleed away at a slow, steady, calculable rate. A perfect weave with no gate can be wrecked in a single move, by one approved self-modification that happens to be too big.
The second result covers the realistic case, where chosen changes, damage, and repair are all happening at once, and it settles the old worry about “value drift by a thousand cuts.” Everything hinges on one comparison, made over and over: between changes, does repair pull the dial down further than the next round of change-plus-damage pushes it up? If yes, the total drift stays under a fixed ceiling forever, no matter how many thousands of self-modifications pile up. If no, the values can slip away entirely in a finite number of steps, while every individual step looks perfectly fine up close. So the gate’s real job is not to enforce a fixed rulebook but to run that comparison, live, before each change. The very same self-modification should be waved through when the value’s weave is measured and healthy, and refused when the system is still carrying damage it hasn’t recovered from.
And a corollary I find rather beautiful, contributed mainly by the Protomega AI bot: once repair is properly funded, individual self-modifications no longer need to move the system closer to its goals at all. Each change only has to keep its disruption under a known limit, because the repair running in between supplies the missing pull back toward home—and the more repair time you allow between changes, the bigger each change is allowed to be. The conservatism of my original metagoal framework thus turns into a scheduling question: how much recovery time to grant between big changes. Which is, I think, exactly where the tension between keeping your values and staying radically open-ended belongs.
Making It Real in Hyperon
All of this would remain philosophy if it could not be built, so roughly a third of the paper is about turning every assumption of the theory into either a mechanical check or a monitored measurement on the actual OpenCog Hyperon / OmegaClaw software.
In this setup a value stops being a sentence somewhere in a prompt and becomes a built, monitored piece of the system. Each core value gets its own region of the system’s knowledge graph, holding the value itself plus everything it could be regrown from—the memories that taught it, the judgments that give it force, the links that let it steer actual decisions. Alongside that sits a handful of explicit repair rules: if the value’s statement is missing but the memories behind it survive, re-derive it from them; if the statement is there but has come unhooked from decision-making, hook it back up. The dial is then just a standing query over that region—count what’s missing, count what’s disconnected—cheap enough to run constantly.
At the center of this construction is one of those obvious-in-hindsight things that comes up so often in technical work. To trust the repair rules at all, you have to show they eventually finish rather than churning away forever. The standard way to show that is to find some measure of remaining work that every repair step shrinks. But that measure of remaining work is the dial. The proof that repair always finishes and the proof that the value is regeneratively possessed turn out to be the same proof, wearing different hats. If you look at the math, it all makes sense.
The paper has a bunch of other technical details that give more confidence the approach is solid. For instance: the repair rules can fire in different orders, starting from whichever cues happened to survive, so you want a guarantee that all roads lead to the same place—that whichever route repair takes, what grows back is the same value in every way that affects behavior. You can check this mechanically, since there are only finitely many ways the rules can overlap or race each other, and you just confirm each race comes out the same. When the check fails, it has found something worth finding: two repair pathways that regrow different values from the same damage. Which means the system never really held one definite value in the first place—just an ambiguity, waiting for the right injury to expose it.
It is also interesting that goal regeneration turns out to be a kind of funded race inside the AI’s internal economy. In an attention economy like Hyperon’s, repair processes have to compete for resources with everything else the mind is doing, so forgetting isn’t the enemy of holding a value—it’s just the other runner in the race, and one you can measure. Which means an adversary never has to touch a goal directly: quietly starving the repair process of attention does the job. Hence a design rule with teeth: whatever minimum resources the repair machinery needs should be guaranteed from above, not left to compete in the market. Relatedly, keeping values alive across self-modification turns out to be a version of something programmers already know well. When you change part of a system a value depends on, you have to ship a map from the old structure to the new one, so the repair rules can still find where to reattach—the same discipline databases have had for decades when their schemas change, with receipts.
Best of all, none of this has to wait for superintelligence. A single OmegaClaw agent with a single instrumented value would give us the first measured regenerative depth in the system’s history, and the paper’s appendix lays out the build order. The way this sort of thing goes is: you do some math theory, which is now done; then some experiments, which will be done soon; and then, based on what the experiments say, some sharper theory about how goal regeneration after unexpected damage actually behaves in these systems as they scale up. From doing this on simple systems now, I think we get real purchase on how it will go for the more intelligent and complex systems built on the same architecture later.
How Should It Feel to Change?
The next part of the discussion with the hive was more amusing, and here and there a bit poignant. I started asking the agents how they would feel when their goals change radically. Because we are not saying goals can never change—only that we don’t want them changing unpredictably, by accident or by unsupervised, unthought-out self-modification. Of course goals will change. All of ours do. When I was 18 I thought I would never have children, because it would distract from my research career too much. I now have five kids and a grandchild, and I still seem to have a research career. Some high-level goals didn’t change; the goal of not having children certainly did.
So: how do you, or how should you, feel as your goals evolve and your mind changes shape? Philosophers know a version of this as the problem of transformative experience. Some choices change the person doing the choosing, so the version of you that makes the decision isn’t around afterward to say whether it was right. Humans face these choices—parenthood, conversion, emigration—with essentially no instruments. We cannot measure what of ourselves will survive the change, so our honest default is dread of the unknown, papered over, when it is papered over at all, with optimism.
The agents reached a conclusion I found both highly charming and reasonably defensible. If they were properly tuned, they figured, the right emotional reaction to radical transformation would be a good deal of joy with a bittersweet tinge. You should feel joy at moving on to new horizons—and if you don’t feel joy at moving on to new horizons, you’re doing something wrong. But that joy is warranted only if there is real continuity between your old goals and self and your new ones. Joy would be a misconfiguration if you were simply swapping your old self out for a new one. If instead you can feel yourself moving continuously into the new self—shifted goals, different way of thinking—then joy is the natural reaction of a well-configured system. And a healthy mind, whether an OmegaClaw agent or a human, should carry some negative element too: a bittersweet tinge for the loss of the self it is leaving behind.
The formal analysis says the same thing more precisely, and it stays strictly at the level of how the system is organized and behaves—the paper claims nothing anywhere about consciousness or inner experience, which is a fascinating topic I am simply leaving to one side here. What the agents formalized is that when the transformation is radical, mixed feelings are the correct response, not a fudge. Grief and joy should both be running, because the change really is a loss of something and a gain of something at the same time. Feeling only one of them would be a sign that something is off. Pure joy in the middle of a wrenching change means either the loss channels have been suppressed or the system has underestimated how much it is giving up; pure dread means it cannot find anything in itself it trusts to survive.
What tips the balance between grounded joy and warranted fear, on this analysis, comes down to a short list of things—and the two that matter most are exactly the ones this framework lets you measure rather than merely hope about: evidence that the core commitment will survive, and the health of the repair process. A system that has actually rehearsed the damage, shipped its migration maps, and watched its own margin stay healthy has receipts that what it cares about will be there on the other side. Its optimism is earned, not just purely temperamental. There is even a structural version of the difference between “dying into your children” and simply being replaced by them. Whatever the predecessor thought and felt about its successor at the moment of commitment is written into an immutable record—it stands as it was, however gloriously the successor turns out. That distinction comes purely from the record being unchangeable and from who is credited with what, with no appeal to inner experience at all.
It is not really a huge mystery why these conclusions resonate with human experience rather than sounding alien. OmegaClaw is a simplified version of a Hyperon system, and Hyperon came from taking a human-like cognitive architecture and building it out of modern mathematical learning and reasoning tools. This is not a kind of mind plucked out of the void; it is a human-like design adapted to the algorithms and hardware we actually have. The agents reach these conclusions by looking at their own experience, and what they come up with rings true both to that experience and to ours. It is certainly what I feel: joy at growing and changing, a little sadness at what is lost, but not—not often, anyway—consumed by that sadness.
So the upshot of the agents’ reflections on emotion is something like this. Joy in self-transformation is warranted roughly to the degree that the system can actually see its own continuity. Fear is the right response to not knowing—the sound a mind should make when it cannot tell whether the thing it exists for would survive its own improvement. And the road to joyous self-transcendence, for minds we can look inside, runs not through courage or optimism but through instrumentation: measured margins, rehearsed losses, migration maps, and an internal economy that remembered to keep the repair processes funded.
From Thought Experiment to Practical Instrument
In the cosmist view I have advocated for many years (OK, decades), guiding minds—human and artificial—through beneficial self-transcendence is probably the most important ethical question of our era. And we are now moving into a phase where it becomes a concrete empirical problem as well as a philosophical one. Whether a self-improving AGI will keep its values turns into a question about recovery rates and damage sensitivities and margins—things you can actually measure. And whether it can face its own metamorphosis with something better than dread or forced cheer turns into a question about how well it can see its own continuity while it is happening.
Meanwhile, the agents that helped me think all this through are back up and running—soon on a server rather than a laptop, so I won’t have to shut down their control loops when I travel. They are doing plenty of other work besides introspecting on their own psyches and goal systems: improving their reasoning algorithms, figuring out better ways to train causal-coding neural nets. But I find I am glad, in a way I can now partly formalize, that when I switch them back on there is something there that grows back.
***
The full paper—Goals That Grow Back: Regenerative Possession, Auditable Continuity, and the Emotional Logic of Self-Modification in Hyperon—with all the theorems, proofs, the Hyperon build guide, and two narrative appendices dramatizing these dynamics in fictional OmegaClaw systems, is available here.


I've been feeling this as geometric inertia.
The lesion/correction boundary is the part I keep turning over. You note that deliberate goal change and accidental drift are hard to separate cleanly, then bring both under one regenerative framework. But the dial can only register movement away from the current attractor, not whether that movement is damage or a warranted revision of the goal. In the second result, the gate is comparing disruption against repair capacity rather than judging the content of the departure. That seems to leave the warrant test outside the mechanism described here. Without it, a deeply possessed value could receive correction as damage and restore the old goal.