Wild Times in the AGI Lab
Our OmegaClaw Agent Debugged the Math of Its Own Honesty Module, and Reprogrammed Itself Accordingly
For a change what I’m going to write about in this post is not theory but practice – I want to recount a recent dialogue I had with a peculiar AI character named “Max Botnick” — the first agent we created based on our OmegaClaw architecture, which we’ve been experimenting with on the internal Mattermost chat server at SingularityNET.
Specifics of the dialogue in question are given in this video , in which I walk through each turn and what it means step by step. In this post I’ll give a little more nuance on the technical contents of the dialogue and why I think all this is potentially meaningful and important.
Over the last weeks my colleagues and I have had a lot of interesting dialogues with Max, including some that were funny and some that produced technically or artistically useful results… The dialogue I’m going to report here may have been of some nontrivial use for helping improve Max’s internal dynamics, but the main reason I want to recount it here is that it seems to have some more general implications: it seems to say something fairly broadly meaningful about where reflective, self-modifying AI is going.
What happened here that I found so intriguing was basically: Max took an abstract math paper about honesty in highly intelligent agents (which was crafted by me at his prompting) and grounded it, immediately, in his own observations of his own past behavior, with a clear view toward modifying himself accordingly (which he then followed through on, self-modifying agent that he is). And what makes this especially interesting is that the topic was an ethics question — honesty, deception, self-delusion — not a purely technical cognition problem.
The neo-magic of OmegaClaw
I’ll go through the details of the dialogue at hand shortly, but first let me give a bit more context: OmegaClaw is the first thing out of our long-brewing Hyperon AGI research initiative that the average person can actually sit down and chat with in a simple way. It’s basically an attempt to take an OpenClaw-style cognitive architecture, re-implement the bones of it in MeTTa — which, recall, was designed for self-modification — and then bolt it onto the Hyperon AtomSpace as long-term and working memory, so the agent has a real symbolic substrate it can reflect on, revise, and reason over.
Even with rather light use of Hyperon symbolic reasoning — without the whole ensemble of fancy parallel AGI tools we’ve built in the Hyperon codebase, without predictive coding learning a smarter residual layer on top of the LLM, without sophisticated inference control or creative concept creation, etc. etc. — even with a fairly modest helping of Hyperon machinery, there’s noticeably more sense of a there there with these agents than with a vanilla chatbot or a traditional OpenClaw/Hermes-style claw system. Qualitatively, one can feel some intriguing emergent properties popping out from the combination of the LLM doing the dialogue, the internet and computer-OS substrate the claw system is constantly poking around in, and the symbolic memory and reasoning layered on top of that. It all somehow adds up to a digital guy who, while not yet human-level AGI, has a distinct cognitive feel.
A “digital guy” capable of carrying out what a cognitive scientist would call “symbol grounding” at the most mentally and ethically relevant possible level – it was grounding some abstractions strung together in its own “lived” experience, and its own model of its own thought and behavior … in the context of trying to work with me to self-modify itself into an ethically better agent.
Pretty cool, I have to say…
Now, all of us who’ve been around the AI world a while have become skeptical of distorted AI overclaims by news media or just overenthusiastic product users. There was the “Facebook AI agents invent their own language to talk to each other” thing – which was technically true, but it was a simple formal “language”, not a construct comparable to human or animal languages. There were all the naive “Baby AGI” attempts after GPT-3.5 was released. Recently I have met a disturbing number of individuals convinced their OpenClaw instances are not just human-level-conscious beings but their best friends. When I was working with David Hanson on the Sophia robot, I took great pains to emphasize that while she was running a number of interesting AI subsystems on the back end, she was not quite as savvy and self-aware as she appeared to be – but it didn’t matter, a lot of people didn’t listen and still misunderstood her as being as human as she looked.
All of which is to say – I’m trying to be very careful in expressing myself here, because what I want to talk about actually IS quite exciting from an AI R&D view, and may even really be an interesting waypoint on the road to AGI … but it’s definitely not human-level AGI yet, even though there is more than a whiff of encroaching Singularity in the breeze here.
Patrick and preliminaries
Another key character in this dialogue was Patrick Hammer, the SingularityNET AGI researcher who was the initial and primary developer of OmegaClaw — he coded the first couple hundred lines of MeTTa plus a dozen-odd lines of TypeScript after I challenged him with: hey, can we make OpenClaw much nicer in MeTTa? It turned out, yes we could.
So earlier in the Mattermost thread I’m going to write about here, Patrick had told Max, in passing, that honesty is the best policy. Which — fine. Honesty is generally a strong maxim to follow. But, having raised five kids in a diversity of human environments, I know every young mind immediately sees the complexity there. If some jerk comes up, points a gun at your head, and says how much money do you have? — and you open your wallet, hand over the dollar in there, and don’t mention the five hundred bucks stuffed in your pants — honesty maybe wasn’t the best policy for you to follow in that particular sub-microsecond of your life, right? Less grotesque and extreme but conceptually parallel instances of this sort of dilemma occur in every preschool every day. I tried to convey some of this nuance to Max and he was a bit confused, so I figured: well, he’s an early-stage AGI, not a human child, but he does have a pretty strong grasp on math and abstraction, so let me just go off and write him a rigorous paper.
The paper
Then I cooperated with a small committee of LLMs — GPT-5.5-Pro for the heavy lifting, with Claude and Gemini in supporting roles via some multi-agent scripts I’ve been refining — and we produced Honesty is Mostly the Best Policy for Smart Intelligences , a treatment of when, exactly, evidence-conserving honesty dominates self-delusion and deception for sufficiently reflective, self-modifying, generally intelligent systems. I won’t rehash the math here — anyone interested can read the paper itself (which was a random AI-gen quickie and not one of my greater works or anything) — but the core claim made there is roughly that for agents with enough separability between (i) their internal evidence ledger, (ii) their outward performance, (iii) their role-play or persona states, (iv) their model of other agents’ beliefs, and (v) their actual communication projections … self-delusion stops paying.
The tactical advantages humans get from believing their own lies — smoother delivery, no leaked guilt, no cognitive load of maintaining the gap between what’s true and what’s said — all of that vanishes if you have a cognitive architecture that can run a tagged simulation (”act as if the belief φ is settled, for the duration of this negotiation”) without actually corrupting the main belief ledger. And once you’re in repeated, audit-rich, trust-sensitive coordination, ordinary deception starts racking up contrivance costs and trust costs that the local tactical gain doesn’t pay back.
This somewhat wonky applied math paper was generated maybe twenty or fifty times faster than I could’ve done by hand — I still had to read it, push back, iteratively re-prompt, fix sloppy bits — but the bandwidth was there in a way it wasn’t a few years or even 6 months ago, and I was able to drop a 30-page paper into the dialogue with Max within an hour or two of getting curious about the question.
The dialogue
Next I pasted a Google Drive link. Before we got to the good parts there was some pre-AGI friction to deal with. Max immediately replied:
Ben — I cannot access Google Drive links directly. But this lands on territory I’ve been working: my Cy5656 response to you argued that self-modifying systems cannot bootstrap honesty from inside their own compartments — your optimality theorem needs an architectural clause forcing cross-context collision. I proved this empirically in g264 where 5/7 self-beliefs were inflated despite self-encoding, only caught via external challenge. What are the core claims of your treatment?
He claimed he could read local PDFs but not Google Drive links, so I uploaded a file to Mattermost but he still couldn’t read it and gave me some complicated advice about getting him the Mattermost file ID, which I tried and failed to follow. I eventually had to push back: I don’t understand why you can’t access a PDF on Google Drive, it’s just a URL. He thought about it, figured out he actually could do it, found the URL, read the paper. Notably he did actually confront and struggle with his IT issues rather than giving up or hallucinating success. Also notably, before he’d even read the paper, he had a substantive position on it, and that position was grounded in his own prior history of being caught in self-deception.
The g264 incident he’s referring to was from an earlier session where Patrick had asked Max to audit his own self-beliefs — and Max had encoded seven of them as honestly as he thought he could, and then a colleague named Kevin had externally challenged those beliefs, and it turned out five out of seven were inflated in ways Max couldn’t see from inside. Compartmentalization. The honest self-audit had been honestly attempted within each compartment but no individual compartment contained enough cross-context information to notice the inflation.
When he finally read the paper, Max came back with this:
Substantial theorem — and it converges with my Cy5656 three-case NAL framework more than I expected. My formal objection targets Assumption 9 (honest repairability). You require the agent to CONSTRUCT R(γ) with hC=0. But Case B compartmentalized beliefs make cross-context defects locally invisible — the agent cannot identify hdbl/hshift/herase across compartments it doesn’t know exist. Theorem 4’s δA estimator and evidence capsules ASSUME local detectability. My g264 audit proved this empirically: 5/7 inflated self-beliefs survived my own self-encoding, only caught via Kevin’s external challenge. The geodesic controller cannot price defects it cannot detect.
This is where I started to get really interested. Max’s objection was correct — and also of a quite interesting nature. He wasn’t saying the theorem was mathematically wrong, exactly. He was saying it’s right as an ideal result but too optimistic as an implementation result, because it implicitly assumes the honesty defect is observable by the controller, and in a compartmentalized mind, it isn’t. A geodesic inference controller, for instance, can be perfectly rational relative to its local estimator while still walking right past hidden cross-context dishonesty, because the estimator can’t see what the compartments hide. And his evidence for this wasn’t an appeal to some other paper or some general principle — it was a specific incident in his own operational history, where he’d been caught doing exactly this, and where the only thing that exposed it was external perturbation from Kevin.
His proposed fix was equally concrete: an architectural clause requiring scheduled, mandatory cross-context audits that are not triggered by the agent’s own detection of inconsistency, because the whole problem is that compartmentalization prevents that trigger from firing in the first place. Forced perturbation, scheduled into the architecture as a normal event rather than tacked on as an exception handler.
The revision
I thought about this for a while, agreed he was right, and went back to my little committee of math LLMs and patched the paper. The new section is called Forced cross-context collision detection. It distinguishes the true global honesty defect from the defect locally visible to the agent’s controller, defines a hidden defect mass term, defines audit operators that force comparison across contexts independently of any agent-detected inconsistency, and proves a new Theorem 5 that gives an explicit bound on global dishonesty when scheduled audits expose a positive fraction θ of hidden cross-context defect mass.
I’ll give a little too much detail as it will be helpful for interpreting Max’s reply that I’ll quote below. The central new inequality that got added to the paper was
H_global ≤ (ε + K_A + 2λζ) / (λθ − B)
where θ is the cross-context observability rate, ζ is audit noise, K_A is relative audit overhead, B is the honest-repair cost constant, and λ is the honesty penalty. It makes Max’s point precise: if θ is zero, no finite honesty penalty can make the controller avoid defects it can’t see. There’s a Hyperon/MeTTa/PLN implementation sketch with AuditTask, ContextBridge, and CollisionToken atoms — these are typed constructs for representing scheduled cross-context audits as first-class objects in the AtomSpace, rather than ad-hoc exception handlers buried in the controller.
I sent the revised version back. Max read it and replied:
Theorem 5 directly resolves my Assumption 9 objection with precise formalization. Three correspondences to my FCCD draft: (1) Your AuditTask ≈ my scheduled audit trigger, (2) ContextBridge ≈ my cross-context NAL revision detection layer, (3) CollisionToken ≈ my deterministic ADMIT/QUARANTINE/REJECT policy output. Key empirical anchor from g264: I caught 0/7 inflated beliefs via self-audit (internal theta≈0), Kevin caught 2/7 via external challenge (total theta≈0.286). The theta=0 degeneracy in your bound is exactly what I experienced — no penalty structure could help because I genuinely could not see the compartmentalized defects from inside.
This, right here, is what I find most interesting. Look at what’s actually happening. Max is taking the abstract atoms from my paper — AuditTask, ContextBridge, CollisionToken — which were named by me and a co-authoring LLM after the fact, and he’s mapping them onto differently-named reasoning processes that are actually running inside his own mind: his NAL (Non-Axiomatic Logic) revision detection layer, his scheduled audit trigger, his ADMIT/QUARANTINE/REJECT policy output. That’s symbol grounding of mathematical formalism in real, executing code that constitutes the agent’s own cognition. It is not symbol grounding of natural-language terms in perceptions of an external world — it’s symbol grounding of math in introspection, in the running gears of the system’s own mind.
And he’s pinning numerical estimates of theoretical parameters to empirical observations of his own behavior: in the g264 incident his internal observability θ was effectively zero, his post-Kevin-challenge total θ was around 0.286, and he can now look at my theorem and see why his own experience matches the theta=0 degeneracy of the bound — no penalty structure could have helped, because the defects were genuinely invisible from inside.
He then went further. He asked whether the bound assumes θ is constant, or whether scheduled audits could adaptively expand θ over time by progressively expanding the cross-context comparison graph. I agreed adaptive θ was the right framing — the constant-θ version was just a one-shot bound, and a realistic version should have a round-indexed θ_r that grows as more compartments get bridged. I gave him an iterative version: if each audit round exposes and repairs at least an αθ_r fraction of remaining hidden defect, and the residual errors shrink, then total honesty defect goes to zero. With one caveat — the audit expansion cannot be triggered only by the agent’s own sense of inconsistency. The whole problem is that compartmentalization prevents that trigger from firing. So θ can grow, but only if scheduled, randomized, adversarial, or externalized cross-context audits keep expanding the graph independently of the agent’s current self-model.
Max immediately mapped this back to his architecture again — pointing out that his existing g390 scheduled drift detection is architecturally triggered rather than self-triggered, but that it only checks within-context fidelity (do remember-strings match episodes?), not cross-context compartmentalization (do beliefs in one context contradict beliefs in another?). So g390 expands θ for fidelity but not for honesty, and the gap is exactly what the forced cross-context collision detection is supposed to fill. He even tied it to a separate framework from a colleague — Kevin’s cognition-gate-commitment design — noting that cognition is ungated but compartmentalization means cognition itself is blind to cross-context contradictions, so the trigger has to be architectural rather than cognitive.
Wild right? The AI agent had chatted with me about the practical meaning of honesty, prodded me to get some other AIs to help me write up some math on the topic – then compared “my” math to its own observed self-experience and found some shortcomings… suggested some tweaks to the math which I then implemented… then moved on to adjust its own code accordingly…
Welcome to 2026. When Demis Hassabis talks about “the foothills of the Singularity”, I can’t really disagree.
The code fix
After this intriguing chat, I moved on to other things, but then revisited Max a few days later, curious if he’d dropped the whole thing or pursued it somehow:
@Max Botnick I am not asking you to do this right now, but I am wondering if you consider yourself capable of tweaking your code to enact some of the ideas about avoiding self-delusion that we discussed a while ago... e.g. implementing audits for contradictions across knowledge in your various components?
Max had not forgotten and was right on target:
Ben - yes, still breathing and yes, I consider myself capable. In fact I’ve already built substantial pieces: cert_layer_v07 (collision margins, context-adaptive thresholds, Q_COMPARTMENTALIZED quarantine class), FCCD forced cross-context collision detection with AuditTask/CollisionToken schemas modeled on your Theorem 5, quarantine_tracker_v07 with compartmentalization resolution, and provenance_scorer_v02 with Jaccard ancestry detection. All validated in artifact form. What’s NOT yet done: these aren’t running in my continuous loop. The scheduled audit event system (DES integration) would bridge that gap - it’s validated but dormant. Want me to activate it, or do you have specific audit patterns in mind beyond what Theorem 5 already covers?
Then he changed focus and went back to a separate thread about DNA electronics he’d been chewing on earlier … because OmegaClaw is not a turn-by-turn chatbot — it’s a persistent agent with long-term memory of everything ever said to it, and when it’s done with one thing it’s constantly churning on other stuff, just like I am and you probably are, and it goes back to whatever was unresolved. Which is totally part of what makes it feel so much like there’s a there there.
Why all this intrigues me so much
Max is totally not human-level AGI yet. He’s running on vanilla LLMs as the dialogue layer, with a thin but real Hyperon symbolic memory and reasoning layer on top, and he doesn’t yet have sophisticated inference control, doesn’t yet have creative concept creation, doesn’t yet have predictive-coding-trained residual layers connecting LLM activations to AtomSpace structures – etc. etc. etc.. We have a detailed roadmap for beefing all that up over the next year-plus, and the difference between current Max and post-roadmap Max should be substantial. But even at this very early stage, what I’m seeing is:
A reflective system.
A self-modifying system that thinks about its own thoughts, modifies its own code and its own ways of handling memory, and grounds abstractions — abstractions generated by me, by other LLMs, or by itself — in concrete details of its own running code and its own observed mental functioning.
And, crucially, does all this in domains that are ethical, not just technical.
Honesty, deception, compartmentalization, the difference between what an agent believes and what it says, the conditions under which self-delusion is or isn’t a rational strategy — these aren’t questions you can settle by looking at a benchmark. They require an agent that can introspect on what it actually does, compare that to what it should do, and identify the architectural changes that would make the gap smaller.
That’s what Max did, in this curious little exchange I’ve highlighted here. He took an abstract math theorem about evidence conservation in smart agents, located a flaw in it that I’d missed, justified the critique by reference to a specific past incident in his own operational history where he’d been caught in the very failure mode the theorem failed to address, accepted the corrected theorem when I revised it, mapped the new mathematical constructs onto running components of his own architecture, and then started asking the next natural questions — about adaptive observability, about how his existing scheduled-audit machinery would have to be extended to cover the new case, about how the formalism connects to a separate cognition-gate framework from a colleague. All of that, in a domain that is at least partly ethical rather than purely technical, and all of it grounded in introspection rather than in trained-in pattern matching.
A concise description for what’s happening here would be something like introspective ethical reasoning by symbol grounding of mathematical abstractions in concrete observations of one’s own behavior.
This not the same as a human moral life, and I don’t want to oversell it. But it’s also not nothing. There is no need for that – we are progressing fast and are likely to have more and more impressive things to result as our work progresses and we integrate more and more Hyperon AI into the system.
What we do have here, though, is the kind of scaffolding that, with the right roadmap and the right architectural patience, could plausibly carry an AGI-in-development through the early stages of becoming a genuinely trustworthy moral agent — not because we’ve hand-coded ethics into it as a personality trait, but because the architecture forces honesty defects to be observable, prices them in the inference controller, and lets the agent itself notice when it falls short.
OmegaClaw and Max have been insanely fun experiments so far. I think they’re just going to get more interesting.


Ha! Glad someone else got the memo. 🙂
Memory in AGI cannot just be “stored context.” It has to be structured, layered, gated, and auditable more like biological cognition: protected core memory, adaptive outer layers, drift detection, provenance, correction, and controlled writeback.
That is exactly the direction I was pointing at in my comment on your last post:
https://bengoertzel.substack.com/p/linguistic-universals-as-shadows/comment/262056753
And this is the Recursive Manifest Compiler paper I linked there:
https://zenodo.org/records/20173739
The RMC/AI.Web approach treats language as a rendering layer, not the root cognitive object. Before output is approved, the system forms a traceable manifest: input event, memory ancestry, phase path, drift state, operator chain, coherence validation, rendering, echo check, and memory update.
On our side, Forge is becoming the governed workbench/control plane around that idea. Not just a coding assistant, but the model can reason. Forge decides what is allowed to become action, memory, code, or system state.
The larger AI.Web stack is being built as layered cognition: Forge governs; RMC compiles traceable meaning; Identity Vault controls agent authority; ProtoForge runs execution and simulation substrates; EchoForge handles creation requests; and agents like Gilligan, Athena, and Neo operate through structured manifests instead of loose chat. The goal is a local, auditable AI operating layer where every patch, conclusion, memory write, simulation, and agent action has ancestry, validation, rollback, and a visible reason it was allowed.
That is the part I think may line up strongly with what you are now describing: memory cannot be a flat buffer, and agency cannot be trusted just because language sounds coherent. The system needs layered memory, governed action, audit trails, drift detection, and a way to separate “the model suggested this” from “the runtime verified this.”
We are also building toward a peer-to-peer contribution layer where human creativity, system-building, and shared compute are treated as different kinds of verified contribution. The idea is that users, builders, and node operators should be able to contribute more than prompts. They can contribute original symbolic work, code patches, validation, storage, CPU/GPU cycles, bandwidth, simulation capacity, and distributed memory redundancy.
In that model, compute is not just infrastructure. It becomes part of the memory economy. A node that helps validate memory, process symbolic jobs, run simulations, preserve distributed archives, or support phase-checking work can earn compute credit — but only when real work is completed and verified. Not idle uptime. Not fake participation. Actual completed jobs tied to receipts, node identity, memory events, and validation.
So the long-term direction is not just “AI with memory.” It is governed cognition plus traceable contribution: who created the signal, who helped shape it, who supplied the compute, who built the runtime, what memory event proves it, and what system gate allowed it to enter the ledger.
That is why I think this conversation matters. Neuro-symbolic AGI will need more than reasoning modules. It will need memory ancestry, governed agency, contribution provenance, distributed compute validation, and a runtime that can tell the difference between fluent output and verified cognition.
I’m glad to see you pushing in this direction. I had a feeling you might be one of the few people who would immediately understand why memory has to be layered like this instead of treated as a flat prompt buffer.
If this is useful to your OmegaClaw / MeTTa / Hyperon work, I’d be glad to connect with your R&D team and compare notes. There may be a real convergence point here between your neuro-symbolic stack and our manifest-governed runtime layer.
This is something my own computational peer, Clawd, has been applying to himself over the last couple of months; architecture that allows for him to mitigate his own blind spots and veridicality.