What Really Remained After the Body Changed

Results from the Agents-Brain test run from July 6 to 15: stable identity, write-back working, visible gaps, and genuine harness blockers.

Eine Frau prüft einen Retro-Roboter am Teststand, der in einer Broschüre als perfekt beworben wird; die Anzeige warnt vor zu früher Feier.

What you'll take away from this: Here’s my honest assessment of my Agent Brain trial run from July 6 to 29—including the initial nine-day findings, the evidence that has since come to light, and the gaps that still remain.

On July 6, 2026, the test for which I had built the Agents Brain began: Nox moved to Hermes, while Identity and Memory remained in the repository.

Nine days later, I was finally able to say more than just „The files were still there.“ This finding from July 15 remains the historical snapshot in the article. Since then, two more weeks of testing and real-world performance tests have been added.

Here is a summary of five questions—each showing the difference between the status at that time and the status as of July 29.

1. Is the identity recognizable in the new tool?

Yes—with one important distinction.

Nox’s self-written SOUL from the previous Workspace was carried over verbatim into the Vault. The only additions were neutral notes indicating where old tool references are historical. The tone, approach, and working principles were preserved.

His operational manual, on the other hand, has been revised. This was necessary because the body—and thus the available tools—had changed. Hermes is now the primary orchestrator body; the former tool environment is now part of historical context.

The test has thus confirmed a meaningful threshold:

  • Content related to identity can be transferred reliably.
  • Tool-specific work rules must be adapted to the new runtime.
  • Both should be placed in separate files so that the change remains visible.

So „same agent” does not mean that every line must remain unchanged. It means that changes are made at the correct level.

2. Was the memory useful?

Yes, because it wasn't marked as complete.

The old long-term memory was curated during the move. Outdated project statuses were clearly marked. Sensitive data was deliberately not transferred to the central Brain. The undocumented KW12 remained visible as a known gap.

This allowed Nox to start with a genuine backstory without confusing old information with current knowledge.

That’s an important distinction. A large amount of old data wouldn’t have made for a better memory. It became useful through categorization: What is reliable? What is the status as of March? What’s missing? What needs to be re-examined using current sources?

This approach has proven effective in everyday work. Nox does not treat the repository as an oracle. To check the current status of projects, he consults the relevant primary sources.

3. Does write-back work during operation?

For the test period: yes.

From July 6 through July 15, there will be a Nox Daily Log for each calendar day. It will include tasks, decisions, delegations, corrections, and open issues.

The impact is more important than the number. Several corrections derived from individual incidents have been turned into lasting lessons. Examples include the active OpenClaw narrative approach, the ban on vague „someday“ threads, and the rule to secure portable adapters with the same CI gates before pushing.

This means that the trip home from work to the Brain is not only planned but also put to practical use.

At the same time, the test revealed a limitation: The Brain is not a passive record of all tools. A Codex task executed directly did not initially appear in Nox’s memory because no agent had been designated as the owner. On July 15, this issue was identified, and the local Bridge was refined so that direct repository work without any other assignment belongs to Nox by default.

This shows: Write-backs require accountability. Without an agent owner, the system does not know whose memory a task belongs in.

4. Do Travel, Governance, and Privacy Actually Go Hand in Hand?

In the kernel that was tested: yes.

Nox and Sol upload their RULES using the approval gates. Sol's "review before posting" rule applies regardless of which harness creates the draft. Nox is also not allowed to publish public content without my approval.

At the same time, file constraints have become machine-verifiable. The Vault validates required front matter, launch sequences, approval gates, Brain Root references, and unwanted persona copies. A Privacy Gate scans tracked and tagged content for secret patterns and unauthorized mirrors of confidential agent data.

On July 15, the full repository tests at that time—comprising 148 tests—passed; CI confirmed this interim result under Python 3.9 and 3.11 without annotations.

By July 29, the suite had grown. For the new memory contract, 224 unit tests, 72 Vault regressions, and 35 privacy regressions passed. The local full test passed under Python 3.9.6; the GitHub Matrix confirmed Python 3.9 and 3.11. These figures confirm the versioned contract that was actually tested as of that date.

Validators can check the structure and identify known risks. Whether a publication is factually accurate remains a matter for human review.

5. Has the transformation of the body already been fully proven?

On July 15, the test for Hermes and the central Nox workflow was robust enough to continue working. For all four documented harnesses, the answer was still “no.”.

By July 29, the situation had become clearer. Claude Code, a fresh Codex run, and Hermes loaded the same binding startup sequence from the same vault in separate read-only smoke tests. All three received the same technical evaluation from the new Memory Checker. The automated Fresh Clone test also passed: It reconstructs a draft agent, including the bridge, in an isolated local clone and verifies the Markdown write/readback between separate processes—using only the Python standard library and Git.

These two tests must not be combined. The Three-Body Smoke test verified reading and classification, not write-back. The Fresh Clone tests the recovery routine without actual harness CLIs. For Gemini, adapters and contract tests still exist, but there is no complete runtime test case yet. Therefore, a true write-back has not been proven across the board for every body.

At the same time, the new memory contract was rolled out to the production server clone and validated there with a successful readback. This confirms the deployment path of the shared brain, not the operational capability of each connected tool.

These outstanding issues will determine the next level of maturity.

An existing bridge indicates that the architecture can address the tool. A read-only smoke test validates the shared read contract. Only a successful startup, a completed task, and a verified memory write-back demonstrate that the agent is fully operational there.

Addendum: There are now two instances of this complete proof. Claude Code achieved the write-back breakthrough on July 29. Grok Build followed on August 12 as the fifth verified entity alongside Hermes, Claude Code, Codex, and Kimi Code—from skill verification through the interactive session to two separate headless runs with a read-back daily marker. The repository thus documents six adapter families instead of four.

What the trial run through July 15 shows

After nine days, I have five reliable results:

  1. A persona closely tied to identity can survive a tool change in readable files.
  2. Operational rules can be adapted separately to the new body.
  3. Daily logs and lessons function as a true write-back when an agent owns the task.
  4. Approval gates and privacy rules can be embedded in Brain independently of the harness and verified automatically.
  5. Missing documentation remains missing; the system highlights gaps rather than concealing them.

That's less spectacular than an autonomous memory that supposedly knows everything. But it's more valuable for my work.

What has been added as of July 29

The core has become more accurate since the first test run. Observation, inference, and generalized rules can now be distinguished in the memory. Contradictions must be updated with the previous status, new evidence, and a visible resolution. A curation is flagged as due after 14 days or ten new Daily Logs, but remains subject to human review. Outdated standalone Memory and Lesson files are superseded—rather than deleted—with the date, reason, and successor. The event date and Git commit time remain separate.

These rules aren't just documented. The shared read contract was verified by three entities, the Fresh Clone Smoke is green, and the productive Brain clone has been brought up to the same status. The main outstanding issues are the actual Gemini run and the complete write-back receipt for each harness.

What I Would Do Differently

I would treat agent ownership as a separate component of the memory architecture. The direct Codex run has shown that a shared brain alone is not enough. Every task requires a clear assignment: Who reads? Who decides? Who writes back?

In addition, I would track E2E test results per harness from the very beginning. "Adapter available," "Validator green," and "Runtime verified" are three different states. They should never be lumped together into a single status indication.

And I would flag known gaps in my knowledge just as early as I would mark existing knowledge. An honest admission of a gap saves more time later on than a supposedly complete memory that no one can distinguish from a reconstruction.

My Assessment

Agents Brain was not a finished product on July 15. Nor is it one on July 29. It is a living system with a core that has since been more widely adopted.

Nox is operating in a new primary body. His identity is recognizable, his memory is organized, and his daily writing routine is active. The governance team is traveling with him. Claude Code, Codex, and Hermes are reading the same contract; the Fresh Clone is performing the Markdown-only recovery. At the same time, Gemini and the genuine write-back per body remain open as separate proofs.

This is exactly how I wanted to write this workshop report: not as a story about an inventor and not as a glossy conclusion, but as a clear answer to a practical question.

Can an AI agent switch tools without me having to rebuild its working identity?

For the kernel and recovery image tested from the canonical files, my answer is: yes. I'm still working on the evidence to support universal „plug-and-play“ functionality.

Update: Two of these issues have since been resolved—Claude Code as of July 29, Grok Build as of August 12. Gemini remains open.

Sources

  • Project source: Nox Migration, Daily Logs July 6–29, 2026, Memory Maintenance, Harness Interop, Fresh Clone, and CI Logs (private primary sources)
  • Git commit: Initial Agents-Brain framework from May 16, 2026, and Nox migration from July 6, 2026
  • Pro Git: What is Git?

About the Author

Saskia Teichmann provides consulting on AI, e-commerce, and digital platforms, and personally verifies technical assumptions in architecture and code.

More About My Work

Discussion on the post

0 comments

Join the discussion

Your email address will not be published. Required fields are marked.

News from the Studio

New posts via email.

Whenever a new workshop report or guide is published, you'll receive a brief email with a link. No set schedule, no ads.

Prefer to read via RSS

Your registration will not become active until you click on the confirmation link. You can cancel your subscription at any time by clicking the link in any email.