NVIDIA's Chief of Staff Recipe: Readable Agent Memory
NVIDIA's NemoClaw recipe makes agent memory readable, correctable and auditable. What the design gets right, what its own benchmark really shows, and what remains unsolved.

Sur cette page
- What NVIDIA actually shipped
- Three stores, three kinds of truth
- Correction is a data path, not a vibe
- Urgency is a claim, not a fact
- Memory cannot authorise anything
- The benchmark, read properly
- The limitations the launch post skips
- What this does and does not answer
- What to do now
- Frequently asked questions
- Is the NemoClaw Chief of Staff a product I can deploy?
- Does the agent act on my messages?
- How is this different from Tencent's Team Memory or Mem0?
- How good is the benchmark evidence?
- The bottom line
- Sources
The strangest thing about NVIDIA's latest agent post is what the memory is made of. Not a vector store, not a fine-tuned model state, not a proprietary context format. Markdown pages you can open in a text editor, and a SQLite database you can query from the command line. The company selling the picks and shovels of the AI build-out has published a memory architecture whose load-bearing property is that a human can read it.
On 4 September, NVIDIA's developer blog published Building a Memory-Driven Agent with NVIDIA NemoClaw, describing a Chief of Staff agent built on NemoClaw, its open-source reference stack for running agents inside sandboxed runtimes. The companion recipe is public in the nemoclaw-community repo, along with the synthetic benchmark behind the post's numbers. I have read the post, the recipe and the benchmark harness rather than deployed the thing, so treat this as a desk read. Even at that distance, this is the most carefully argued piece of agent-memory engineering a major vendor has put in public this year, and its argument is worth taking seriously: persistent memory only stays trustworthy if it stays inspectable and correctable.
Key Takeaways - The recipe splits agent memory into three stores with three governance regimes: source evidence it never modifies, a human-readable Markdown "self model" of derived knowledge, and a SQLite obligation ledger holding judgements, corrections and an append-only audit trail. - Correction is a designed data path, not a re-prompt. User corrections persist across runs, land in the audit trail, and repeated patterns compile into a small preference policy the user can inspect, edit or delete. - The model never writes to the memory directly. It returns a versioned JSON envelope that deterministic code validates and applies transactionally, and an "intent gate" stops inbound urgency from outranking work the user actually chose. - NVIDIA's own benchmark shows a real but modest overall gain (+8.1 points against agentic RAG on 186 synthetic questions), with the headline temporal gains resting on five- and six-question slices, and two categories regressing. The vendor published the regressions. - This is a reference recipe for one person's work stream, not a workplace deployment. It sends no messages, its fixtures are synthetic, and its own limitations list includes a credential-provenance gap the blog post does not mention.
What NVIDIA actually shipped
The blog post is the wrapper. The substance is the Memory-Driven Chief of Staff recipe, proposed in nemoclaw-community issue #122 on 18 August and now merged, which packages the agent as a profile for Hermes, one of the agent runtimes NemoClaw supports. The agent, called NemoHermes in the recipe, watches a user's Slack and Outlook, maintains a structured memory of people, projects, priorities and working patterns, and produces ranked recommendations about where the user's attention is needed.
Context matters for how to read this. NemoClaw is NVIDIA's alpha-stage, Apache-2.0 reference stack for running OpenClaw, Hermes or LangChain Deep Agents inside OpenShell, its sandboxed agent runtime, with managed inference and network policy. The main repo sat at 22,378 GitHub stars when I checked today. NVIDIA has been building toward this since the NemoClaw launch at GTC Taipei in June, but the earlier material was about sandboxing and deployment. This recipe is the first time NVIDIA has published a worked answer to a harder question: what should an agent's memory actually look like once it persists?
The post's five "design lessons" are easy to skim past as vendor homily. The recipe underneath them is not.
Three stores, three kinds of truth
The architectural move is a separation most agent memory systems blur. The recipe keeps three kinds of information in three different places, each with its own rules.
First, evidence. Slack messages and Outlook mail stay in their source systems, and the collectors are read-only by construction: the recipe never posts, replies, edits, deletes or marks anything as read. When the agent later claims something, the claim can be traced back to a source that the agent itself cannot have touched.
Second, knowledge. The "self model" is a set of structured Markdown pages about people, projects, goals, patterns, concepts and current attention, with a schema that defines indexing, cross-references, provenance and growth limits. Provenance here is concrete: non-obvious claims carry their source, evidence date and a trust level. NVIDIA is explicit that the self model stores a derived interpretation rather than replacing source evidence, precisely so that when an answer is wrong you can work out which stage failed: the evidence, the memory maintenance, the retrieval, or the model's final decision.
Third, judgement. Whether an item needs attention, where it ranks, whether the user ignored it: all of this lives in a SQLite ledger at a known path, alongside corrections and an append-only audit trail. The recipe is careful never to write judgements back into the source systems as read flags, labels or folders. Your inbox does not become the agent's scratchpad.
The detail I like most is who holds the pen. The model does not write SQL. It returns a versioned JSON envelope of proposed decisions, and a deterministic script, apply_decisions.py, validates that envelope and applies it transactionally. The model proposes; code disposes. If the envelope is malformed, nothing happens, and the audit trail says why. Anyone who has watched an agent fumble a stateful write will recognise the value of that boundary immediately.
Maintenance runs on a schedule rather than a vibe: intake every half hour, obligation review every six hours, then a nightly chain of memory writing, repair, consolidation and preference update between 01:00 and 04:30. A "wake gate" means the selectors check for work first and the model is only called when there is some. Memory upkeep that costs nothing on a quiet night is a small, boring, excellent decision.
Correction is a data path, not a vibe
Here is the piece that connects this recipe to a question this publication has been circling for weeks. When I compared shared agent memory systems in August, the unanswered question was correction: what happens when a memory is wrong, and who gets to fix it. Neither Tencent's Team Memory nor Asana's Agentic Work Management documents a real process. The NemoClaw recipe answers the question at single-agent scope, and the answer is structural rather than procedural.
A user correction is an explicit command, correct.py priority <item> low or correct.py ignore <item>, not a hopeful note in the next prompt. Corrections survive later runs, so the agent does not resurface something you already demoted. Repeated corrections are no-ops, and corrections against completed or incompatible rows fail with a specific exit code and an explanation of the required state transition. Every change lands once in an append-only audit trail.
Then the interesting part. When corrections against the agent's judgement repeat past a fixed threshold, they compile into a small, readable preference policy. One bad Monday does not rewrite the rules; a pattern does. And the policy is a file you can inspect, edit or delete, not a nudge buried in model state you will never see. The loop is visible end to end: agent judgement, user correction, audit event, preference update.
There is honest machinery elsewhere in the lifecycle too. Message bodies are cleared after 30 days by default while metadata, obligation state and audit history remain. Recipient lists are reduced to "direct", "mentioned" or "broadcast" and never stored. Export gives you the complete store, the memory and the learned policy in Markdown and JSON. Deletions in Outlook are reconciled through Graph and tombstoned; Slack deletions age out through retention because a bounded history read cannot distinguish "deleted" from "older than the window", and the recipe says so. These are the decisions of a team that has thought about what happens to this data on a bad day.
Urgency is a claim, not a fact
The intent gate deserves its own moment, because it is a small piece of defensive design wearing a productivity costume. The highest attention tier is reserved for obligations with proof in memory that the user chose the work: the item connects to a stated priority. An urgent expense-policy attestation stays visible but ranks below a quieter request tied to something the user actually committed to. The walkthrough fixtures test exactly this: an urgent deadline unrelated to the user's chosen work stays in medium.
Now read that as a security property. Every inbound message is untrusted input that would love to seize the top of the queue, and "urgent" is the cheapest word in business. The intent gate means a message cannot promote itself by tone; it can only be promoted by a relationship to memory the user owns. And the ranking itself is not the model's call. The model may interpret the relationship, but deterministic code enforces tier size, overflow behaviour and ranking order. An agent whose attention can be hijacked by the loudest sender is a liability; this one is structurally unimpressed.
Memory cannot authorise anything
The post's best sentence is the one every agent builder should steal: "Context can inform an action, but it cannot authorize one."
The mechanism behind it is the NemoClaw and OpenShell boundary. Memory and retrieved content are inputs to the model, not trusted security policy. If the agent misinterprets its memory, or follows malicious instructions embedded in something it ingested, it still operates inside operator-defined runtime boundaries: OpenShell sandboxes the filesystem, process and network access, and credentials for managed inference and MCP connections stay outside the sandbox entirely. This is the same direction of travel as the turn-scoped permissions work in OpenAI's Codex runtime: authority is something the runtime grants per decision, not something the agent accumulates by remembering.
Worth pausing on why this matters for memory specifically. A memory page noting that a colleague prefers Slack can inform a recommendation. What it must never do is lower the bar for sending the message. Keeping the knowledge layer readable and the action boundary enforced elsewhere is one design, and it is the right one.
The benchmark, read properly
NVIDIA built a synthetic benchmark, mnemo, to back the claim that the self model improves task performance. Before the numbers, the harness deserves credit. All 186 questions are graded deterministically, with no model anywhere in the scoring path. The two corpora are fully synthetic: 425 documents from a platform-engineering team, and 173 from a construction programme manager generated with a different model, ingested in ordered halves so later information supersedes earlier claims. Abstention questions, where the corpus simply does not say, are scored wrong if the system answers confidently. Corpus, questions, gold answers and scorer are fingerprinted so runs only compare when nothing moved. The harness even refuses to print a dollar cost it cannot measure, on the stated grounds that a cost nobody measured must not read as zero. I wish that sentence were unnecessary. It is not.
The published comparison runs an agentic RAG baseline against the self-model configuration, both on NVIDIA Nemotron 3 Ultra, so the model is held constant and memory architecture is the variable. The overall result is a genuine but modest gain. The category rows need more care than they will get on social media, so here they are with the question counts the post includes but the headlines will drop:
|
Category |
Questions |
Agentic RAG |
Self model |
Difference |
In answers |
|---|---|---|---|---|---|
|
Overall accuracy |
186 |
82.8% |
90.9% |
+8.1 pp |
+15 |
|
Hard questions |
31 |
67.7% |
87.1% |
+19.4 pp |
+6 |
|
Facts that changed over time |
5 |
60.0% |
100.0% |
+40.0 pp |
+2 |
|
Point-in-time reasoning |
6 |
33.3% |
66.7% |
+33.3 pp |
+2 |
|
Entity disambiguation |
15 |
66.7% |
86.7% |
+20.0 pp |
+3 |
|
Multisource synthesis |
73 |
87.7% |
94.5% |
+6.8 pp |
+5 |
|
Faithfulness to corpus |
13 |
100.0% |
92.3% |
-7.7 pp |
-1 |
|
Single-hop lookup |
30 |
86.7% |
83.3% |
-3.3 pp |
-1 |
|
Citation coverage |
186 |
92.5% |
97.8% |
+5.4 pp |
+10 |
Three honest observations. First, the two most dramatic rows, the temporal ones, rest on five and six questions; the +40-point "changed over time" result is two additional correct answers. Directionally encouraging, statistically thin. Second, the regressions are real and the vendor printed them: the self model got one faithful-abstention answer wrong that the baseline got right, and one single-hop answer worse. The abstention slip is the one I would watch. A memory that knows things is slightly more willing to answer when it should say the corpus does not say, and that is exactly the failure mode you do not want compounding. Third, the repo's own results README states limits the blog post omits: the published runs cover corpus A only, on one base model, the answers predate a rename at publication, and the accounting cannot support a cost comparison.
And then the frame around all of it. This is a synthetic benchmark against a synthetic workplace. The recipe's offline walkthrough substitutes recorded model envelopes at the two points where inference would otherwise run. The recipe does not send messages or modify source systems at all. Nothing here demonstrates performance against a live inbox with a live adversary on the other end. Vendor evidence, usefully self-consistent, still self-asserted. No independent replication exists yet; the post is two days old.
The limitations the launch post skips
The blog post is candid by vendor standards, but the recipe's own limitations section goes further, and one entry deserves daylight. The Graph collector's credential provenance is not enforced at runtime. The setup script refuses a provider that does not match the recipe's read-only profile, but ingest_graph.py trusts whichever attached provider supplies the access token; a different provider exposing the same key can bypass the setup-time refusal. NVIDIA's own mitigation advice is to inspect the attached providers and not run Graph intake unless the effective provider is the enforced read-only profile. That is a real gap in the exact boundary the post celebrates, and to the recipe's credit it is written down where an operator will find it.
The rest of the list sets scope honestly: scheduled jobs need the Linux sandbox; live connectors cover only Slack and one Outlook mailbox; the scheduled writer currently maintains only people and attention pages, with the other page types having schemas but no automatic writer; memory compaction is model-guided with a checker that detects invariant violations but does not invent a safe consolidation; and the whole thing is described as "a reference recipe for one user's work stream on a machine they control, not a hosted service or a production-readiness claim."
What this does and does not answer
Set against the memory stories this site has already run, the contribution comes into focus. The Tencent Team Memory piece was about team-scale governance: owners, visibility tiers, review before share. The shared-memory comparison found correction and conflict resolution undocumented by everyone. NemoClaw's recipe is deliberately not playing that game. It is one principal's memory on one machine. Nothing in it addresses propagation of a wrong fact to other people's agents, or whose version wins when two agents disagree, because nothing in it shares. Those problems remain exactly as open as they were in August.
What it does answer is the prior question, the one a team memory system would need answered first: how does a memory stay correct for the person it serves? The answer offered here is that correctness is maintained by a correction path, and you can only correct what you can read. Judgement becomes data with an audit trail, knowledge becomes Markdown with provenance, and evidence stays where the agent cannot touch it. If a team-scale system ever grows out of this lineage, those are the right primitives to inherit.
There is also a quiet connection to the agent mind-virus work covered here last week. That paper showed ideas propagating between agents through their own memory files, and its practical lesson was that the write path to persistent memory is where the security model lives. NemoClaw's design reads almost like a response: the write path is narrow, transactional, audited, and the thing being written is legible to the human it serves. Readable memory is not a complete defence against a poisoned entry, but an entry you can read is an entry you can argue with.
What to do now
If persistent agent memory is on your roadmap, three steps, in order.
Today: run the offline walkthrough. It needs only Python, makes no network requests, and exercises the ranking, the intent gate, correction durability and the memory checker against synthetic fixtures, including a deliberate defect the checker is shown catching. It is the rare vendor demo you can run without handing over a credential.
This week: open the state. After the walkthrough, read the ledger with sqlite3 and the memory pages in your editor, then correct something and confirm the correction survives a later run. The entire value proposition of this design is that you can do exactly that, so do it.
Before any real deployment: resolve the Graph provenance caveat explicitly, decide your retention numbers rather than inheriting the 30-day default, and get your own abstention measurement on your own corpus. The synthetic benchmark tells you the architecture can beat agentic RAG on a made-up workplace. Your inbox is not made up.
Frequently asked questions
Is the NemoClaw Chief of Staff a product I can deploy?
Not really, and it does not claim to be. It is an Apache-2.0 reference recipe in the nemoclaw-community repo, described by its own documentation as one user's work stream on a machine they control, running inside a NemoClaw-managed OpenShell sandbox. Live connectors currently cover Slack and a single Outlook mailbox. Treat it as an architecture you can run and adapt, not a service you can buy.
Does the agent act on my messages?
No. The current recipe reads Slack and Outlook but sends nothing and modifies nothing in the source systems. It maintains memory, ranks obligations and recommends where your attention is needed. NVIDIA frames that scope as deliberate: live action connectors would require separate handling for credentials, privacy, retention and deletion.
How is this different from Tencent's Team Memory or Mem0?
Scope. Tencent's Team Memory and Asana's Agentic Work Management govern memory shared across a team, with owners and access tiers. Mem0, Zep's Graphiti and Letta centre on one agent remembering its own past. NemoClaw's recipe is also single-principal, but its distinguishing move is internal: separating evidence, knowledge and judgement into three stores, and making user correction a durable, audited data path rather than a prompt-level plea.
How good is the benchmark evidence?
Methodologically careful, statistically thin, and synthetic. The deterministic grading, abstention scoring and comparability fingerprinting are better practice than most vendor benchmarks. But the headline temporal improvements rest on five- and six-question slices, two categories regressed, the published runs cover one corpus and one model, and no independent replication exists as of 6 September 2026.
The bottom line
The field has spent two years treating agent memory as a storage problem: bigger windows, better retrieval, more vectors. NVIDIA's recipe is interesting because it treats memory as a governance problem at the smallest possible scale, one person's working life, and answers with legibility, a judgement ledger, and a correction loop that survives contact with the next run. The benchmark is honest enough to publish its own regressions, and the limitations list is honest enough to include a hole in the boundary the marketing celebrates. Both kinds of honesty are rarer than they should be.
What it has not done is solve shared memory, and it has not proven any of this against a live workplace. The open question I am left with is whether the preference-policy threshold, the point where repeated corrections rewrite the rules, is calibrated for how people actually correct agents: rarely, reluctantly, and usually once too late. If a vendor ever gets that dial right at team scale, memory stops being the feature everyone demos and becomes the infrastructure everyone trusts.
If you are working out where persistent memory fits in your own agent stack, that is a conversation I have with clients regularly. Get in touch.
Sources
NVIDIA Developer Blog, "Building a Memory-Driven Agent with NVIDIA NemoClaw": https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/ (published 2026-09-04, retrieved 2026-09-06)
NVIDIA, nemoclaw-community repository, Memory-Driven Chief of Staff recipe (README, docs/data-lifecycle, limitations): https://github.com/NVIDIA/nemoclaw-community/tree/main/examples/recipes/nvidia/memory-driven-chief-of-staff (retrieved 2026-09-06)
NVIDIA, nemoclaw-community issue #122, recipe design proposal: https://github.com/NVIDIA/nemoclaw-community/issues/122 (opened 2026-08-18, retrieved 2026-09-06)
NVIDIA, Agent Memory Benchmark (mnemo), including corpus READMEs and results caveats: https://github.com/NVIDIA/nemoclaw-community/tree/main/examples/tools/agent-memory-benchmark (retrieved 2026-09-06)
NVIDIA, NemoClaw repository: https://github.com/NVIDIA/NemoClaw (retrieved 2026-09-06; star count checked 2026-09-06)
NVIDIA, OpenShell repository: https://github.com/NVIDIA/OpenShell (retrieved 2026-09-06)
NVIDIA NemoClaw documentation: https://docs.nvidia.com/nemoclaw/latest/ (retrieved 2026-09-06)
Continuer la lecture
Agent Field Notes
Recevez le prochain numéro.
Harnesses d’agents, environnements d’exécution, sécurité et gouvernance, expliqués pour celles et ceux qui doivent exploiter ces systèmes.
Vous faites face à une décision de ce type ?
Nous réalisons des revues d'architecture, des évaluations de gouvernance et des comparaisons de frameworks à versions figées pour les équipes confrontées à des décisions déterminantes sur les systèmes d'agents.