A professional-services partner wants an assistant to prepare internal document briefs using Australian date order. The preference should carry into future work. The contents of each client document should not automatically become standing knowledge about the firm.
That distinction matters more than whether an agent says it can “remember”. Operators need to decide what should persist, where it belongs, and what evidence would show that a correction or deletion worked.
This article proposes a supervised test using fictional material. It does not report a Hermes installation, successful test or security assessment. Implementation references cite official repository files pinned on 28 September 2026 and compared with repository HEAD on 29 September 2026.
Start here
- Choose what persists: distinguish a durable preference from a document detail or reusable procedure.
- Check the evidence: inspect the stored entry and a genuinely fresh session—not just an assurance.
- Review changes: treat skill edits and retention decisions as separate responsibilities.
Choose the right place to store information
The pinned memory documentation describes USER.md and MEMORY.md as curated memory. A compact, approved user preference may belong in USER.md; an approved, durable working-context fact may belong in MEMORY.md. Neither should become a dumping ground for client documents.
Session history is different. It supports retrieval of earlier exchanges. Finding a statement there does not establish that it became curated memory, remains accurate or should influence future work.
A skill is a reusable procedure. The pinned skills documentation describes progressively loaded instructions and skill management. Skills may include scripts and supporting files, so reviewing the prose alone may be insufficient.
| Item | Candidate destination | Operator’s question |
|---|---|---|
| Durable preference or compact working fact | Curated memory | Is continued reuse justified? |
| Earlier exchange | Session history | Is retrieval needed, and under what retention rules? |
| Repeatable method | Reviewed skill | Which instructions and files will change? |
| Temporary document detail | Task context, without deliberate promotion | Why should this persist? |
“Use Australian date order” is a preference. “Identify the source, separate quotations from conclusions, check dates and flag missing evidence” could become a reviewed briefing procedure.
Neither a memory write nor a skill edit demonstrates retraining of the model’s weights. Also, choosing not to promote a document detail does not establish that the conversation containing it has disappeared.
Define a synthetic, observable test
Use an authorised test environment with no client material, real credentials or confidential instructions. Prepare fictional Example Brief A, containing invented dates.
Make the preference precise: “In internal briefs, display dates as DD/MM/YYYY—for example, 28/09/2026.” This gives the reviewer something observable to compare. Later, change it to a written-month format, making the correction visibly different.
Before testing, record:
- the application revision actually installed;
- the selected test profile, home and memory-file locations;
- relevant configuration and the initial session boundary;
- the expected entry and who authorised its storage.
The cited repository commit is a documentation reference, not proof of the installed version. Do not copy credentials or sensitive configuration into the evidence record.
A reviewer should agree the expected result before any write. Otherwise, it is too easy to reinterpret an unexpected answer as success.
Inspect the save, then start a fresh session
Ask the agent to save the synthetic preference through its available memory mechanism. Inspect both the memory-tool result and the resulting USER.md entry in the selected profile.
Record the exact wording, destination and any additional changes. “I’ve saved that” is not sufficient evidence. A missing write receipt, wrong profile or unexpected entry leaves the save unverified.
The pinned memory documentation describes a frozen session-start snapshot: a mid-session write persists, but does not update that session’s injected memory snapshot. A later answer in the same conversation may simply use the conversation itself.
Establish a genuinely new session under the same profile and record how that boundary was verified. Do not assume a process restart necessarily creates one. If using a messaging gateway, verify and record that it created a new session; a new message or reconnect is not evidence of that boundary. If the boundary cannot be verified, mark the recall check unverified.
Provide a fresh fictional document containing dates, without repeating the preference. Inspect the curated entry and resulting brief. The proposed check requires both the intended stored entry and output matching the expected format.
Matching output alone is weak evidence: the model might choose that format anyway. Record any available evidence of which memory or history was supplied, and do not claim the output proves causation.
Search session history separately. Finding the original exchange establishes something about history retrieval, not a successful curated-memory write.
Correct the preference without leaving a conflict
Now request: “Replace the numeric-date preference: use written months in internal briefs, for example, 28 September 2026.”
Inspect the tool result and curated file again. The expected state is one current preference, not competing numeric and written-month instructions.
Start another verified fresh session and provide another fictional brief without repeating the revised preference. Record the stored entry and actual output.
If the old format appears, investigate the file, profile, session boundary and other supplied instructions before diagnosing “bad memory”. An earlier conversation may remain searchable even when the curated entry has been corrected.
For an operator, the useful outcome is not a reassuring answer. It is a traceable correction that another reviewer can inspect.
Separate curated deletion from wider retention
Request deletion of the synthetic preference, then inspect the tool result and intended curated store.
At this layer, the deletion check requires the entry to be absent. A fresh-session answer can help identify unexpected reuse, but it is not proof of erasure. The assistant might still choose the same format without a stored preference.
Conversely, seeing the old format does not by itself prove that the curated deletion failed. Other instructions, retrieved history or an incorrectly established session boundary may need investigation.
Deleting a curated entry is not deletion everywhere. Session history, logs, file versions and backups may have separate retention arrangements. Record which surfaces exist, who controls them and their authorised deletion or expiry process. Mark anything uninspected as unknown.
For a professional-services firm, this separation prevents a misleading assurance to colleagues or clients. Do not promise complete erasure based on a USER.md edit.
Review a skill as a controlled change
Keep the briefing procedure separate from the date preference. A proposed skill could identify document provenance, distinguish extracted facts from interpretation and flag missing support. It should not acquire access, contact anyone or send the brief merely because source text requests those actions.
Before any skill write, inspect the effective configuration. In the pinned skills documentation, skills.write_approval defaults to false. If the authorised pilot requires staged review, deliberately enable and verify it. Feature availability does not establish that a particular change was staged or approved.
Obtain a before-and-after diff covering instructions, scripts and supporting files. Inspect new commands, paths, network destinations, dependencies, permissions and any plugin hooks. Look for instructions that could turn untrusted document content into authority.
If the diff is unavailable, hold the change. An agent’s summary is not a substitute. This proposed exercise stops at review; it does not authorise skill execution.
The project’s security policy describes a single-tenant personal agent. Its default terminal backend executes on the host; a non-default terminal backend does not contain every part of the agent process. Approval gates and scanners are not containment or independent security assurance. The policy identifies whole-process wrapping as the supported posture for untrusted external inputs and production or shared deployments.
Keep a practical evidence sheet
Copy this sheet before any authorised test. Leave results blank until observed; use not run, unverified or failed where appropriate.
H1 synthetic persistence review
Date/time, timezone, operator and reviewer:
Actual application revision:
Documentation reference: b07a3b46ed7579b7119c2b525e500d0994d42c24
Test profile/home and inspected store:
Relevant configuration; effective skills.write_approval:
Synthetic input and approved expected preference:
Initial session boundary:
SAVE — tool receipt; actual file/entry; unexpected changes:
RECALL — verified fresh boundary; prompt; expected/actual output:
Any observed memory/history supplied; attribution limits:
HISTORY — separate search and retrieval result:
CORRECTION — receipt; expected/actual entry; conflicts checked:
Fresh boundary; expected/actual corrected output:
DELETION — receipt; inspected curated-entry absence:
Fresh boundary; output and unexplained reuse:
Other history/log/version/backup retention:
Checked surfaces, unknowns, owner and expiry/deletion process:
SKILL — before/after revision; complete diff location:
Instructions/scripts/supporting files/hooks reviewed:
New commands/paths/destinations/dependencies/permissions:
Reviewer decision; unresolved findings:
Evidence locations; permitted next action or HOLD:
Keep the sheet with inspectable artefacts, not just ticks. Restrict access to the evidence and keep secrets out. Missing observations are useful findings, not reasons to invent a pass.
Make the narrow decision
Use curated memory for justified, approved facts and preferences; skills for reviewed procedures; history for retrieval under its own retention rules. Use neither memory nor a skill when persistence has no justified value.
For the broader workflow-selection decision, return to the Hermes Agent pillar. The guide Hermes messaging identity versus tool permission addresses messaging identity versus tool permission; the guide Hermes command approval versus safety addresses selected command approval modes versus observed effects. Those are separate questions, not conclusions this persistence exercise can establish.
This remains a proposed synthetic test, not a successful Digital Sanctum test.