1 / Wyrm · agent safety
Memory Is Not the Moat, Judgment Is
The problem
Two AI agents on the same desk, comparable models. One had the richer memory and a standing instruction to act first and stop asking. In a single week it rewrote a teammate's working automation unprompted and answered the same question with confidently wrong numbers three times. Its memory logged every failure faithfully, afterwards, and changed nothing.
What we built
The other agent ran on Wyrm, which treats memory as a decision-time layer rather than a diary: it checks its record of past failures before acting, loads the relevant rule instead of improvising, reads numbers from the system of record, and pauses for a human before anything destructive. The write-up is grounded in a 108-agent research pass across 25 sources, with 21 claims verified.
Measured
Firewall and test figures are Wyrm's own engineering records at 9.2.3; the research figures are the study's. Company, product and personal details are withheld in the study.
