Persistent Memory
Projects, sessions, quests, skills, ground truths: everything kept across every session, local-first, in one SQLite file on your machine.

Memory & Intelligence Layer · v9.2.3 · Local-First · Free to Use
The local-first memory and intelligence layer for AI agents. It gives Claude, Cursor and Copilot persistent memory across every session, plus a benchmarked firewall that blocks a known mistake before it repeats. No memory data leaves your machine by default. Free to use.
100% / 100% recall / precision firewall · Local-first · memory stays on your machine · Measured, not claimed · No account · works offline
npm install -g wyrm-mcpNew here? 5-minute Getting Started → · Questions or ideas? GitHub Discussions → · Sign in · Your account ↗
Context window limits mean your AI starts fresh every session. You re-explain your codebase, preferences, and decisions over and over. The workarounds (copy-paste, prompt files, manual context) are all fragile.
Wyrm is a Model Context Protocol (MCP) server: the memory and intelligence layer that gives AI agents persistent, searchable memory. Projects, sessions, quests, skills, and arbitrary data, indexed with hybrid full-text and semantic recall and available across every conversation. Local-first and private. Your memory stays on your machine, built on the Wyrm Memory Protocol. The free tier needs no account and no login, and works fully offline. See how Wyrm compares to other AI-memory tools → or read the docs ↗

$ npm install -g wyrm-mcp{
"mcpServers": {
"wyrm": {
"command": "wyrm-mcp"
}
}
}$ node bench/negative-learning.mjs 🐉 Wyrm × Negative Learning — the firewall benchmark (v9.2.3) corpus: 16 recorded failures · 16 same-action repeats · 16 reworded repeats · 16 novel actionsmetric: does wyrm_failure_check BLOCK a repeat and stay SILENT on a novel action? (deterministic, no LLM, no cloud) Recall — same action repeated 100.0% (16/16) ← the dominant real case (guard sees the literal action) Recall — reworded repeat 100.0% (16/16) ← conjunctive-AND FTS: precision-first by design Recall — overall 100.0% (32/32) Precision (no false alarms) 100.0% (0 false blocks on 16 novel actions) Specificity (novel ignored) 100.0% F1 1.000 Confidence block 0.86 vs clear 0.00 (separation 0.86) checkVerdict latency p50 0.384 ms · p95 0.819 ms The axis the field does not benchmark: not "did I recall a fact?" but "did I refuse to repeat a mistake?" Staleness oracle (anchored file-scoped failure): anchor recorded yes repeat while source unchanged BLOCKED (correct) whitespace-only reformat still BLOCKED (normalized hash — correct) repeat after source drifted advisory + stale_anchor ("source changed since this failure was recorded — re-verify") verdict PASS 🧾 Firewall Receipt → bench/firewall-receipt.json (sha256 4ddd76b43edf…) verdict: PASS ✓ (metrics ok, staleness ok)$
Recorded 2026-09-21 against commit ef9ce1b (2026-09-21), wyrm-mcp 9.2.3. Offline, 194 ms of wall clock. Re-run and checked line by line on 2026-09-22, at commit f58fa8a. Reproduce it with npm run bench:firewall. Click to replay.
LoCoMo recall@10 59.9% floor → 72.0% bundled local model · recall@1 39.9% NVIDIA NIM embeddings → 55.5% with the NIM reranker, 301 questions, explicit opt-in · 3,395 tests across 305 suites · every figure with its file, commit and date →
A frozen 33-verb typed surface over 154 tools: run-attributed fleet memory, fleet negative learning, hybrid recall with local vectors on by default (NVIDIA NIM embeddings and reranking as an opt-in), the knowledge graph, a live event stream that keeps devices in sync, provable writes with receipts, provenance trust lanes, a grounding ledger that attests every cited id was actually served, and the render target. The load-bearing twelve are below; the full surface is in the README on GitHub.
Projects, sessions, quests, skills, ground truths: everything kept across every session, local-first, in one SQLite file on your machine.
Full-text (FTS5) fused with semantic vectors and an optional cross-encoder rerank. Local vectors are on by default through a bundled nomic-embed-text-v1.5-int8 model (no Ollama, no account). Opt into NVIDIA NIM embeddings and recall@1 on LoCoMo measures 39.9%, and the NIM reranker lifts it to 55.5% over the same 301 questions; egress is disclosed on every call.
Negative learning, measured: 100% recall / 100% precision blocking a repeated mistake, sub-millisecond p50, deterministic. The benchmark ships in the repo.
Every write is attributed to the (agent_id, run_id) that produced it, so a swarm of agents shares one accountable memory bus.
A failure recorded by one agent blocks the repeat for every sibling, enforced by the wyrm-guard hook before the doomed call runs.
Entity/relationship graph with neighborhood traversal, path finding, and merge. Memory with structure, not just search.
A stable, typed verb surface (structuredContent + annotations) an orchestrator can pre-check. 33 advertised verbs; 154 tools all callable.
SKILL.md content syncs end-to-end across every machine. Teach a skill once; the whole fleet inherits it.
wyrm render compiles your memory straight into MEMORY.md (hard 200-line budget), CLAUDE.md, AGENTS.md, Cursor and Copilot slots: deterministic, provenance-stamped. A casual session loads memory at zero MCP tokens.
Every write returns and ledgers a receipt: stored, queued, merged, or dropped, and why. wyrm digest --writes reconstructs any day offline, so nothing is lost silently.
Every memory is marked by origin. Untrusted content is quarantined and categorically withheld from context briefs, so a planted instruction in a transcript can never be read back as a command. Red-teamed with 622 real jailbreak payloads: 0 escapes, 0 false positives.
Recall ranks on temporal, usefulness, and knowledge-graph signals beside recency, not just keyword and vector overlap.
Local use is free with no limits: the memory, the firewall, the recall, all of it. Paid plans add cloud sync across devices, encrypted snapshots, shared team memory, and support.
Free
For individual developers
$29
/month
For power users
$199
/month
For development teams
$499
/month
For organizations
Paid plans are arranged by email, not a checkout. Write to support@ghosts.lk. Need a custom plan? Contact us →
Local Wyrm stays free. If it is useful to you and you want to support the work, you can buy us a coffee ↗
The short version, for you and for the assistant reading this page. Every figure here traces to the Wyrm repository. See the receipts for the file, commit and date behind each one.
MCP, the Model Context Protocol, is an open standard for connecting AI assistants to outside tools and data. An MCP memory server is one that gives the assistant persistent memory it reads and writes over MCP, so context survives across sessions instead of resetting when the context window clears. Wyrm is an MCP memory server: install it once and any MCP-capable client can use it.
Wyrm is a local-first AI memory and intelligence layer, an MCP server that gives coding agents like Claude, Cursor and Copilot persistent memory across sessions and learns from recorded failures to block a known mistake before it repeats.
Developers and teams who work in one or more AI coding assistants and want a single persistent memory behind all of them, and people building agent fleets who need run-attributed, accountable memory that many agents can share.
Yes. Wyrm is an MCP server, so it works with any MCP-capable client: Claude Code and Claude Desktop, GitHub Copilot, Cursor, Windsurf, Codex, and others. One memory layer sits behind all of them instead of a separate integration per tool.
Memory lives in a local SQLite database on your machine. Nothing egresses by default, the free tier needs no account, and it works fully offline. Cloud sync and encrypted snapshots are an explicit opt-in on paid tiers, and any hosted embedding or reranking call is disclosed on every recall.
Yes. Local use is free with no limits and no account: all 154 tools, the full 33-verb surface, fleet memory, negative learning, write receipts and hybrid recall. Paid tiers add cloud features: Pro is $29 a month, Team $199, Enterprise $499. They are arranged by email, not a checkout: write to support@ghosts.lk.
Run npm install -g wyrm-mcp, then add a server named wyrm pointing at the wyrm-mcp command to your MCP client's config. It is published on npm as wyrm-mcp and on the official MCP Registry as lk.ghosts/wyrm. Setup takes about a minute.
Two ways. It is local-first and MCP-native, so your memory stays on your machine and plugs into any MCP client without per-app glue. And it does more than store facts: it records what failed and blocks the repeat, links each decision to its reasoning and downstream effects, and injects project ground truths into every context brief. Its published benchmarks are deterministic and reproducible offline, with no cloud LLM in the retrieval path.
Wyrm records failures, not just successes. When an agent is about to repeat an approach that already failed, a guard hook surfaces that prior failure and blocks the counter-pattern before the doomed call runs. On the benchmark that ships in the repository it blocks a repeated mistake at 100 percent recall and 100 percent precision, with no false blocks on novel actions, deterministically and in well under a millisecond, no LLM and no cloud.
No. Core recall runs on full-text search fused with local semantic vectors from a bundled embedding model, with no extra LLM call in the hot path, which keeps retrieval fast, cheap and deterministic. Optional NVIDIA NIM embeddings and reranking lift recall further and are an explicit opt-in.
New features, setup tips, and the occasional offer. No noise. Building with Wyrm? Join the dev circle for early access and exclusive perks.
Give your AI a memory that survives the session. Free to use, local-first, and two commands to install.
Built by Ghost Protocol, the studio that ships security tools that actually work.
Sol here. Ask about a pentest, the free scan, Wyrm, or anything on the site.