Claude Code Persistent Memory
Overcoming CLAUDE.md & Auto Memory Limits.
TL;DR
Give Claude Code persistent memory across every session. The Memwyre plugin auto-injects past project context when you start a session and auto-captures decisions when you exit — no manual CLAUDE.md management needed.
Answer: By default, Claude Code uses CLAUDE.md (static guidelines) and Auto Memory (local MEMORY.md). However, native Auto Memory is strictly capped at 200 lines / 25KB and silently truncates older notes when exceeded. The Memwyre Plugin solves this by using deterministic lifecycle hooks (SessionStart and Stop) to auto-inject dynamic, semantically ranked memories on launch and auto-capture architectural decisions on exit across Claude Code, Cursor AI, and VS Code via MCP.
The Problem: Claude Code Forgets Everything
Every time you close a Claude Code session, the context window resets. Your debugging breakthroughs, architecture decisions, database schema notes, and style conventions — all gone. You spend the first 5 minutes of every session re-explaining your project.
Anthropic provides CLAUDE.md and auto-memory as built-in solutions, but they require manual maintenance, are limited to flat Markdown files, and don't share context across tools. If you use Claude Code and Claude Desktop and Cursor, each tool maintains its own isolated silo.
The Complete Guide to CLAUDE.md: Hierarchy, Syntax & Best Practices
CLAUDE.md is the foundational instruction file for Anthropic's Claude Code CLI. Whenever you initiate a session, Claude reads this file and injects its contents into the root system prompt. Mastering CLAUDE.md is the single most effective way to eliminate repetitive prompting for static project rules.
Directory Hierarchy & Rule Precedence
Claude Code searches for instruction files across a deterministic directory hierarchy. Understanding this resolution order prevents conflicting instructions:
- Global Configuration (
~/.claude/CLAUDE.md): Applies across every repository and terminal session on your computer. Use this exclusively for developer-specific habits, such as preferred shell shortcuts, default terminal flags, or global Git commit author styles. - Project Root (
./CLAUDE.md): Placed at the root of your Git repository. This file is shared among your team and defines project-level architecture, build commands, testing frameworks, and linting rules. - Subdirectory Scopes (
./apps/web/CLAUDE.md): In monorepos, you can place localized files within package directories. When Claude runs commands inside that directory, it merges the package-level rules with the root rules.
Production CLAUDE.md Master Template
The most effective CLAUDE.md files are concise, imperative, and structured with clear Markdown headers. Avoid narrative paragraphs; use bullet points and exact shell syntax:
# Project: E-Commerce Microservices Platform
## Tech Stack & Package Manager
- Runtime: Node.js 20 LTS with pnpm (NEVER use npm or yarn)
- Framework: Next.js 15 (App Router only, Server Actions for mutations)
- Database: PostgreSQL 16 with Drizzle ORM
- Styling: Tailwind CSS v4 with Shadcn UI
## Core Build & Test Commands
- Dev Server: `pnpm run dev` (runs on http://localhost:3000)
- Typecheck: `pnpm run type-check` (TypeScript strict mode)
- Test Suite: `pnpm test` (Vitest unit tests)
- Integration Tests: `pnpm test:e2e` (Playwright)
- Database Migration: `pnpm drizzle-kit push`
## Architectural Invariants
- Never use standard API route handlers (`/api/...`) for form submissions; always use Server Actions with `zod` validation.
- All database queries must run through repository modules in `lib/db/repositories/`.
- Zero `any` policy in TypeScript. Use unknown + narrowing or generic constraints.
## Git & Workflow Rules
- Commits must adhere to Conventional Commits format: `feat(auth): add OAuth provider`.
- Never commit directly to `main`. Create feature branches (`feat/`, `fix/`).What NEVER to Put in CLAUDE.md (Context Inflation Anti-Pattern)
Because CLAUDE.md is injected verbatim into every single turn of your conversation, bloating it directly degrades performance:
- Do NOT include raw database schemas: Dumping 500 lines of SQL or Prisma models into
CLAUDE.mdconsumes 3,000+ tokens on every interaction. Claude can read schema files on demand via its file tools when needed. - Do NOT include ephemeral debugging fixes: Notes like "Fixed Redis port mismatch on line 88" become obsolete in days and clutter Claude's reasoning window.
- Do NOT paste conversation logs or changelogs: Historical changelogs belong in Git history, not in the active LLM context.
Anthropic Prompt Caching vs. Persistent Memory
Many developers confuse Anthropic's Prompt Caching with long-term memory. While both optimize context, they serve fundamentally different functions:
| Attribute | Anthropic Prompt Caching | Persistent Cross-Session Memory |
|---|---|---|
| Primary Purpose | Reduces latency & input token cost for repetitive prompt prefixes. | Preserves architectural decisions & solutions across terminal restarts. |
| Time-to-Live (TTL) | 5 Minutes (Ephemeral) | Permanent (Cross-Session) |
| Session Boundary | Dies when terminal closes or prompt pauses >5 min. | Persists across terminal sessions, reboots, and days. |
| Cross-Tool Sharing | ❌ Locked to active API request prefix. | ✅ Shared between Claude Code, Cursor, and VS Code. |
In short: Prompt caching saves tokens within a rapid 5-minute typing loop. But as soon as you step away for coffee, switch git branches, or close your terminal, the cache evaporates. Persistent memory stores lessons forever and injects them only when semantically relevant.
How Claude Code Memory Works: Deconstructing Native Limits
To understand why Claude Code developers frequently experience memory loss, we inspected Claude Code's memory engine (located in src/memdir/). Claude Code provides three native mechanisms for managing context:
1. CLAUDE.md (Static Rules)
CLAUDE.md is a Markdown file placed at your project root or in ~/.claude/CLAUDE.md. On startup, Claude loads this file verbatim into its context window. It is ideal for permanent guidelines (e.g., "Always use pnpm", "Follow TypeScript strict mode"), but it requires 100% manual updating and does not learn dynamically from your terminal sessions.
2. Auto Memory & MEMORY.md (Version 2.1.59+)
Anthropic introduced Auto Memory to let Claude write notes automatically to a local Markdown file at ~/.claude/projects/<project-name>/memory/MEMORY.md. While an improvement, it suffers from three critical architectural constraints:
- The 200-Line Silent Truncation Cap:
MEMORY.mdis hardcoded to truncate after 200 lines. When your project accumulates more than 200 lines of facts, older entries are silently dropped from Claude's prompt injection. Claude has no awareness that earlier debugging lessons have vanished. - The 25KB File Size Limit: Auto Memory enforces a 25KB binary ceiling. Large code snippets or database schemas quickly trigger this limit.
- Exact Keyword Search Only: Native Auto Memory relies primarily on exact text matching rather than two-stage cross-encoder semantic search. If your prompt phrases a problem differently than how Claude originally saved it, the memory is missed.
3. Auto Dream (Idle Consolidation)
Claude Code includes an experimental background process called Auto Dream that reorganizes memories while the agent is idle. However, it remains machine-locked: it cannot synchronize observations to your laptop, your teammates, or your other editors (Cursor, VS Code, or Claude Desktop).
Claude Code CLI Commands: Managing Working Context
During active development, Claude Code accumulates terminal output, file reads, and tool execution traces. Anthropic provides several built-in slash commands in the CLI to inspect and control working context:
| Command | Action | Context Impact |
|---|---|---|
| /memory | Opens the interactive memory viewer to review or edit stored notes. | Allows manual pruning of stale or inaccurate auto-memories. |
| /compact | Forces immediate context compression by summarizing conversational history. | Frees up 40–70% of context tokens. Warning: fine-grained error logs are lost. |
| /clear | Wipes all working session history and restarts the context window. | 100% reset. Only CLAUDE.md and auto-memory persist. |
| /context | Displays a visual bar chart of active tokens (system prompt, tools, chat). | Diagnostic tool to pinpoint files or tools causing prompt bloat. |
| /cost | Outputs dollar expenditure and cumulative token counts for current session. | Monitors prompt token burn rate in real time. |
The Team Silo Problem: Why Local Memory Fails Collaborative Engineering
By default, Claude Code writes notes to ~/.claude/projects/<project-name>/memory/MEMORY.md. Because this path is stored in your personal home directory on your local workstation, it is completely invisible to your teammates.
This creates severe organizational context fragmentation across development teams:
- Redundant Problem Solving: When Senior Engineer Sarah spends 45 minutes debugging an esoteric Docker networking issue on macOS and Claude notes the fix, Junior Engineer Alex encounters the exact same failure the next day and spends another 45 minutes rediscovering the solution from scratch.
- Git Merge Conflicts from Shared Files: Teams that try to solve this by committing dynamic notes to
CLAUDE.mdin Git quickly suffer from branch collision. Every developer's branch modifiesCLAUDE.md, generating messy merge conflicts on pull requests. - Cross-Editor Disconnect: Even on a single developer's laptop, if you alternate between Claude Code in your terminal, Cursor for quick edits, and Claude Desktop for strategic planning, none of them know what the other two discovered.
Four Approaches to Claude Code Memory
There is no single "right" approach — each method suits a different workflow. If you've researched this space, you've likely seen claude-mem (89K+ GitHub stars) alongside Anthropic's built-in options. Here is an honest breakdown of all four approaches.
| Feature | CLAUDE.md / Auto-Memory | claude-mem (OSS) | MCP Memory Server | Memwyre Plugin |
|---|---|---|---|---|
| Automation | Manual edits | Predictive (LLM decides when to save) | Predictive (LLM decides when to use tool) | Deterministic (Always runs on SessionStart/Stop) |
| Storage | Flat Markdown file | Local SQLite + ChromaDB | Varies (cloud or local) | Cloud vault + entity graph |
| Setup | Manual file creation | npx claude-mem install | JSON config + API key | claude plugin install |
| Cross-Session | ✅ Loads on startup | ✅ Local DB persistence | ✅ Via tool calls | ✅ Auto-injected on startup |
| Cross-Tool | ❌ Claude Code only | ⚠️ SSH sync (claude-mem-sync) | ✅ Any MCP client | ✅ Shared vault (Cursor, VS Code, Claude Desktop) |
| License | N/A (built-in) | AGPL-3.0 | Varies | Apache-2.0 |
| Best For | Static project rules | Local-first power users | Real-time tool access | Hands-free cross-tool memory |
These approaches are complementary, not mutually exclusive. Many production teams use CLAUDE.md for permanent guidelines and an external memory layer for dynamic session memory.
Approach 1: Pure Native (CLAUDE.md + Auto Memory)
Anthropic's built-in combination requires no external tools or API keys. You write static guidelines into CLAUDE.md, and Claude Code automatically records notes to MEMORY.md.
- Best for: Solo developers working on single repositories who don't mind manually curating Markdown files.
- Trade-off: Suffers from the 200-line silent truncation limit and zero cross-tool synchronization.
Approach 2: Local Open-Source (claude-mem by David Dias)
claude-mem is a popular open-source project (89K+ GitHub stars) that runs an autonomous memory server locally on your machine. It installs via npm and manages context using SQLite and ChromaDB:
# Install claude-mem CLI
npx claude-mem install Under the hood, claude-mem hooks into Claude's execution stream to summarize observations and store vector embeddings in ChromaDB.
- Strengths: Completely local-first, zero cloud reliance, open-source codebase.
- Limitations: Single-machine lock-in. To share context across multiple computers or team members, developers must configure
claude-mem-syncover SSH tunnels. Additionally, its AGPL-3.0 copyleft license creates compliance concerns for proprietary enterprise codebases.
Approach 3: Generic MCP Memory Servers
The Model Context Protocol (MCP) allows Claude Code to connect to external servers exposing memory tools (e.g., save_memory and search_memory):
# Add a custom MCP server to Claude Code
claude mcp add memory-server -- node /path/to/server.js- Strengths: Standardized protocol supported across Claude Code, Cursor, and Claude Desktop.
- Limitations: Predictive tool invocation. The LLM must actively decide when to query or update memory during conversation turns. This consumes reasoning tokens on every turn and frequently fails when the model doesn't realize relevant memory exists.
Approach 4: Dedicated Unified Context Engine (Memwyre)
Memwyre approaches memory deterministically rather than predictively. Instead of forcing the LLM to remember to make tool calls, Memwyre intercepts Claude Code's native lifecycle hooks (SessionStart and Stop). Context is injected before prompt typing begins and extracted when the terminal closes, synchronized across Claude Code, Cursor AI, and VS Code (MCP).
How the Memwyre Claude Code Plugin Works
The plugin hooks into Claude Code's native lifecycle events — two hooks, zero configuration after install. It reads your project directory name and handles context retrieval and capture automatically.
Developer runs claude in terminal. Plugin intercepts initialization before prompt input, identifies git workspace, and queries Memwyre API.
Top-k memories scored by cross-encoder relevance and recency are formatted into a clean XML block and injected silently into Claude's prompt.
On terminal exit (or /exit), full JSONL transcript sends to Memwyre worker. Model strips ephemeral shell noise and indexes structural decisions.
① SessionStart — Context Injection
When you open a Claude Code session, the plugin fires before the first prompt. It reads your working directory, queries the Memwyre retrieval engine for past memories matching that project, and injects them directly into Claude's system prompt.
<memwyre-context>
## Past Memories for my-project
- Database uses PostgreSQL 15 with pgvector
- Auth flow: JWT + refresh tokens in httpOnly cookies
- Fixed: race condition in worker queue (use Redis lock)
</memwyre-context>② Stop — Session Capture
When you exit Claude Code (or the session ends), the plugin reads the full JSONL session transcript, sends it to Memwyre's background worker, and extracts structured memories — architecture decisions, debugging solutions, code patterns, and configuration choices.
Troubleshooting & Edge Cases
No extraction model is perfect. Here's how to handle the edge cases:
- Misclassified memory: If the extraction model captures something irrelevant or incorrect, you can view, edit, or delete any individual memory from the Memwyre dashboard or via the API (
DELETE /api/v1/memories/:id). Every memory is individually addressable. - Stale facts: Decided to switch from PostgreSQL to CockroachDB? The Ebbinghaus logarithmic recency decay model automatically deprioritizes older, superseded facts. The most recent observation wins in retrieval ranking — you don't need to manually clean up outdated context.
- Deduplication: If consecutive sessions produce near-identical observations (e.g., "project uses Tailwind" captured in sessions #4, #5, and #6), the extraction model deduplicates them during ingestion. Your vault stays lean.
- Project exclusion: Don't want to capture sessions for a specific repo? Unset the
MEMWYRE_API_KEYenvironment variable for that terminal session, or configure project-level exclusions in your Memwyre dashboard settings.
Install in 60 Seconds
Choose between standard CLI plugin installation or manual hooks configuration:
Option A: Direct Plugin Install (Recommended)
- 1. Install the plugin package:
claude plugin install @memwyre/claude-code-plugin - 2. Export your API key:
export MEMWYRE_API_KEY="bv_sk_your_api_key_here"Add to your
~/.zshrc,~/.bashrc, or system environment variables.
Option B: Manual Hooks Configuration (Custom Setup)
If configuring hooks manually in ~/.claude/hooks.json:
{
"description": "Memwyre: Persistent autonomous memory",
"hooks": {
"SessionStart": [{ "hooks": [{ "type": "command", "command": "node \"/path/to/node_modules/@memwyre/claude-memwyre/dist/inject-memory.cjs\"", "timeout": 30 }] }],
"Stop": [{ "hooks": [{ "type": "command", "command": "node \"/path/to/node_modules/@memwyre/claude-memwyre/dist/capture-session.cjs\"", "timeout": 30 }] }]
}
}What Claude Remembers With Memwyre
- 🧠 Architecture Decisions: Database choices, API patterns, deployment configs, and framework decisions from past sessions.
- 🐛 Debugging Solutions: Race conditions fixed, environment variable gotchas, and edge cases you already solved once.
- 🔗 Entity Relationships: Connections between database tables, files, services, and APIs — enabling multi-hop reasoning.
- ✂️ Dynamic Pruning: Filters out duplicate CLI logs, compiler errors, and noise — keeping memory lean and token costs low.
Monorepo & Package Boundary Scoping
In multi-package monorepos (Turborepo, Nx, pnpm workspaces), native CLAUDE.md files often create context collision: Claude Code running in a backend service directory pulls in frontend styling rules, polluting the prompt and wasting context. Memwyre applies path-aware AST bounding to ensure strict context isolation.
Claude Code receives only relevant Next.js and shared UI tokens. No backend noise.
Claude Code receives pure backend security & database rules without CSS clutter.
Benchmark: Why Retrieval Quality Matters
Generic memory plugins dump raw vectors into Claude's context window. That approach fails on the queries that actually matter in a codebase — temporal reasoning ("when did we switch from REST to GraphQL?"), multi-hop connections ("which services depend on the auth token format we changed last week?"), and adversarial edge cases ("we never discussed Redis" → the system should abstain, not hallucinate).
We evaluated Memwyre's retrieval engine against a flat vector RAG baseline on the LoCoMo-10 benchmark (Snap Research, ACL 2024) — 1,540 questions across 10 long conversations spanning 200,000+ tokens:
| Category | Flat Vector RAG | Memwyre Engine | Improvement |
|---|---|---|---|
| Single-Hop Recall | 53.0% | 80.0% | +51% |
| Multi-Hop Reasoning | 24.0% | 45.0% | +87.5% |
| Temporal Alignment | 48.0% | 74.0% | +54% |
| Open-Domain Reasoning | 50.0% | 76.0% | +52% |
| Overall Accuracy | 43.7% | 70.5% | +61% |
| Context Tokens Sent | ~26,000 | ~3,000 | −81% |
The improvement comes from three architectural choices: dynamic context pruning during ingestion (strips conversational filler), two-stage cross-encoder re-ranking (high-recall vector fetch → precision cross-encoder scoring), and Ebbinghaus logarithmic recency decay (automatically deprioritizes stale observations).
View the full LoCoMo-10 benchmark results →Read our Vector DB vs. Agent Memory deep dive →
Real-World Workflow: Large Codebase Refactoring
To understand the depth of this integration, consider a common scenario: migrating a large React SPA to Next.js App Router.
In a standard Claude Code setup, you might tackle routing on Monday, API endpoints on Tuesday, and state management on Wednesday. By Wednesday, Claude has forgotten that you decided to use Server Actions for mutations instead of traditional API routes. It will start generating standard REST calls, requiring you to manually correct it and burn through tokens.
With the Memwyre Plugin, Monday's architectural decision ("we are exclusively using Server Actions for Next.js mutations") is extracted during the session stop hook. When you open Claude on Wednesday, that context is automatically injected. Claude inherently knows the project's boundaries, saving you countless prompt-correction cycles and significantly reducing your token usage by preventing hallucinated code paths.
Token Costs, Noise, & Security
Sending large conversational contexts directly impacts both billing and privacy. The plugin employs several strategies to mitigate this:
Algorithmic Noise Filtration
Not every CLI error needs to be remembered. Memwyre's backend uses a specialized extraction model that differentiates between ephemeral noise (e.g., a typo in a git commit command) and structural knowledge (e.g., adding a new enum to a Prisma schema). Only the structural knowledge is saved to your vector vault, keeping your persistent memory highly relevant and dense.
Predictable Token Usage
By summarizing and deduplicating past sessions, the plugin injects a concise <memwyre-context> block that rarely exceeds 1,500 tokens. Compared to manually pasting in megabytes of old transcript logs, this targeted injection saves Anthropic API costs while providing superior context.
Pasting old CLAUDE.md files, chat logs, and manual specs.
Dynamic cross-encoder top-k injection (~1,500 tokens max).
Calculated across 22 work days at 5 Claude Code sessions/day.
Zero-Retention & Self-Hosting
Your code is yours. The Memwyre extraction engine uses zero-retention policies—meaning your transcripts are processed in memory and immediately discarded. For enterprise environments with strict compliance requirements, the entire Memwyre backend can be self-hosted behind your firewall, ensuring your proprietary source code never leaves your VPC.
License & Requirements
- License: Apache-2.0 — no copyleft obligations. You can use, modify, and deploy Memwyre in proprietary environments without open-sourcing your changes. This matters for teams: claude-mem's AGPL-3.0 license requires that any modifications served over a network be made open-source, which creates real compliance friction in enterprise environments.
- System requirements: Node.js 18+ (for the plugin runtime). Works on macOS, Linux, and Windows. No local database required — unlike claude-mem (SQLite + ChromaDB dependency), Memwyre stores data in a managed cloud vault.
- Self-hosting: For teams that need full data sovereignty, the entire Memwyre backend is available for self-hosted Docker deployment. See the GitHub repo for instructions.
Why Not Just Use CLAUDE.md?
CLAUDE.md is genuinely useful — and you should keep using it. It's the right place for static project rules like "use TypeScript strict mode" or "prefer Tailwind over inline styles."
But it has real limitations:
- Manual maintenance — you have to remember to update it.
- Flat storage — everything goes into one Markdown file. No semantic search or entity relationships.
- Tool-locked — CLAUDE.md only works in Claude Code. Your Cursor and Claude Desktop sessions can't access it.
Explore More Agent Memory Integrations
FAQ
Does Claude Code have built-in memory?
~/.claude/projects/.../memory/. Both are useful for static rules but limited for dynamic, searchable, cross-tool memory.How do I give Claude Code persistent memory across sessions?
claude plugin install @memwyre/claude-code-plugin. Set your MEMWYRE_API_KEY environment variable, and every session will automatically load relevant past context on start and save new insights on exit.What is the best Claude Code memory plugin?
How is Memwyre different from claude-mem?
Does the plugin work with Claude Desktop too?
What license is Memwyre released under?
What happens if Memwyre captures something wrong?
DELETE /api/v1/memories/:id). The Ebbinghaus decay model also automatically deprioritizes stale facts over time, so outdated context naturally fades from retrieval results.Why does Claude Code forget instructions despite having Auto Memory and CLAUDE.md?
MEMORY.md file that is strictly hardcapped at 200 lines / 25KB. When your notes exceed 200 lines, older memories are silently truncated without notifying Claude. Additionally, native memory uses exact keyword matching rather than semantic vector search, so Claude fails to retrieve notes unless your query uses identical wording. Memwyre solves this with automated vector and entity graph search that scales to millions of tokens.What is the 200-line / 25KB limit in Claude Code native memory, and how does Memwyre bypass it?
src/memdir/), MEMORY.md is limited to 200 lines and 25KB to prevent system prompt bloat. When exceeded, facts at the top of the file fall off silently. Memwyre bypasses this by moving storage out of the local Markdown file and into an external, cloud-hosted semantic memory vault. Instead of dumping raw files, Memwyre retrieves only the top-k relevant memories (~1,500 tokens) on session start using two-stage cross-encoder re-ranking.Is my memory data private?
Give Claude Code Persistent Memory
One plugin install. Automatic context injection. Automatic session capture. Memory that works across Claude Code, Claude Desktop, Cursor, and VS Code.
Start Free