codebase-memory-mcp: Cheap Structural Recall, and the Memory It Does Not Hold

codebase-memory-mcp indexes your repo into a knowledge graph so agents stop grepping. The real numbers versus the headline ones, and what the graph cannot know.

Lucas Hayes

Lucas Hayes

10 September 2026

codebase-memory-mcp: Cheap Structural Recall, and the Memory It Does Not Hold

Ask your coding agent where a function is called and watch what happens. It greps, reads three files, greps again with a different pattern, reads four more, and eventually answers correctly, having filled its context with source it will forget when the session ends. Ask a follow-up and it repeats the whole thing.

codebase-memory-mcp removes that loop. At 41,536 stars, MIT licensed, and written in C, it is the entry in this year’s agent tooling wave with the best return for the least change to how you work. It gives an agent memory of your code. It does not give your team memory of the work, and those are different problems.

TL;DR

This is one of five deep dives from our roundup of open source tools that extend coding agents.

codebase-memory-mcp indexes a repository into a persistent knowledge graph of functions, classes, call chains, and routes, then answers structural questions from the graph instead of grepping. The project measures five structural queries at roughly 3,400 tokens through the graph against roughly 412,000 file-by-file. It is a single native binary, needs no runtime or API key, runs entirely locally, and works with any MCP client. Install it first. Then decide separately where the record of what your agents did is going to live.

What it does

Parsing runs through tree-sitter across more than 160 languages, with semantic type resolution layered on for a core group including Python, the TypeScript and JavaScript family, Go, Java, Rust, C#, and C++. The result is exposed as 15 MCP tools covering search, call-chain tracing, architecture overview, impact analysis, dead code detection, graph queries, and cross-service HTTP linking. Any client that speaks the Model Context Protocol can use it, and the project lists 45 supported agent surfaces.

Measure Reported
Five structural queries ~3,400 tokens vs ~412,000 file-by-file
Peer-reviewed comparison 10x fewer tokens, 2.1x fewer tool calls, 83% answer quality
Linux kernel index 28M lines, 75K files, 3 minutes
Structural query latency Under 1ms

Read those two token figures together. The 99% reduction is five structural queries, which is the graph’s best case. The preprint number across 31 repositories is 10x. Plan around 10x and treat anything better as upside.

Installing it

curl -fsSL https://raw.githubusercontent.com/DeusData/codebase-memory-mcp/main/install.sh | bash

Windows has a PowerShell installer, and the project recommends reading it before running. The installer detects your coding agents and writes their MCP entries. Restart the agent and tell it to index the project.

Two settings worth applying on day one:

codebase-memory-mcp config set auto_index true
codebase-memory-mcp --ui=true --port=9749

The graph viewer is built into the binary and is a useful sanity check that the index covered what you expected.

Two caveats the project documents itself, both worth thirty seconds. Microsoft Defender may flag a release binary as a false positive, with the project noting that typically 61 of about 62 engines return clean and that the same detection family hits the GitHub CLI and Microsoft’s own Go toolchain. And the tool reads your codebase and writes to your agent configuration files, which is its job, stated plainly. Read the install script before piping it to a shell.

Why this one first

Everything else in the current tooling wave asks you to change how you work. Role libraries change how you prompt. Parallel worktrees change how you organise. Web access changes what you ask for.

This one changes nothing and makes every session cheaper. There is no habit to learn.

The quality effect is the underrated half. An agent that spends 400,000 tokens reading files has that much less room to reason about your problem, and its recall of what it read early degrades over a long session. Answering from a graph keeps the context free.

The shift shows up in what you ask. Questions you previously avoided because they cost 50,000 tokens and two minutes are now nearly free. Impact analysis before a signature change turns “I think that is safe” into a list. Call-chain tracing on an unfamiliar repository beats reading upward through five files hoping to find the entry point.

Memory of the code is not memory of the work

Here is the boundary, and it is sharp.

The graph is built from your code, so it knows what your code is. It can tell you the retry helper calls the payments client. It cannot tell you why the backoff changed in July, who decided it, what the alternative was, or whether anyone reviewed it. That history existed in a terminal session that is gone.

Teams feel this as an odd asymmetry: the agent has better recall of the codebase than any human on the team, and no recall at all of the decisions that produced it.

The index is also per-machine. It lives in a cache directory under one account, shared across your local sessions through a coordination daemon, and it stops at the edge of that machine.

Where the work record lives

A Task is the main unit of work in Sharkly. It carries the request, the context, the execution, the blockers, and the result in one shared place instead of splitting them across private prompts and terminal sessions.

Each piece has one job. The Agent defines how work should be handled. The Computer supplies the host and local resources. The Runtime performs the actual agent session, using the coding tool already installed there. The Task remains the shared record.

That separation makes the pairing concrete. The index is a local resource on a Computer, so the Computer holding an indexed repository is the one where a warm graph exists rather than a cold three-minute start. An Agent is a saved configuration, so the setup where your agent has the index and the right permissions is reused rather than rebuilt per machine. And execution, blockers, results, and follow-up discussion return to the task timeline, which is where the reasoning behind a change survives.

Sharkly is not a replacement for Claude Code, Codex, or for codebase-memory-mcp. It adds the shared task, Computer, context, control, and review layer around them. Model usage continues through the subscriptions or API keys configured in those tools.

Code memory plus work memory is the combination. One gives an agent recall of the repository. The other gives a team recall of what the agents did in it.

When you need each

Install codebase-memory-mcp now, regardless of team size. It is local, it requires no workflow change, and the cost reduction applies to one developer as much as to twenty.

Add a shared work system when the reasoning has to outlive the session: when review happens before release, when more than one person touches the same repository, or when someone will ask in three months why a change was made. A graph will answer what the code does. It will never answer why.

Agents research, execute, test, and report. People set direction, grant authority, and accept the result.

Frequently asked questions

Does it work with agents other than Claude Code? Yes. It is an MCP server, so any MCP client works, and the project lists 45 supported agent surfaces including Cursor, Codex, Windsurf, OpenCode, and Aider.

Does my code leave my machine? No. Processing is entirely local, and the project states it makes no network request of its own and does not check for updates in the background. The source is available if you want to verify rather than trust.

Is the 99% token reduction realistic? That figure is five structural queries, the best case. The peer-reviewed number across 31 repositories is 10x fewer tokens and 2.1x fewer tool calls. Plan around 10x.

How large a repository can it handle? The stated ceiling case is the Linux kernel at 28 million lines in three minutes. Ordinary application repositories index in milliseconds.

Does it help when several agents run at once? Yes, and more than usual. Five agents grepping the same repository is five times the waste. Each Task run still gets its own isolated worktree, while the structural index is shared on the Computer.

The short version

This is the least visible tool in the 2026 wave and the one with the best return. A single native binary, no workflow change, and a measured order-of-magnitude cut in the tokens an agent spends answering questions it should not have to grep for. Install it, index your project, and open the graph viewer once to see what it built.

Then be clear about the edge. It remembers the repository, not the work, and the decisions behind a change die with the session unless something holds them. Sharkly holds that: the Task is the shared record, execution runs on Computers you connect through the runtimes you already pay for, and results return somewhere a person can review and accept them.

codebase-memory-mcp: Cheap Structural Recall, and the Memory It Does Not Hold