A streamable-HTTP code intelligence MCP server for LLM clients.
The point: offer a network-based code insights for both humans & machines. Although probably more suitable for humans (talk to your code) rather than specific compiler-based code insights like those of Serena.
Able to handle very large codebases.
Re-uses ripgrep crates just like ripgrep itself.
Warning
Authentication & Authorization are outside the scope of this project.
Run this on a private LAN only.
Anyone who can reach the bind address can use the tools but --project scopes what they can read.
Run this as docker OR with chroot for extra jailing.
NOTE: https://modelcontextprotocol.io/docs/tutorials/security/authorization is what you should be using regardless.
Q: What is this - ever heard of Serena MCP?
A: Yes but this is very different. Serena uses the power of language servers & therefore compilers to answer questions. It is incredibly powerful but also really slow. Plus, Serena struggles with 100s of repos because a single LSP can only handle one project at a time. This goes through 100s of repos like knife through butter.
Q: How can you get any intelligence out of cat, find, and grep - you mad?
A: You'd be surprised! Modern SOTA models (Sonnet, GLM, Kimi, Codex, etc.) are really good at navigating large codebases, provided they have the tools to do so. This project gives them these tools.
Q: No auth??
A: No. Auth is hard and you shouldn't be relying on my auth anyway. MCP standard defines auth, use that.
Download the latest release from the Releases page.
The included Dockerfile uses a multi-stage build with cargo-chef for optimal layer caching. The final image is based on debian:bookworm-slim (~80 MB) and contains only the stripped binary and git (needed by the ignore crate for .gitignore traversal).
Multi-arch images (linux/amd64, linux/arm64) are published to GHCR automatically on every release. Pull the pre-built image — no local build needed:
docker pull ghcr.io/devfire/code-mcp:latest# Run with defaults (bind 0.0.0.0:8080, project /project)
docker run -p 8080:8080 -v /path/to/repo:/project:ro ghcr.io/devfire/code-mcp:latest# Override any flag — all CLI args are supported natively via ENTRYPOINT
docker run -p 9090:9090 \
-v /path/to/repo:/project:ro \
-v /path/to/memories:/memories:ro \
ghcr.io/devfire/code-mcp:latest \
--bind 0.0.0.0:9090 \
--project /project \
--memory-dir /memories \
--max-sessions 128 \
--initialize-rate-per-min 20 \
--session-idle-timeout-secs 3600# With debug logging
docker run -p 8080:8080 -e RUST_LOG=debug,rmcp=info \
-v /path/to/repo:/project:ro ghcr.io/devfire/code-mcp:latestThe ENTRYPOINT is the binary itself, so any arguments after the image name replace the default CMD and go directly to clap.
Pass --help to see all options:
docker run --rm ghcr.io/devfire/code-mcp:latest --helpBuild locally instead
If you want to build from source (e.g. for an unreleased commit or a fork):
docker build -t code-mcp .
docker run -p 8080:8080 -v /path/to/repo:/project:ro code-mcpFlags:
Details
--bind <addr:port>— default0.0.0.0:8080.--project <path>— required. Every path the tools touch is canonicalized and required to lie within this directory; anything outside is rejected withinvalid_params. Symlinks in input paths are resolved before the check, socat /proj/link-to-etc-passwdis rejected because its canonical form is/etc/passwd. The server refuses to start without it.--memory-dir <path>— optional. If set, enables thememoriestool and reads<path>/instructions.md(if present) into theInitializeResult.instructionspayload sent to the model on connect.--max-sessions <N>— default64. Hard cap on concurrent stateful sessions in the rmcpLocalSessionManager. New initialize POSTs are rejected with503 Service Unavailable+Retry-After: 5once the cap is met. Existing-session traffic (any POST carryingMcp-Session-Id) passes through untouched.--initialize-rate-per-min <R>— default12. Per-peer cap on new initialize requests, expressed as a per-minute token bucket (capacity =R, refilling continuously over 60 s). When exhausted, new initializes from that peer return429 Too Many Requests+Retry-After: <secs>. A misconfigured client that reconnects in a tight loop gets throttled here instead of pinning unbounded session state. Default12/min ≈ one fresh session every 5 s sustained — well above any healthy reconnect rate.--trust-forwarded-for— defaultfalse. When set, the gate uses the rightmost entry ofX-Forwarded-Foras the peer IP for rate-limiting. This assumes a single trusted proxy hop (e.g. AWS ALB) that appends the real client IP; entries to the left of the last hop are client-supplied and forgeable. Only enable when the server sits behind a reverse proxy you control.--session-idle-timeout-secs <N>— default1800(30 min). Idle timeout for stateful sessions. A background reaper task closes any session whose last observed request is older than this, so abandoned clients (process killed, network gone, no DELETE sent) don't pin slots against--max-sessionsindefinitely. The cap defends against bursts; the reaper handles long-lived zombies.--session-sweep-interval-secs <N>— default60. How often the reaper sweeps for idle sessions.
cargo install --git https://github.com/devfire/code-mcp.gitAll tools return structured ToolResponse objects with metadata (truncation status, error counts, match counts) rather than plain strings. This allows clients to programmatically detect truncation and other conditions.
Details
### `grep` Regex search across files using parallel directory traversal (`ignore` + `grep-searcher`).| arg | type | default | notes |
|---|---|---|---|
directory |
string |
— | required |
pattern |
string |
— | required; Rust regex flavor — no lookaround/backrefs |
output_mode |
string |
files_with_matches |
files_with_matches (list matching files; fast for broad scans), content (matching lines with context), or count (per-file match tally) |
before_context |
int |
0 |
lines of context before matches (ignored in files_with_matches and count modes) |
after_context |
int |
0 |
lines of context after matches (ignored in files_with_matches and count modes) |
max_results |
int |
100 |
exact cap (no over-shoot); for files_with_matches, caps the number of files; for content, caps the number of matching lines; for count, caps the number of files |
case_insensitive |
bool |
false |
equivalent to (?i) prefix in pattern |
include_hidden |
bool |
false |
|
follow_symlinks |
bool |
false |
|
respect_gitignore |
bool |
true |
|
file_extensions |
string[] |
[] |
e.g. ["rs", "toml"]; empty = all |
max_bytes |
int |
~5 MiB | hard cap on response size |
Output modes:
files_with_matches(default): Returns only file paths that contain matches. Each path appears once (on first match), then the file's search stops early — efficient for broad reconnaissance queries.max_resultscaps the number of files.content: Returns matching lines with optional context (before/after). The classic grep output mode, useful when line-level detail is needed.max_resultscaps the number of lines.count: Returns per-file match tallies aspath: Nlines, sorted by path. Useful for understanding distribution of matches across files.
Walker errors and search errors are tallied and returned in the response metadata rather than silently dropped.
Find files by regex.
| arg | type | default | notes |
|---|---|---|---|
directory |
string |
— | required |
pattern |
string |
— | required |
max_results |
int |
100 |
|
include_hidden |
bool |
false |
|
respect_gitignore |
bool |
true |
|
match_basename |
bool |
true |
when false, the regex matches the full path |
Load persisted context (conventions, project facts, prior feedback) for this server. Available only when --memory-dir was set at startup; otherwise returns an invalid_params error.
| arg | type | default | notes |
|---|---|---|---|
name |
string |
— | Optional filename within the memory dir (e.g. area_network.md). Plain basename only — .., /, \ are rejected. |
Without name: returns the contents of <memory-dir>/MEMORY.md if present, otherwise a listing of *.md files in the dir. Re-reads on every call, so edits made on disk are picked up live.
The model is told about this tool via InitializeResult.instructions whenever the server is launched with --memory-dir. The expected pattern is:
- On session start, the model calls
memorieswith no args → gets the index. - The index points to specific memory files via
cat-able paths or names. - The model loads what's relevant via
cat(or anothermemories(name=...)call).
This mirrors Claude Code's auto-memory pattern.
Read file contents with pagination.
| arg | type | default | notes |
|---|---|---|---|
file_path |
string |
— | required |
offset |
int |
0 |
0-based line number to start from |
max_lines |
int |
2000 |
maximum lines to return per call |
max_bytes |
int |
~5 MiB | hard cap on response size (UTF-8-safe cut at line boundary) |
Use offset to page through large files: if the response indicates truncation, call again with offset = previous_offset + max_lines. The response will include metadata indicating whether the result was truncated and the reason.
With --project ./my/repo set. What's ok and what isn't:
cat ./my/repo/src/main.rs— inside the root, allowedgrep ./my/repo --pattern foo— directory inside the root, allowedcat /etc/passwd— outside the root, rejected. Nice try lolcat ./my/repo/../../etc/passwd— canonicalizes to/etc/passwd, rejected. Samecat ./my/repo/link-to-secret(symlink to/etc/passwd) — symlink resolves outside root, rejected. Same same.
The --memory-dir is not required to be inside --project — it's server-side config, not user-driven file access.
<memory-dir>/
├── instructions.md # appended to InitializeResult.instructions on connect
├── MEMORY.md # index — returned by memories() with no args
├── area_network.md # returned by memories(name="area_network.md")
├── area_crypto.md
├── area_commands.md
└── ... # one area_*.md per functional area
The instructions.md file is read once at startup. The other files are read on demand by the memories and cat tools, so editing them does not require a restart.
code-mcp is a read-only, LLM-free tool server: it can serve a memory directory but it cannot create one. Generating memories needs read access to the code, write access to the memory dir, and a model — and all three only coexist on the box where the files live, not in a remote client connecting over the network.
So bootstrapping is a separate, co-located step: run a local coding agent against the repo to produce the MEMORY.md index + per-area files, then point --memory-dir at the output. The included script does this:
# Uses `claude -p` by default; the agent reads/writes the local filesystem directly.
scripts/generate-memories.sh ./my/repo ./memories
# Any stdin-driven agent with local fs access works:
AGENT="codex exec" scripts/generate-memories.sh /srv/monorepo /srv/memoriesThe script feeds the agent a prompt that builds a functional-area mental model — the major subsystems, their entry points, how they talk, and the non-obvious gotchas — rather than transcribing code the client can already grep. The goal is to orient a cold client so it doesn't burn calls rediscovering structure every session. Review the generated files before serving them, then start the server with --memory-dir ./memories.
Logging via RUST_LOG:
RUST_LOG=debug,rmcp=info ./target/release/code-mcpDefault level is info,rmcp=info. Press Ctrl-C for graceful shutdown (cancels the rmcp CancellationToken, drains active sessions).
Add to ~/.claude.json (or use claude mcp add):
{
"mcpServers": {
"code-mcp": {
"type": "http",
"url": "http://your-dev-box.lan:8080/"
}
}
}Other MCP clients with streamable-HTTP support (Cursor, Zed, etc.) take similar config.
cargo test # tools, scope, limiter, and gate-middleware tests
cargo clippy --all-targets -- -D warnings- No auth — by design (LAN deployment). For path scoping, use
--project. - The
regexcrate has no lookaround or backreferences. Patterns that need them won't compile and you'll get aninvalid_paramsMCP error. .gitignoreis honored only inside a directory tree that contains a.git/directory (this isignorecrate behavior, not ours).- Each parallel-walker worker keeps a thread-local
Stringbuffer and ships it to the main thread viampsc; counter isAtomicUsizewithfetch_add-based exact capping. There is noArc<Mutex<...>>on the hot path.