Local-first code intelligence for OpenCode
Project Page · GitHub · Changelog
lgrep gives AI agents a better first move in codebases they already have on disk.
It is built for active local development, where files are changing, multiple agent sessions may be exploring the same repo, and the first challenge is often finding the right implementation before anyone knows the right symbol name.
Instead of starting with glob, grep, and random file reads, agents can:
- search by meaning when they do not know the symbol yet
- search by symbol when they know the name
- inspect file and repo structure before opening code
- reuse one warm local server across multiple sessions and agents
That is the whole pitch: fewer bad searches, less wasted context, faster understanding for both humans and agents working in local repos.
lgrep combines two complementary engines in one MCP server for local repositories:
- Semantic engine - natural-language code search using Voyage Code 3 embeddings with local LanceDB storage
- Symbol engine - exact symbol, outline, and text tools using tree-sitter parsing with a local JSON index
Use the semantic engine to answer questions like:
- "where is auth enforced between route and service?"
- "how does retry logic work for failed requests?"
- "where are permissions checked?"
Use the symbol engine to answer questions like:
- "find the
authenticatefunction" - "show me the outline for
src/auth.py" - "get the symbol source for
UserService.login"
Your working tree stays local and searchable as it evolves. Only short semantic queries and indexing payloads go to Voyage. Symbol lookup stays fully local and works without an API key.
AI coding agents usually fail early, not late.
They miss because they start with the wrong retrieval primitive:
grepcannot answer concept questions- full-file reads waste tokens on irrelevant code
- repeated local exploration across parallel agents duplicates work
lgrep fixes that by giving agents a search stack that matches how they actually reason:
- find the implementation by intent
- narrow to the right file or symbol
- retrieve only the code that matters
For heavy OpenCode users, this is not a convenience plugin. It is search infrastructure.
- Intent-first search - agents can ask by meaning before they know names
- Exact structure tools - file outlines, repo outlines, symbol lookup, and text search are in the same server
- Local-first storage - vectors and indexes live on disk, not in someone else's SaaS
- Shared warm process - one HTTP MCP server can serve multiple concurrent OpenCode sessions against the same local repos
- Commercially usable -
lgrepis MIT licensed, so commercial use is allowed
grep and rg are still the right tool for exact text and regex lookups. They are not good at intent discovery.
If the code says jwt.verify() and your agent asks "where is authentication enforced?", text search often misses the right entry point. Semantic search closes that gap.
mgrep is the closest semantic-search comparison point.
mgrepis semantic-onlylgrepcombines semantic search with symbol and structure toolsmgrepis cloud-orientedlgrepkeeps vectors local and shares one warm server across agents
lgrep is strongest when an agent is already inside a local repo and needs to move from intent to source without wasting context:
- ask by meaning when the right name is unknown
- inspect the matching files and outlines
- retrieve the exact symbol or text that matters
- keep the same warm local server available as the working tree changes
flowchart LR
A[OpenCode Session 1] --> M[lgrep MCP Server]
B[OpenCode Session 2] --> M
C[OpenCode Session N] --> M
M --> S[Semantic Engine\nVoyage Code 3 + LanceDB]
M --> Y[Symbol Engine\ntree-sitter + JSON index]
S --> V[(Local vector store)]
Y --> J[(Local symbol store)]
S -. query embeddings .-> Q[Voyage API]
Agent -> lgrep MCP server -> semantic engine + symbol engine
- Discover files while respecting
.gitignore - Chunk code with AST-aware boundaries
- Embed chunks with Voyage Code 3
- Store vectors locally in LanceDB
- Search with hybrid retrieval and reranking
- Parse source with tree-sitter
- Extract functions, classes, methods, and related structure
- Store a local symbol index
- Serve symbol search, outlines, and source retrieval without an API call
- Python 3.11+
- a Voyage API key if you want semantic search
pip install git+https://github.com/Sharper-Flow/lgrep.gitgit clone https://github.com/Sharper-Flow/lgrep.git
cd lgrep
pip install .stdio is the local default for single-session / single-user setups — no server process needed. For shared or multi-session deployments, see Scale-up: shared HTTP server below.
Create a key at dash.voyageai.com.
You only need this for the semantic engine. The symbol engine works without it.
For single-user, single-session setups, stdio is the local default. Add this to ~/.config/opencode/opencode.json:
{
"instructions": [
"~/.config/opencode/instructions/lgrep-tools.md"
],
"mcp": {
"lgrep": { "type": "local" }
}
}If you prefer to run a shared HTTP server (see section 3), swap the mcp.lgrep block for:
{
"mcp": {
"lgrep": {
"type": "remote",
"url": "http://localhost:6285/mcp",
"enabled": true
}
}
}Or let the installer wire the shared-HTTP path automatically:
lgrep install-opencodeThat installer will:
- create
~/.cache/lgrep/for indexes and logs - add a
type: "remote"MCP entry pointing athttp://localhost:6285/mcp - copy the packaged
lgrep-tools.mdinstruction andskills/lgrep/SKILL.mdinto your OpenCode config - append the instruction file to the
instructionsarray so agents preferlgrepfirst
To use stdio with lgrep install-opencode, run the installer first and then change mcp.lgrep in opencode.json to { "type": "local" }.
Important: the active agent must also expose lgrep_* tool definitions in its
tool manifest. If an agent profile only allows read/glob/grep, the model
cannot choose lgrep even when the MCP server is configured and the instruction
policy is present.
The installed files land at:
~/.config/opencode/instructions/lgrep-tools.md~/.config/opencode/skills/lgrep/SKILL.md
For shared or multi-session deployments, run lgrep as a persistent HTTP server instead of stdio:
VOYAGE_API_KEY=your-key \
LGREP_WARM_PATHS=/path/to/project-a:/path/to/project-b \
lgrep --transport streamable-http --host 127.0.0.1 --port 6285Why HTTP instead of stdio?
With stdio, each OpenCode session spawns its own server process. With streamable-http, one warm server handles all sessions. After starting the HTTP server, use the type: "remote" MCP config from section 2 above.
lgrep init-ignore /path/to/projectlgrep prune-orphans --dry-run
lgrep prune-orphans --execute --cache-dir /path/to/cacheprune-orphans is dry-run by default. Use --execute to actually delete orphaned semantic cache directories. --cache-dir overrides LGREP_CACHE_DIR for a single run. --execute and --dry-run are mutually exclusive; passing both exits with an error. Agents can call the same workflow via the lgrep_prune_orphans MCP tool listed in Symbol tools; that path also skips projects currently loaded in the running server.
Grace window. Recently modified cache dirs are preserved for 1 hour by default so the pruner cannot race a live indexer. Override with LGREP_PRUNE_MIN_AGE_S=<seconds> (0 disables grace entirely). The missing_meta and project_path_enoent reasons bypass the grace check because they are unambiguous.
Transport-aware MCP safety. When lgrep is reached over a shared transport (for example streamable-http), the MCP tool coerces dry_run=True regardless of the caller's request. Destructive prunes on shared deployments must go through the CLI (lgrep prune-orphans --execute) so the operator is explicit.
Each orphan is deleted independently. If shutil.rmtree fails for one entry (for example a lingering file lock or permission issue), the batch continues and the failure is recorded in the response under failures[] as {path, error}; the rest of the reclaim still lands. Re-run lgrep prune-orphans --execute after addressing the error, or inspect with --dry-run first to confirm the orphan is still present.
Deletion is refused for any path outside the resolved cache directory (path-confinement guard) and for any symlinked cache entry (TOCTOU guard) — both show up in failures[] rather than as successful deletes.
lgrep prune-symbols --dry-run
lgrep prune-symbols --execute --storage-dir /path/to/storageprune-symbols is dry-run by default. Use --execute to actually delete stale symbol-store index files (index_<hash>.json). --storage-dir overrides LGREP_SYMBOLS_DIR for a single run (default: ~/.cache/lgrep/symbols/). --execute and --dry-run are mutually exclusive; passing both exits with an error. Agents can call the same workflow via the lgrep_prune_symbols MCP tool listed in Symbol tools; that path also skips projects currently loaded in the running server.
Grace window. Recently modified index files are preserved for 1 hour by default so the pruner cannot race a live indexer. Override with LGREP_PRUNE_MIN_AGE_S=<seconds> (0 disables grace entirely). Only the unreadable_index_json reason is grace-eligible; the repo_path_enoent and missing_repo_path_field reasons bypass the grace check because they are unambiguous.
Transport-aware MCP safety. When lgrep is reached over a shared transport (for example streamable-http), the MCP tool coerces dry_run=True regardless of the caller's request. Destructive symbol-store prunes on shared deployments must go through the CLI (lgrep prune-symbols --execute) so the operator is explicit.
Each stale index is deleted independently. If an unlink fails for one entry (for example a lingering file lock or permission issue), the batch continues and the failure is recorded in the response under failures[] as {path, error}; the rest of the reclaim still lands. Re-run lgrep prune-symbols --execute after addressing the error, or inspect with --dry-run first to confirm the stale index is still present.
Deletion is refused for any path outside the resolved storage directory (path-confinement guard) and for any symlinked index file (TOCTOU guard) — both show up in failures[] rather than as successful deletes.
Typical OpenCode flow:
- Ask an intent question with
lgrep_search_semantic - Inspect structure with
lgrep_get_file_outlineorlgrep_get_repo_outline - Retrieve exact symbols with
lgrep_search_symbolsandlgrep_get_symbol
Examples:
lgrep_search_semantic(query="authentication flow", path="/path/to/project")
lgrep_get_file_outline(path="/path/to/project/src/auth.py")
lgrep_index_symbols_folder(path="/path/to/project")
lgrep_search_symbols(query="authenticate", path="/path/to/project")
lgrep_get_symbol(symbol_id="src/auth.py:function:authenticate", path="/path/to/project")
High-value prompts:
- "Where do we enforce auth between route and service?"
- "Find the
authenticatefunction" - "What are the main symbols in
src/auth.py?" - "Show me the repo structure around billing"
- "Find references to
verifyToken"
| Task | Best tool | Why |
|---|---|---|
| Intent or concept discovery | lgrep_search_semantic |
Search by meaning |
| Find a function or class by name | lgrep_search_symbols |
Exact symbol lookup |
| Inspect a single file's structure | lgrep_get_file_outline |
Fast AST outline |
| Inspect repo structure | lgrep_get_repo_outline |
Symbol-level overview |
| Find exact text or identifiers | lgrep_search_text or grep |
Literal match |
| Retrieve exact source for a symbol | lgrep_get_symbol |
Targeted code retrieval |
| Read a known file directly | Read |
No search needed |
As of 3.0.0, every lgrep MCP tool returns a structured dict matching a declared TypedDict in src/lgrep/server/responses.py. Clients should consume responses as native dicts — no json.loads is needed.
Example — lgrep_search_semantic:
{
"query": "authentication flow",
"path": "/path/to/project",
"engine": "hybrid",
"total": 3,
"results": [
{"file_path": "src/auth.py", "line_number": 42,
"content": "...", "score": 0.91,
"start_line": 42, "end_line": 87, "match_type": "hybrid"},
# ...
],
}engine is "hybrid" when hybrid=true (the default) or "vector" when hybrid=false.
Error responses use the shared ToolError shape:
{"error": "VOYAGE_API_KEY not set. Cannot perform semantic search."}Before 3.0.0, tools returned these objects as json.dumps(...) strings. If you upgrade from 2.x, remove any json.loads(response) wrappers on tool output. See the Upgrade from 2.x notes in the changelog for the full migration path.
| Tool | Purpose |
|---|---|
lgrep_search_semantic(query, path, limit=10, hybrid=true) |
Search code by meaning |
lgrep_index_semantic(path) |
Build or refresh a semantic index |
lgrep_status_semantic(path?) |
Show semantic index and watcher status |
lgrep_watch_start_semantic(path) |
Start background semantic re-indexing |
lgrep_watch_stop_semantic(path?) |
Stop the watcher |
| Tool | Purpose |
|---|---|
lgrep_index_symbols_folder(path, max_files=500, incremental=True) |
Index symbols in a local folder |
lgrep_index_symbols_repo(repo, ref="HEAD") |
Index symbols from a GitHub repo |
lgrep_list_repos() |
List indexed symbol repos |
lgrep_get_file_tree(path, max_files=500) |
Show repo file tree |
lgrep_get_file_outline(path) |
Show symbol outline for one file |
lgrep_get_repo_outline(path, max_files=500) |
Show symbol outline for a repo |
lgrep_search_symbols(query, path, limit=20, kind?) |
Search symbols by name |
lgrep_search_text(query, path, max_results=50) |
Search literal text |
lgrep_get_symbol(symbol_id, path) |
Retrieve one symbol |
lgrep_get_symbols(symbol_ids, path) |
Retrieve multiple symbols |
lgrep_invalidate_cache(path) |
Drop the symbol index for a repo |
lgrep_prune_orphans(dry_run=True) |
Report (or with dry_run=False, delete) orphan semantic cache dirs; skips active projects and the symbols/ cache |
lgrep_prune_symbols(dry_run=True) |
Report (or with dry_run=False, delete) stale symbol-store index files; skips active projects and non-local github: entries |
Symbol IDs use this deterministic format:
file_path:kind:name
Examples:
src/auth.py:function:authenticate
src/auth.py:class:AuthManager
src/auth.py:method:login
lgrep supports both stdio and streamable-http. Use stdio for the
local single-session default; use streamable-http only when you intentionally
want one shared local daemon for multiple OpenCode/Vision sessions.
lgrep --transport streamable-http --host 127.0.0.1 --port 6285Security notes:
- default host is
127.0.0.1 - there is no built-in auth layer on the HTTP transport
lgrepdoes not set CORS headers, and browser-based clients should not connect directly to the streamable HTTP endpoint- if you do put it behind a proxy, enforce your own authentication and origin controls there
- exposing
0.0.0.0is a non-default, explicit opt-in; do not do it without a reverse proxy or firewall
| Variable | Required | Default | Description |
|---|---|---|---|
VOYAGE_API_KEY |
For semantic search | none | Voyage API key |
LGREP_LOG_LEVEL |
No | INFO |
Log verbosity |
LGREP_CACHE_DIR |
No | ~/.cache/lgrep |
Cache directory |
LGREP_WARM_PATHS |
No | none | Colon-separated projects to warm on startup |
LGREP_AUTO_WARM_DISK |
No | true |
Auto-load all discoverable disk caches on startup when no explicit warm paths are set. Set false for large shared machines. |
LGREP_AUTO_WATCH |
No | false |
Auto-start file watchers for warmed projects |
LGREP_TOOL_TIMEOUT_S |
No | 45 |
Per-tool server-side timeout (seconds). Bounds each MCP tool invocation. |
LGREP_WORKER_MAX_THREADS |
No | 4 |
Max worker threads for supervised blocking daemon jobs. |
LGREP_PRUNE_MIN_AGE_S |
No | 3600 |
Grace window (seconds) before prune-orphans will treat an ambiguous orphan (unreadable meta / missing chunks) as prunable. 0 disables grace. |
LGREP_SYMBOLS_DIR |
No | ~/.cache/lgrep/symbols |
Symbol index storage directory used by lgrep index-symbols and lgrep prune-symbols. |
LGREP_WORKTREE_DEDUP |
No | unset | When set (any value), git worktrees sharing a common .git directory resolve to the same semantic cache key, eliminating duplicate embeddings and disk usage across worktrees. |
LGREP_TRANSPORT |
No (auto-set) | unset | Transport kind (stdio/streamable-http) populated by lgrep run_server. Tools use this to apply transport-aware safety. Do not set manually. |
For agent-heavy local setups that route lgrep through a Vision-managed MCP server, prefer explicit warm paths over warming every cached repository:
lgrep:
port: 6278
command: /home/you/.local/bin/lgrep
env:
VOYAGE_API_KEY: "${VOYAGE_API_KEY}"
LGREP_WORKTREE_DEDUP: "1"
LGREP_WARM_PATHS: "/home/you/dev/primary:/home/you/dev/tooling"
LGREP_AUTO_WARM_DISK: "false"
LGREP_TOOL_TIMEOUT_S: "8"
LGREP_WORKER_MAX_THREADS: "4"LGREP_WORKTREE_DEDUP=1avoids duplicate semantic caches for git worktrees.LGREP_WARM_PATHSshould name only repos agents actively search.LGREP_AUTO_WARM_DISK=falseprevents surprise startup work from old cache entries.- Set
LGREP_TOOL_TIMEOUT_Sbelow the MCP proxy/client timeout so callers get a structured lgrep error before a transport deadline. - Keep
LGREP_WORKER_MAX_THREADSsmall for shared daemons so concurrent agents cannot create unbounded blocking work. - Use
lgrep_diagnosticswhen investigating high CPU/thread count. It reports PID, uptime, loaded projects, worker limit, active jobs, recent abandoned/finished jobs, and full local project paths without exposing API keys or environment values. lgrep_status_semantic(path="")is intentionally cheap and memory-only. Pass a specificpathwhen you need deep file/chunk counts.- Destructive cache cleanup over shared HTTP is forced to dry-run; run
lgrep prune-orphans --execute(orlgrep prune-symbols --executefor symbol indexes) from a local shell when an operator intentionally wants deletion.
Agent fallback rule: if a default hybrid lgrep_search_semantic call times out
or hits a deadline, retry once with hybrid:false and a small limit such as
limit=5, then fall back to lgrep_search_symbols, lgrep_search_text, or
direct file reads.
.gitignoreis respected automatically.lgrepignorelets you exclude additional paths
Example .lgrepignore entries:
src/generated/
docs/site/
*.test.data
Resource use depends on repository size, enabled engines, and cache history:
- RAM/CPU - the shared server keeps warmed project state available and does most local work during indexing or search requests
- Disk - semantic vectors and symbol indexes are stored locally and grow with indexed projects
- Network - semantic indexing and semantic queries call Voyage; symbol lookup, outlines, and text search stay local
- Semantic engine - AST-aware chunking for 30+ languages, with text fallback when needed
- Symbol engine - tree-sitter-language-pack support across 165+ languages
When using git worktrees (e.g., ADV's per-change worktree isolation), multiple checkouts of the same repository can accumulate duplicate semantic indexes — one per worktree path. This wastes disk space (hundreds of MB per worktree) and Voyage API tokens.
Enable worktree dedup by setting LGREP_WORKTREE_DEDUP=1 in your environment:
export LGREP_WORKTREE_DEDUP=1With this flag, lgrep resolves each project path through git rev-parse --git-common-dir and uses the repository root (parent of .git) as the cache key. All worktrees of the same repository share:
- One LanceDB cache directory (disk dedup)
- One
ProjectStateobject in memory (memory dedup — RSS stays bounded regardless of worktree count) - One
project_meta.jsonwithalias_pathstracking every worktree that uses the canonical cache
Concurrency: Cross-process alias updates to project_meta.json are guarded by a POSIX advisory lock (fcntl.flock) on a dedicated .meta.lock file, so simultaneous writes from multiple lgrep instances do not lose alias entries.
Tradeoff: When dedup is enabled, stale-file cleanup during indexing is skipped. Deleted-file chunks remain in the shared cache until a full rebuild. This is benign (extra search results, not wrong results) and avoids cross-worktree corruption.
ADV integration: Call the invalidate_worktree_cache MCP tool during /adv-archive Phase 9 (before adv_worktree_delete) to remove the worktree's alias from the shared cache metadata:
invalidate_worktree_cache(paths: ["/path/to/worktree"])
Garbage collection: Run lgrep gc --execute periodically (or via systemd timer). This runs three passes:
prune_orphans— deletes whole cache directories whose project root no longer exists on diskgc_worktree_meta— removes stale alias entries fromproject_meta.jsonfiles (worktrees that were deleted without callinginvalidate_worktree_cache)prune_symbols— deletes stale symbol-store index files (index_<hash>.json) whoserepo_pathis missing, unreadable, or absent from the JSON
The prune_orphans and gc_worktree_meta passes respect the 1-hour grace window (configurable via LGREP_PRUNE_MIN_AGE_S) and skip active in-memory projects. The prune_symbols pass respects the same grace window, but only for the unreadable_index_json reason; the repo_path_enoent and missing_repo_path_field reasons bypass grace, and non-local github: entries are skipped.
VOYAGE_API_KEY not set
Set the key in your MCP server environment. The symbol engine still works without it.
Slow first semantic index
The first run embeds the whole project. Later runs skip unchanged files using hashes.
Repository not indexed from symbol tools
Run lgrep_index_symbols_folder(path=...) first.
Stale semantic results
lgrep_search_semantic runs an auto-staleness check before every search. When
file mtimes have moved past the index timestamp and content hashes have drifted,
the search serves the current (possibly slightly stale) index immediately and
triggers a background single-flight refresh — it never blocks on a full re-embed.
The next search observes fresh results; freshness converges automatically with no
operator configuration. lgrep_watch_start_semantic(...) (LGREP_AUTO_WATCH) is an
incremental-freshness optimization (per-file background re-index on edit), not a
correctness dependency. lgrep_index_semantic(...) is only needed for first-time
setup or to force a guaranteed-fresh refresh before a specific query.
Native dependency build issues
lgrep depends on packages with native extensions. Prebuilt wheels usually work; otherwise install the required compiler toolchain.
The semantic tools were renamed in v2.0.0:
| v1.x | v2.x |
|---|---|
lgrep_search |
lgrep_search_semantic |
lgrep_index |
lgrep_index_semantic |
lgrep_status |
lgrep_status_semantic |
lgrep_watch_start |
lgrep_watch_start_semantic |
lgrep_watch_stop |
lgrep_watch_stop_semantic |
git clone https://github.com/Sharper-Flow/lgrep.git
cd lgrep
pip install -e ".[dev]"
pytest -vMIT - see LICENSE.