# MCP Local RAG (MCP server)

> MCP Local RAG is an MCP server that adds search and knowledge tools to AI assistants such as Claude Desktop, Claude Code and Cursor. Easy-to-setup local RAG server with minimal configuration. It has 407 GitHub stars, is released under the MIT license and runs locally with npx -y mcp-local-rag.

English | 简体中文 | Deutsch | Español | Português (Brasil) | Français Search private documents from an MCP client or the terminal without sending them to an embedding API.

## Key facts

| Fact | Value |
|---|---|
| Repository | https://github.com/shinpr/mcp-local-rag |
| GitHub stars | 407 |
| License | MIT |
| Language | TypeScript |
| Transport | stdio |
| Packages | npm: mcp-local-rag |
| Remote URL | none |
| Needs API key | no |
| Official | no |
| Works with | Claude Desktop, Claude Code, Cursor, VS Code |
| Category | AI & search |
| Latest release | v0.21.0 (Oct 3, 2026) |
| Last commit | Oct 3, 2026 |
| MCP registry name | io.github.shinpr/mcp-local-rag |

## Install

### Claude Desktop (claude_desktop_config.json)

```json
{
  "mcpServers": {
    "mcp-local-rag": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-local-rag"
      ],
      "env": {
        "BASE_DIR": "your-value",
        "BASE_DIRS": "your-value",
        "DB_PATH": "your-value",
        "CACHE_DIR": "your-value",
        "HF_ENDPOINT": "your-value",
        "MODEL_NAME": "your-value"
      }
    }
  }
}
```

Settings > Developer > Edit Config. macOS: ~/Library/Application Support/Claude/, Windows: %APPDATA%\Claude\. Restart Claude Desktop afterwards.

### Claude Code

```sh
claude mcp add --env BASE_DIR=your-value --env BASE_DIRS=your-value --env DB_PATH=your-value --env CACHE_DIR=your-value --env HF_ENDPOINT=your-value --env MODEL_NAME=your-value --transport stdio mcp-local-rag -- npx -y mcp-local-rag
```

### Cursor (.cursor/mcp.json)

```json
{
  "mcpServers": {
    "mcp-local-rag": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "-y",
        "mcp-local-rag"
      ],
      "env": {
        "BASE_DIR": "your-value",
        "BASE_DIRS": "your-value",
        "DB_PATH": "your-value",
        "CACHE_DIR": "your-value",
        "HF_ENDPOINT": "your-value",
        "MODEL_NAME": "your-value"
      }
    }
  }
}
```

Project file; use ~/.cursor/mcp.json to enable it in every project.

### VS Code (.vscode/mcp.json)

```json
{
  "servers": {
    "mcp-local-rag": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "-y",
        "mcp-local-rag"
      ],
      "env": {
        "BASE_DIR": "your-value",
        "BASE_DIRS": "your-value",
        "DB_PATH": "your-value",
        "CACHE_DIR": "your-value",
        "HF_ENDPOINT": "your-value",
        "MODEL_NAME": "your-value"
      }
    }
  }
}
```

Config formats checked against the official docs on 2026-10-07.

## Environment variables

- `BASE_DIR`: Base directory for document storage (defaults to current working directory). Ignored when BASE_DIRS is set.
- `BASE_DIRS`: JSON array of base directories (e.g. '["/a","/b"]'). Takes precedence over BASE_DIR.
- `DB_PATH`: Path to LanceDB database directory (defaults to ./lancedb/)
- `CACHE_DIR`: Directory where Transformers.js models are cached (defaults to ./models/)
- `HF_ENDPOINT`: Hugging Face model download endpoint. Set this to a mirror URL when direct downloads are blocked (defaults to https://huggingface.co).
- `MODEL_NAME`: Embedding model name (defaults to Xenova/all-MiniLM-L6-v2)
- `MAX_FILE_SIZE`: Maximum file size in bytes (defaults to 104857600 / 100MB)
- `RAG_MAX_DISTANCE`: Maximum distance threshold for filtering search results. Results with distance greater than this value will be excluded. Lower values mean stricter filtering…
- `RAG_GROUPING`: Grouping mode for quality filtering. 'similar' returns only the most similar group (stops at first distance jump). 'related' includes related groups (stops at…
- `RAG_MAX_FILES`: Maximum number of files to keep in search results. Results are filtered to include only chunks from the top N best-scoring files. For example, 1 returns only…
- `CHUNK_MIN_LENGTH`: Minimum chunk length in characters (1-10000, defaults to 50). Chunks shorter than this threshold are filtered out during ingestion.
- `STORE_IMAGES`: Store supported PDF and DOCX images during ingestion and return them with matched chunks (defaults to false).
- `EMBED_TITLE_PREFIX`: Embed each chunk together with its document title, which can help when passages don't restate the topic the title names (defaults to false). After changing…
- `EMBED_HEADING_PREFIX`: Add section headings to chunk embeddings when they fit (defaults to false). Independent of EMBEDTITLEPREFIX. Re-ingest documents after changing it.
- `RAG_DEVICE`: Execution device for the embedder (defaults to cpu). Passed straight to ONNX Runtime; see the Transformers.js device source for the supported backend names…
- `RAG_DTYPE`: Embedding quantization dtype for the embedder (defaults to fp32). Opt-in and pass-through; accepts any dtype the chosen model provides (fp32, fp16, q8, int8…
- `RAG_HYBRID_WEIGHT`: Keyword boost factor for hybrid search (0.0-1.0, defaults to 0.6). 0 means semantic similarity only; higher values increase the keyword-match contribution to…
- `RAG_RERANK_CMD`: External reranker command template. Use {query} and {top} for query text and result count; unset disables reranking.
- `RAG_RERANK_TIMEOUT_MS`: Time budget per rerank call in milliseconds (100-600000, defaults to 10000). On timeout the spawned command is killed and the pre-rerank ordering is returned.

## Similar MCP servers

- [Context7](https://appsgit.com/mcp-servers/context7): Up-to-date code docs for any prompt. (62,752 stars, MIT, needs API key)
- [FunASR](https://appsgit.com/mcp-servers/funasr): Transcribe local audio with FunASR and SenseVoice using private, on-device inference. (20,596 stars, MIT)
- [Bifrost](https://appsgit.com/mcp-servers/bifrost): Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS. (8,591 stars, Apache-2.0)
- [Honcho](https://appsgit.com/mcp-servers/honcho): Memory that reasons: continual learning for stateful agents. (7,494 stars, AGPL-3.0, needs API key)
- [Semble](https://appsgit.com/mcp-servers/semble): Fast and Accurate Code Search for Agents. (6,185 stars, MIT)
- [Exa](https://appsgit.com/mcp-servers/exa): Connect AI agents to Exa for web search, content fetching, and multi-step research. (5,088 stars, MIT, official)

---

Canonical page: https://appsgit.com/mcp-servers/mcp-local-rag
Source: appsgit (https://appsgit.com), the app store for github. Data from the GitHub API, refreshed nightly.
Machine access: JSON API https://appsgit.com/api/v1/apps (OpenAPI: https://appsgit.com/openapi.json), MCP server https://mcp.appsgit.com/mcp, full index https://appsgit.com/llms-full.txt.
