Vmlx
Vmlx is an MCP server that adds search and knowledge tools to AI assistants such as Claude Desktop, Claude Code and Cursor. vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm. It has 886 GitHub stars, is released under the Apache-2.0 license and runs locally with uvx vmlx.
- AI & search
- Apache-2.0
- Actively maintained
Install Vmlx
Generated from the server's published package. Replace your-value with your own values.
Claude Desktop
{
"mcpServers": {
"vmlx": {
"command": "uvx",
"args": [
"vmlx"
]
}
}
}Settings > Developer > Edit Config. macOS: ~/Library/Application Support/Claude/, Windows: %APPDATA%\Claude\. Restart Claude Desktop afterwards.
Claude Code
claude mcp add --transport stdio vmlx -- uvx vmlxCursor
{
"mcpServers": {
"vmlx": {
"type": "stdio",
"command": "uvx",
"args": [
"vmlx"
]
}
}
}Project file; use ~/.cursor/mcp.json to enable it in every project.
VS Code
{
"servers": {
"vmlx": {
"type": "stdio",
"command": "uvx",
"args": [
"vmlx"
]
}
}
}Config formats checked against the official docs on Oct 7, 2026: modelcontextprotocol.io (opens in a new tab), code.claude.com (opens in a new tab), cursor.com (opens in a new tab), code.visualstudio.com (opens in a new tab).
About Vmlx
Self-hosted inference server for LLMs, VLMs, and image generation on Apple Silicon. OpenAI + Anthropic + Ollama compatible HTTP API. Self-hosted; no third-party API keys required. Native MTP artifact detection and family-specific cache policy gates keep speculative/cache settings explicit and model-safe.
- llm
- lmstudio
- macbook
- mlx
- mlxllm
- mlxstudio
- anthropic-api
- kvcache-compression
- kvcache-optimization
- kvcache-reuse
Similar MCP servers
More ai & search MCP servers
Context7
upstash/context7
Up-to-date code docs for any prompt.
AI & searchTypeScriptFunASR
modelscope/FunASR
Transcribe local audio with FunASR and SenseVoice using private, on-device inference.
AI & searchPythonBifrost
maximhq/bifrost
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
AI & searchGoHoncho
plastic-labs/honcho
Memory that reasons: continual learning for stateful agents.
AI & searchPythonSemble
MinishLab/semble
Fast and Accurate Code Search for Agents.
AI & searchPythonExa
exa-labs/exa-mcp-server
Connect AI agents to Exa for web search, content fetching, and multi-step research.
OfficialAI & searchTypeScript
What is Vmlx?
Vmlx is an MCP server that adds search and knowledge tools to AI assistants such as Claude Desktop, Claude Code and Cursor. vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm. It has 886 GitHub stars, is released under the Apache-2.0 license and runs locally with uvx vmlx. The source code is at github.com/jjang-ai/vmlx.
How do I install the Vmlx MCP server?
Add the command uvx vmlx to your MCP client: put it in claude_desktop_config.json for Claude Desktop, run claude mcp add for Claude Code, or add it to .cursor/mcp.json (Cursor) or .vscode/mcp.json (VS Code). The snippets on this page are ready to paste.
Is Vmlx free?
The server is open source under the Apache-2.0 license, so running it is free. It does not declare any required API key.
Is Vmlx actively maintained?
The most recent commit was on Oct 6, 2026. The latest release is v1.6.75, published Oct 6, 2026. appsgit only lists MCP servers with a commit in the last six months and re-checks every server daily.