LLMKube
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and….
- AI & LLM Tools
- Go
- Apache-2.0
- Docker
- Actively maintained
GitHub stats
- GitHub stars
- 231
- Forks
- 34
- Open issues
- 101
- Commits, last 12 months
- 1,039
- Last commit
- Oct 3, 2026
- Latest release
- v0.10.1Sep 30, 2026
- Created
- Nov 12, 2025
Data from the GitHub API, refreshed daily.
About LLMKube
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and…. LLMKube is open source under the Apache-2.0 license and written mainly in Go, so you can read the code, run it on your own hardware and keep full control of your data.
The project has 231 stars and 34 forks on GitHub, with 1,039 commits over the last year. It ships a Docker image, which makes it a good fit for a small VPS or a home server running Docker Compose.
Topics
- ai
- gguf
- gpu
- inference
- kubernetes
- kubernetes-operator
- llama-cpp
- llm
- local-llm
- nvidia
Repository preview
Similar apps
More apps like LLMKube
Other projects in AI & LLM Tools, ranked by GitHub stars.
Ollama
ollama/ollama
Get up and running with Llama 3.3, DeepSeek-R1, Phi-4, Gemma 3, and other large language models.
AI & LLM ToolsGoDockerStable Diffusion WebUI
AUTOMATIC1111/stable-diffusion-webui
Browser interface for Stable Diffusion image generation.
AI & LLM ToolsPythonLangflow
langflow-ai/langflow
Visual framework for building and deploying AI agents and RAG workflows.
AI & LLM ToolsPythonDockerOpen-WebUI
open-webui/open-webui
User-friendly AI Interface, supports Ollama, OpenAI API.
AI & LLM ToolsPythonDockerComfyUI
Comfy-Org/ComfyUI
Node-based interface and engine for diffusion image and video generation.
AI & LLM ToolsPythonRAGFlow
infiniflow/ragflow
Open-source RAG engine with deep document understanding.
AI & LLM ToolsGoDocker
What is LLMKube?
LLMKube is a free, open-source app you can host on your own server, listed in our AI & LLM Tools category. Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and…. The source code lives at github.com/defilantech/LLMKube, where it has 231 stars.
Is LLMKube free and open source?
Yes. LLMKube is open-source software released under the Apache-2.0 license, so you can self-host it at no cost. You only pay for the server it runs on, and some projects also offer an optional paid cloud plan.
Can I run LLMKube with Docker?
Yes. LLMKube provides a Docker image, so you can run it on any Linux server with Docker and Docker Compose installed. The project README lists the image name and the environment variables to set.
What are the best alternatives to LLMKube?
Popular open-source alternatives to LLMKube include Ollama, Stable Diffusion WebUI and Langflow.
Is LLMKube actively maintained?
The most recent commit to LLMKube was on Oct 3, 2026 with 1,039 commits in the past twelve months. The latest release is v0.10.1, published Sep 30, 2026. Check the open issues on GitHub to gauge how quickly maintainers respond.