Skip to content
appsgit

Deploy guide

How to self-host Ollama with Docker Compose

Run local LLMs with Ollama and Docker Compose: CPU or NVIDIA GPU setup, pulling models, securing the API behind HTTPS with auth, backups and upgrades.

  • Updated
  • Intermediate
  • About 15 minutes

You will need

  • 4 vCPU / 8 GB RAM for 7-8B models (16 GB+ for larger ones)
  • Docker + Docker Compose v2
  • Optional: NVIDIA GPU with the NVIDIA Container Toolkit
  • Disk space for models (4-40 GB each)

What is Ollama?

Ollama is a tool for running large language models such as Llama, Gemma, Qwen, Mistral and DeepSeek locally, with a simple CLI and an HTTP API. It is open source under the MIT license, handles model downloads and quantized formats for you, and exposes both its own API and an OpenAI-compatible endpoint. It is the backend behind many self-hosted AI tools, including Open WebUI.

Requirements

  • A Linux server with at least 4 vCPU and 8 GB of RAM for 7-8B models.
  • Docker Engine and Docker Compose v2.
  • For GPU acceleration: an NVIDIA GPU with current drivers and the NVIDIA Container Toolkit, or an AMD GPU supported by ROCm.
  • Fast disk space for models.

Step 1: Prepare the server

This guide assumes Ubuntu 24.04 with Docker installed. If you need Docker, follow the Docker Engine install guide. For NVIDIA, check that nvidia-smi works on the host and that you have run sudo nvidia-ctk runtime configure --runtime=docker followed by a Docker restart.

mkdir -p ~/ollama && cd ~/ollama

Step 2: Create the Docker Compose file

Save this as docker-compose.yml:

services:
  ollama:
    image: ollama/ollama:0.35.1
    container_name: ollama
    restart: unless-stopped
    ports:
      - "127.0.0.1:11434:11434"
    environment:
      OLLAMA_KEEP_ALIVE: "10m"
      OLLAMA_MAX_LOADED_MODELS: "1"
      OLLAMA_NUM_PARALLEL: "2"
    volumes:
      - ./models:/root/.ollama
    # NVIDIA GPU: remove the comment markers below.
    # deploy:
    #   resources:
    #     reservations:
    #       devices:
    #         - driver: nvidia
    #           count: all
    #           capabilities: [gpu]

The port is bound to 127.0.0.1 on purpose. Ollama has no login or API key, so publishing 11434 on a public interface lets anyone on the internet use your hardware. OLLAMA_KEEP_ALIVE controls how long a model stays in memory after the last request. OLLAMA_NUM_PARALLEL sets how many requests one model serves at the same time, and each parallel slot uses extra memory. For AMD GPUs, use the ollama/ollama:0.35.1-rocm image and pass /dev/kfd and /dev/dri as devices instead.

There are no secrets in this file. You add authentication at the reverse proxy in Step 4. Pin the image to a version from the releases page, or use latest if you prefer automatic upgrades.

Step 3: Start and open the app

docker compose up -d
docker compose exec ollama ollama pull llama3.2
docker compose exec ollama ollama run llama3.2 "Say hello in five words."

Ollama has no web interface of its own. Check the API from the server:

curl http://127.0.0.1:11434/api/tags

The response lists your installed models. Browse the Ollama model library for more, and use docker compose exec ollama ollama ps to see which models are loaded and whether they run on GPU or CPU. For a chat interface, add Open WebUI (see our Open WebUI guide).

Step 4: Put it behind HTTPS

Only expose Ollama beyond the server if a remote client really needs it, and always with authentication. With Caddy, generate a password hash with caddy hash-password, then:

ollama.example.com {
    basic_auth {
        apiuser PASTE_HASH_FROM_caddy_hash-password
    }
    reverse_proxy 127.0.0.1:11434
}

Clients then call https://apiuser:password@ollama.example.com, or send an Authorization: Basic header. Better still, keep Ollama private and reach it over a VPN such as WireGuard or Tailscale. Nginx Proxy Manager can do the same with an Access List.

Backups and upgrades

The ./models folder holds downloaded models and your custom Modelfiles. Models can always be pulled again, so backups are optional: save the output of ollama list and any custom Modelfiles you wrote.

To upgrade, change the version tag (or keep latest), then:

docker compose pull && docker compose up -d

New Ollama releases often add support for new model architectures, so upgrade before pulling a brand-new model.

Troubleshooting

  • Models run on CPU despite a GPU: the NVIDIA Container Toolkit is missing or the deploy block is commented out. docker compose exec ollama nvidia-smi must list your GPU.
  • "model requires more system memory": the model does not fit. Pick a smaller or more heavily quantized variant, for example a q4 tag.
  • Other containers cannot reach Ollama: on the same Compose network use http://ollama:11434, not localhost.
  • First response is slow: the model is loading into memory. Raise OLLAMA_KEEP_ALIVE to keep it warm.

Next steps

Connect Open WebUI for a chat interface, try embedding models for RAG, write a Modelfile to customise system prompts, and point editors or n8n at the OpenAI-compatible endpoint at /v1.

Spotted something out of date? Tell us and we will update the guide.

FAQ

Ollama questions

Still curious? Email info@appsgit.com.

What port does Ollama use?

Ollama's HTTP API listens on port 11434. Clients such as Open WebUI, editors and scripts send requests to that port.

Is Ollama free?

Yes. Ollama is free and open source under the MIT license. The models you download have their own licenses, so check each model's terms for commercial use.

Does Ollama have authentication?

No. The Ollama API has no built-in authentication, so anyone who can reach port 11434 can use your models. Keep it on localhost or a private network, or put a reverse proxy with authentication in front.

How much RAM do I need for Ollama?

As a rule of thumb, 8 GB of RAM or VRAM runs 7-8B models, 16 GB runs 13-14B models, and 32 GB or more is needed for 30B-class models at 4-bit quantization.

Can Ollama run without a GPU?

Yes. Ollama runs on CPU only, but responses are much slower than on a GPU. Small models in the 1-4B range are usable on modern CPUs.

Ollama vs LM Studio?

Ollama is a headless server with an API, built for Docker, servers and integrations. LM Studio is a desktop app with a built-in chat UI; both can expose an OpenAI-compatible endpoint.