# How to self-host Ollama with Docker Compose

> Run local LLMs with Ollama and Docker Compose: CPU or NVIDIA GPU setup, pulling models, securing the API behind HTTPS with auth, backups and upgrades.

## Key facts

| Fact | Value |
|---|---|
| App | Ollama (https://appsgit.com/apps/ollama) |
| Difficulty | intermediate |
| Time | about 15 minutes |
| Requirements | 4 vCPU / 8 GB RAM for 7-8B models (16 GB+ for larger ones); Docker + Docker Compose v2; Optional: NVIDIA GPU with the NVIDIA Container Toolkit; Disk space for models (4-40 GB each) |
| Last updated | 2026-10-06 |

## What is Ollama?

Ollama is a tool for running large language models such as Llama, Gemma, Qwen, Mistral and DeepSeek locally, with a simple CLI and an HTTP API. It is open source under the MIT license, handles model downloads and quantized formats for you, and exposes both its own API and an OpenAI-compatible endpoint. It is the backend behind many self-hosted AI tools, including Open WebUI.

## Requirements

- A Linux server with at least 4 vCPU and 8 GB of RAM for 7-8B models.
- Docker Engine and Docker Compose v2.
- For GPU acceleration: an NVIDIA GPU with current drivers and the [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html), or an AMD GPU supported by ROCm.
- Fast disk space for models.

## Step 1: Prepare the server

This guide assumes Ubuntu 24.04 with Docker installed. If you need Docker, follow the [Docker Engine install guide](https://docs.docker.com/engine/install/ubuntu/). For NVIDIA, check that `nvidia-smi` works on the host and that you have run `sudo nvidia-ctk runtime configure --runtime=docker` followed by a Docker restart.

```bash
mkdir -p ~/ollama && cd ~/ollama
```

## Step 2: Create the Docker Compose file

Save this as `docker-compose.yml`:

```yaml
services:
  ollama:
    image: ollama/ollama:0.35.1
    container_name: ollama
    restart: unless-stopped
    ports:
      - "127.0.0.1:11434:11434"
    environment:
      OLLAMA_KEEP_ALIVE: "10m"
      OLLAMA_MAX_LOADED_MODELS: "1"
      OLLAMA_NUM_PARALLEL: "2"
    volumes:
      - ./models:/root/.ollama
    # NVIDIA GPU: remove the comment markers below.
    # deploy:
    #   resources:
    #     reservations:
    #       devices:
    #         - driver: nvidia
    #           count: all
    #           capabilities: [gpu]
```

The port is bound to `127.0.0.1` on purpose. Ollama has no login or API key, so publishing 11434 on a public interface lets anyone on the internet use your hardware. `OLLAMA_KEEP_ALIVE` controls how long a model stays in memory after the last request. `OLLAMA_NUM_PARALLEL` sets how many requests one model serves at the same time, and each parallel slot uses extra memory. For AMD GPUs, use the `ollama/ollama:0.35.1-rocm` image and pass `/dev/kfd` and `/dev/dri` as devices instead.

There are no secrets in this file. You add authentication at the reverse proxy in Step 4. Pin the image to a version from the [releases page](https://github.com/ollama/ollama/releases), or use `latest` if you prefer automatic upgrades.

## Step 3: Start and open the app

```bash
docker compose up -d
docker compose exec ollama ollama pull llama3.2
docker compose exec ollama ollama run llama3.2 "Say hello in five words."
```

Ollama has no web interface of its own. Check the API from the server:

```bash
curl http://127.0.0.1:11434/api/tags
```

The response lists your installed models. Browse the [Ollama model library](https://ollama.com/library) for more, and use `docker compose exec ollama ollama ps` to see which models are loaded and whether they run on GPU or CPU. For a chat interface, add Open WebUI (see our Open WebUI guide).

## Step 4: Put it behind HTTPS

Only expose Ollama beyond the server if a remote client really needs it, and always with authentication. With [Caddy](https://caddyserver.com/docs/), generate a password hash with `caddy hash-password`, then:

```caddyfile
ollama.example.com {
    basic_auth {
        apiuser PASTE_HASH_FROM_caddy_hash-password
    }
    reverse_proxy 127.0.0.1:11434
}
```

Clients then call `https://apiuser:password@ollama.example.com`, or send an `Authorization: Basic` header. Better still, keep Ollama private and reach it over a VPN such as WireGuard or Tailscale. Nginx Proxy Manager can do the same with an Access List.

## Backups and upgrades

The `./models` folder holds downloaded models and your custom Modelfiles. Models can always be pulled again, so backups are optional: save the output of `ollama list` and any custom Modelfiles you wrote.

To upgrade, change the version tag (or keep `latest`), then:

```bash
docker compose pull && docker compose up -d
```

New Ollama releases often add support for new model architectures, so upgrade before pulling a brand-new model.

## Troubleshooting

- **Models run on CPU despite a GPU:** the NVIDIA Container Toolkit is missing or the `deploy` block is commented out. `docker compose exec ollama nvidia-smi` must list your GPU.
- **"model requires more system memory":** the model does not fit. Pick a smaller or more heavily quantized variant, for example a `q4` tag.
- **Other containers cannot reach Ollama:** on the same Compose network use `http://ollama:11434`, not `localhost`.
- **First response is slow:** the model is loading into memory. Raise `OLLAMA_KEEP_ALIVE` to keep it warm.

## Next steps

Connect Open WebUI for a chat interface, try embedding models for RAG, write a Modelfile to customise system prompts, and point editors or n8n at the OpenAI-compatible endpoint at `/v1`.

## FAQ

### What port does Ollama use?

Ollama's HTTP API listens on port 11434. Clients such as Open WebUI, editors and scripts send requests to that port.

### Is Ollama free?

Yes. Ollama is free and open source under the MIT license. The models you download have their own licenses, so check each model's terms for commercial use.

### Does Ollama have authentication?

No. The Ollama API has no built-in authentication, so anyone who can reach port 11434 can use your models. Keep it on localhost or a private network, or put a reverse proxy with authentication in front.

### How much RAM do I need for Ollama?

As a rule of thumb, 8 GB of RAM or VRAM runs 7-8B models, 16 GB runs 13-14B models, and 32 GB or more is needed for 30B-class models at 4-bit quantization.

### Can Ollama run without a GPU?

Yes. Ollama runs on CPU only, but responses are much slower than on a GPU. Small models in the 1-4B range are usable on modern CPUs.

### Ollama vs LM Studio?

Ollama is a headless server with an API, built for Docker, servers and integrations. LM Studio is a desktop app with a built-in chat UI; both can expose an OpenAI-compatible endpoint.

Prefer not to do it yourself? [appsgit installation help](https://appsgit.com/services/install) installs it on your server for a fixed quote.

---

Canonical page: https://appsgit.com/guides/ollama
Source: appsgit (https://appsgit.com), the app store for github. Data from the GitHub API, refreshed nightly.
Machine access: JSON API https://appsgit.com/api/v1/apps (OpenAPI: https://appsgit.com/openapi.json), MCP server https://mcp.appsgit.com/mcp, full index https://appsgit.com/llms-full.txt.
