Self-hosted LLMs on a Swarm cluster, GPU or CPU
Most Open WebUI guides assume Docker Compose on a single box. I run my homelab on Docker Swarm, and getting OpenWebUI, Ollama, and ChromaDB working as Swarm services -- especially with a GPU -- took more digging than it should have. Here is the stack file I ended up with, plus the GPU configuration that actually makes Swarm schedule onto your NVIDIA card.
Overview#
Three services in one stack: OpenWebUI for the front end, Ollama for model hosting, and ChromaDB as the vector database for RAG. Everything below works on CPU too; the GPU parts are clearly marked so you can strip them out.
Prerequisites#
- Basic understanding of Docker Swarm.
- Docker Swarm configured on your host system.
- Proper directories or volumes created for data storage:
mkdir -p data/open-webui data/chromadb data/ollama - GPU support (optional) with NVIDIA Container Toolkit installed and configured.
Deployment Steps#
Step 1: Prepare Your Environment#
-
With GPU Support: Ensure your host system has GPU support configured:
- Enable CUDA for your OS and GPU.
- Install the NVIDIA Container Toolkit.
- Edit
/etc/docker/daemon.jsonto advertise the GPU to Swarm (get the UUID prefix fromnvidia-smi -a | grep UUID):{ "runtimes": { "nvidia": { "path": "nvidia-container-runtime", "runtimeArgs": [] } }, "default-runtime": "nvidia", "node-generic-resources": ["NVIDIA-GPU=GPU-<YOUR_GPU_UUID_PREFIX>"] } - Enable GPU resource advertising in
/etc/nvidia-container-runtime/config.tomlby uncommenting:swarm-resource = "DOCKER_RESOURCE_GPU" - Restart the Docker daemon:
sudo service docker restart
-
With CPU Support: Remove the
deploy.resources.reservations.generic_resourcesblock from theollamaservice in the stack file below. That is the only GPU-specific part.
Step 2: Configure the Docker Stack#
Below is a sample docker-stack.yaml file for deploying the three services:
version: '3.9'
services:
openWebUI:
image: ghcr.io/open-webui/open-webui:main
depends_on:
- chromadb
- ollama
volumes:
- ./data/open-webui:/app/backend/data
environment:
DATA_DIR: /app/backend/data
OLLAMA_BASE_URLS: http://ollama:11434
CHROMA_HTTP_PORT: 8000
CHROMA_HTTP_HOST: chromadb
CHROMA_TENANT: default_tenant
VECTOR_DB: chroma
WEBUI_NAME: Awesome ChatBot
CORS_ALLOW_ORIGIN: "*"
RAG_EMBEDDING_ENGINE: ollama
RAG_EMBEDDING_MODEL: nomic-embed-text-v1.5
RAG_EMBEDDING_MODEL_TRUST_REMOTE_CODE: "True"
ports:
- target: 8080
published: 8080
mode: ingress
deploy:
replicas: 1
restart_policy:
condition: any
delay: 5s
max_attempts: 3
chromadb:
hostname: chromadb
image: chromadb/chroma:0.5.15
volumes:
- ./data/chromadb:/chroma/chroma
environment:
- IS_PERSISTENT=TRUE
- ALLOW_RESET=TRUE
- PERSIST_DIRECTORY=/chroma/chroma
ports:
- target: 8000
published: 8000
mode: ingress
deploy:
replicas: 1
restart_policy:
condition: any
delay: 5s
max_attempts: 3
healthcheck:
test: ["CMD-SHELL", "curl localhost:8000/api/v1/heartbeat || exit 1"]
interval: 10s
retries: 2
start_period: 5s
timeout: 10s
ollama:
image: ollama/ollama:latest
hostname: ollama
ports:
- target: 11434
published: 11434
mode: ingress
deploy:
resources:
reservations:
generic_resources:
- discrete_resource_spec:
kind: "NVIDIA-GPU"
value: 0
replicas: 1
restart_policy:
condition: any
delay: 5s
max_attempts: 3
volumes:
- ./data/ollama:/root/.ollama
Step 3: Deploy the Stack#
The command is the same for GPU or CPU -- just make sure you modified the stack file first if you are CPU-only:
docker stack deploy -c docker-stack.yaml -d super-awesome-ai
