Self-hosted AI: local LLMs with Ollama + Open WebUI #33

Open
opened 2026-08-19 02:26:06 +00:00 by itguyeric · 1 comment
Owner

Run local models on the fleet instead of reaching for a cloud API. Ollama to manage and serve models, Open WebUI as the chat front end. Uses the RTX for inference, so plan GPU allocation since Plex and a possible Immich ML container also want it.

Doubles as the local engine for other services: Paperless-AI extraction (#28) and anything we wire up later. New bootc image with GPU passthrough like the Plex box, behind SWAG, LAN-only or over Tailscale.

Run local models on the fleet instead of reaching for a cloud API. Ollama to manage and serve models, Open WebUI as the chat front end. Uses the RTX for inference, so plan GPU allocation since Plex and a possible Immich ML container also want it. Doubles as the local engine for other services: Paperless-AI extraction (#28) and anything we wire up later. New bootc image with GPU passthrough like the Plex box, behind SWAG, LAN-only or over Tailscale.
itguyeric added this to the Homelab project 2026-08-19 02:26:06 +00:00
Author
Owner

Blocked by: Need AI-capable server

Blocked by: Need AI-capable server
Sign in to join this conversation.
No description provided.