Monitoring and alerting #2
Labels
No labels
area/ci
area/media
area/network
area/observability
area/platform
area/security
area/storage
area/web
type/project
type/task
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
itguyeric/infra#2
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Highest-leverage item on the list. Three separate silent degradations were found by accident in a single day:
trusted.confprotected the arr web UIs while leaving/apiwide open, because the include sat insidelocation /.Every one of those looked healthy from outside. The common shape is that nothing was checking the thing that was believed.
Minimum useful version: host up/down, systemd unit failures, disk and thin-pool usage, and a check that each host's booted bootc digest matches what the pipeline last pushed. Tailscale can webhook on node-offline for roughly ten minutes of setup and would have caught one of the three today.
Blocks nothing formally, but makes DR, backups and the Kubernetes experiment far less scary.
Folding web analytics into this project: privacy-friendly, cookie-free analytics for the Hugo site with Umami or Plausible (self-hosted, behind SWAG, same Postgres pattern as everything else).
Also, given the Victoria Metrics contacts, worth evaluating VictoriaMetrics as the metrics backend/TSDB here instead of (or fronting) Prometheus, with Grafana on top. Companion pieces to slot in as we build this out: