Tooling & maintenance

Tooling & maintenance

How to update the automation behind the AV systems. Most of it lives in the repository under Automation/, and runs on docker-server, which sits on all the AV VLANs and is the operations hub (monitoring, config backups, fleet management).

Each tool has a detailed README next to its code; this page is the map and the “how do I change it” workflow. The golden rule: change things in git, not on the box.

The monitoring stack (Grafana, Prometheus, Companion, Uptime-Kuma)

docker-server runs these as a Docker Compose project called services, and git is the single source of truth. GitHub notifies docker-server the moment Live changes — a webhook, published through Tailscale Funnel — and the box reconciles the running stack to match the repo within seconds. A systemd timer (spactech-deploy.timer) also checks Live every ~2 minutes, so a missed notification only costs time. There is no Portainer and nothing to edit on the box — a merged commit is the deploy.

Everything for it is in Automation/server/; the mechanics and the data-safe runbook are in Automation/server/deploy/README.md.

To change it: edit under Automation/server/, open a PR, merge to Live. It deploys itself within ~2 minutes. Common changes:

  • Add a device to monitor — add a target to Automation/server/prometheus.yml (copy an existing job_name block). node_exporter must be running on the target (see Signage fleet below for the bare-metal install). The deploy validates the new file and recreates the Prometheus container to load it — a SIGHUP would re-read a stale copy, because prometheus.yml is a single-file bind mount.
  • Change a service version — bump the image: tag in Automation/server/docker-compose.yml.
  • Dashboards & alerts — these live in Grafana’s own volume and are edited in the Grafana UI; the repo keeps exported copies under Automation/server/backups/grafana/. Re-export after UI changes so the repo stays current.

Prometheus keeps 180 days of full-resolution data, with no size cap (the --storage.tsdb.retention.time flag in the compose file). At roughly 1.3 GB per 15 days that grows to about 16 GB, so it is worth watching docker-server’s free space until it levels off. Prometheus drops data older than the limit once it gets there.

Deploy now instead of waiting (on docker-server): ~/SPACTech/Automation/server/deploy/deploy.sh

Roll back: revert the commit — it deploys like any other merge. Data volumes are external, so container recreation never touches Grafana/Prometheus/Kuma/Companion data. Volume backups and a restore recipe are in the deploy README.

Host-level setup on docker-server

The services stack deploys itself, but docker-server’s own host packages and network config do not — they are applied by script and captured here:

  • setup-docker-server-comms-lldp.sh — two fixes that both need root on the box. It gives the comms VLAN interface an address (the EdgeRouter already reserves 10.0.24.200 for it; the interface just never asked over DHCPv4), and installs lldpd scoped to the enp3s0f0 trunk so the box identifies itself in the switches’ LLDP neighbour tables instead of appearing as an anonymous trunk port. Run it on docker-server with sudo; DRY_RUN=1 shows what it would do. It writes a netplan drop-in rather than editing 00-installer-config.yaml, so it is reversible by deleting one file.
  • setup-docker-server-rx-buffers.sh — receive buffers sized for SVSi streams, both needing root. It raises net.core.rmem_max to 8 MB through /etc/sysctl.d, so a socket that asks for a large receive buffer gets one, and grows the enp3s0f0 RX ring from 200 to 511 descriptors with a oneshot unit (spactech-rx-ring.service) that runs whenever the NIC appears. Without them, anything receiving a 1080p60 stream here lost a burst of packets every few seconds; see Receiving a stream on docker-server. Run it with sudo; DRY_RUN=1 shows what it would do. Resizing the ring resets the NIC (traffic pauses for under 0.1 s), so it skips that when the ring is already 511.
  • deploy/install.sh — installs the deploy’s systemd units (the timer, the GitHub webhook and the path unit it drives) and publishes the webhook through Tailscale Funnel, then sends it a signed push to prove the chain deploys. Re-run it with sudo whenever a unit file in Automation/server/deploy/ changes — a merge updates the checkout, never /etc/systemd/system. webhook.py reloads on its own.

The signage appliance image

Reproducible Ubuntu image for a new/replacement digital-signage box (DS1–DS4), in Automation/signage-image/ — full guide in its README.md.

  • Build / write a stick (on a Linux host with the toolchain — docker-server): make iso then make usb.
  • Update the OptiSigns player — bump OPTISIGNS_VERSION and OPTISIGNS_SHA256 in Automation/signage-image/Makefile (one reviewed commit; the download is version-pinned and verified fail-closed, so a new OptiSigns release can’t silently change the image).
  • Change the appliance config (autologin, node_exporter, watchdog, update-alignment) — edit Automation/signage-image/provision.sh. It is idempotent, so sudo ./provision.sh also brings an existing box up to standard.

The signage fleet (DS1–DS4)

Helper scripts under Automation/server/scripts/ run from docker-server and push config to the playback PCs (they use sshpass + sudo, prompting once for the spacadmin password):

  • align-signage-updates.sh — keeps unattended-upgrades aligned with the overnight reboot (upgrade 03:00, reboot 03:30) so a box never runs all day on half-applied updates.
  • silence-signage-update-nag.sh — stops Ubuntu’s “Software Updater” mapping over the player, and pins the other apt timer. align-signage-updates.sh moved apt-daily-upgrade.timer to 03:00 but left apt-daily.timer — the package-list refresh — at its stock 12-hour random delay, so discovery still fired at random times of day and update-notifier popped a GUI prompt onto the wall when it found something. Unlike the others this also runs from the tech workstation over the DS1–DS4 ssh aliases, not just docker-server.
  • setup-signage-watchdog.sh — arms the hardware watchdog so a hung boot auto-recovers instead of needing an on-site power cycle.
  • setup-signage-watchdog-metrics.sh — exports watchdog state to the node_exporter textfile collector so Grafana can alert when the watchdog is not armed. Added because that failure was silent: RuntimeWatchdogUSec reads 30s even when no device exists.
  • setup-signage-kiosk-metrics.sh — exports kiosk state to the node_exporter textfile collector, so Grafana alerts when the OptiSigns player is not running (node_optisigns_running), when the update-nag suppression regresses (node_update_notifier_suppressed), or when a prompt is on the wall right now (node_update_manager_running). Added for the same reason as the watchdog metrics: each of these fails silently, and the detector would otherwise be someone noticing a blank wall or a window on it. The player check is the one with teeth — a box can ping, report SMART and keep its watchdog armed while showing nothing at all.
  • check-signage-apt-sources.sh — audits /etc/apt/sources.list.d/ubuntu.sources for the -updates pocket (FIX=1 repairs it). DS1 was missing it entirely and sat seven months behind on point releases while reporting 0 upgradable; the unattended-upgrades Allowed-Origins fix could not help, because that setting filters what apt already sees rather than adding a repository.
  • setup-signage-screenshot.sh — installs spac-screenshot, the flash-free way to see what a sign is showing, on a box not built from the image. It runs the image’s own installer (signage-image/collectors/install-screenshot.sh), then takes one capture and checks it came through GNOME Shell, the faithful route. Never use gnome-screenshot on a live sign; it flashes the wall.

Add a new box to monitoring: install the host agent — sudo apt install prometheus-node-exporter smartmontools (bare metal, :9100; docker-server itself is set up the same way) — then add it to the signage job in prometheus.yml and merge. The Grafana signage alerts (root-fs read-only, host down, SMART failing) then cover it automatically. The watchdog and update-nag alerts need setup-signage-watchdog-metrics.sh and setup-signage-kiosk-metrics.sh run against the box as well — they key off textfile-collector metrics, not anything node_exporter exposes on its own. Neither is part of provision.sh, so a freshly imaged box is monitored for disk and reachability but silently unmonitored for the watchdog, the player and the update nag until someone runs the two scripts — which is exactly how a box provisioned on 2026-09-15 sat on a blank wall with a dead player and every signage alert green.

Switch & router config backups

Pull the persisted config from every managed switch and the EdgeRouter into Networking/ so git diff shows real changes. Scripts under Automation/server/scripts/ (run from docker-server):

  • make switch-backups (from Automation/server/) — everything: the NETGEAR M4250s, the Cisco SG300/SG500, and the EdgeRouter.
  • Individually: fetch-switch-configs.sh (M4250s), backup-cisco-configs.sh (Cisco), backup-router-config.sh (EdgeRouter).

Commit the resulting changes under Networking/ to keep the saved configs current.

A backup saves the switch first. The M4250 path runs write memory before it pulls startup-config, so whatever is running is persisted and captured together, and it fails outright if the switch does not confirm the save. It used to send copy running-config startup-config, which this firmware rejects as invalid input; the failure was swallowed, so every backup silently captured the last-saved config instead. Anything changed on a switch but never saved was neither persisted nor backed up.

That makes running a backup a save, so only run it against a switch whose running config is what should survive a reboot. Check first, without saving anything, by comparing show running-config against show startup-config — an unsaved experiment on any one switch will otherwise be made permanent by the next make switch-backups.

PoE on/off is left out of the backups. Companion’s system on/off buttons turn PoE on and off on AV-CORE 0/19–0/28 and a dozen Control Room ports, and they do it through the port’s PoE config. That’s the same setting as the no poe line: the switch’s PoE API has no runtime-only off. So fetch-switch-configs.sh drops every no poe line; otherwise each backup would commit whatever state the system was left in. The save still happens, so a switch backed up while the system is off boots with those ports unpowered until the system is switched on. A port whose PoE was turned off on purpose isn’t recorded in git either. Other PoE settings (poe power limit none) are.

Where the backups live

Every destructive change this tooling makes is preceded by a backup on docker-server under ~/volume-backups/ (Docker volume tarballs, retired stack definitions, extracted secrets). Restore recipes are in the deploy README.

Some of this tooling is still landing via open pull requests — if a path above doesn’t exist on Live yet, check the open PRs.