Tooling & maintenance
Tooling & maintenance
How to update the automation behind the AV systems. Most of it lives in the
repository under
Automation/, and runs on docker-server, which sits on all the AV VLANs and
is the operations hub (monitoring, config backups, fleet management).
Each tool has a detailed README next to its code; this page is the map and the
“how do I change it” workflow. The golden rule: change things in git, not on the
box.
The monitoring stack (Grafana, Prometheus, Companion, Uptime-Kuma)
docker-server runs these as a Docker Compose project called services, and git
is the single source of truth. GitHub notifies docker-server the moment Live
changes — a webhook, published through Tailscale Funnel — and the box reconciles
the running stack to match the repo within seconds. A systemd timer
(spactech-deploy.timer) also checks Live every ~2 minutes, so a missed
notification only costs time. There is no Portainer and nothing to edit on the
box — a merged commit is the deploy.
Everything for it is in Automation/server/; the mechanics and the data-safe
runbook are in
Automation/server/deploy/README.md.
To change it: edit under Automation/server/, open a PR, merge to Live. It
deploys itself within ~2 minutes. Common changes:
- Add a device to monitor — add a target to
Automation/server/prometheus.yml(copy an existingjob_nameblock). node_exporter must be running on the target (see Signage fleet below for the bare-metal install). The deploy validates the new file and recreates the Prometheus container to load it — a SIGHUP would re-read a stale copy, becauseprometheus.ymlis a single-file bind mount. - Change a service version — bump the
image:tag inAutomation/server/docker-compose.yml. - Dashboards & alerts — these live in Grafana’s own volume and are edited in
the Grafana UI; the repo keeps exported copies under
Automation/server/backups/grafana/. Re-export after UI changes so the repo stays current.
Prometheus keeps 180 days of full-resolution data, with no size cap (the
--storage.tsdb.retention.time flag in the compose file). At roughly 1.3 GB per 15
days that grows to about 16 GB, so it is worth watching docker-server’s free space
until it levels off. Prometheus drops data older than the limit once it gets there.
Deploy now instead of waiting (on docker-server):
~/SPACTech/Automation/server/deploy/deploy.sh
Roll back: revert the commit — it deploys like any other merge. Data volumes are
external, so container recreation never touches Grafana/Prometheus/Kuma/Companion
data. Volume backups and a restore recipe are in the deploy README.
Host-level setup on docker-server
The services stack deploys itself, but docker-server’s own host packages and
network config do not — they are applied by script and captured here:
setup-docker-server-comms-lldp.sh— two fixes that both need root on the box. It gives thecommsVLAN interface an address (the EdgeRouter already reserves10.0.24.200for it; the interface just never asked over DHCPv4), and installslldpdscoped to theenp3s0f0trunk so the box identifies itself in the switches’ LLDP neighbour tables instead of appearing as an anonymous trunk port. Run it on docker-server withsudo;DRY_RUN=1shows what it would do. It writes a netplan drop-in rather than editing00-installer-config.yaml, so it is reversible by deleting one file.setup-docker-server-rx-buffers.sh— receive buffers sized for SVSi streams, both needing root. It raisesnet.core.rmem_maxto 8 MB through/etc/sysctl.d, so a socket that asks for a large receive buffer gets one, and grows theenp3s0f0RX ring from 200 to 511 descriptors with a oneshot unit (spactech-rx-ring.service) that runs whenever the NIC appears. Without them, anything receiving a 1080p60 stream here lost a burst of packets every few seconds; see Receiving a stream on docker-server. Run it withsudo;DRY_RUN=1shows what it would do. Resizing the ring resets the NIC (traffic pauses for under 0.1 s), so it skips that when the ring is already 511.deploy/install.sh— installs the deploy’s systemd units (the timer, the GitHub webhook and the path unit it drives) and publishes the webhook through Tailscale Funnel, then sends it a signed push to prove the chain deploys. Re-run it withsudowhenever a unit file inAutomation/server/deploy/changes — a merge updates the checkout, never/etc/systemd/system.webhook.pyreloads on its own.
The signage appliance image
Reproducible Ubuntu image for a new/replacement digital-signage box (DS1–DS4),
in Automation/signage-image/ — full guide in its
README.md.
- Build / write a stick (on a Linux host with the toolchain — docker-server):
make isothenmake usb. - Update the OptiSigns player — bump
OPTISIGNS_VERSIONandOPTISIGNS_SHA256inAutomation/signage-image/Makefile(one reviewed commit; the download is version-pinned and verified fail-closed, so a new OptiSigns release can’t silently change the image). - Change the appliance config (autologin, node_exporter, watchdog,
update-alignment) — edit
Automation/signage-image/provision.sh. It is idempotent, sosudo ./provision.shalso brings an existing box up to standard.
The signage fleet (DS1–DS4)
Helper scripts under Automation/server/scripts/ run from docker-server and
push config to the playback PCs (they use sshpass + sudo, prompting once for
the spacadmin password):
align-signage-updates.sh— keeps unattended-upgrades aligned with the overnight reboot (upgrade 03:00, reboot 03:30) so a box never runs all day on half-applied updates.silence-signage-update-nag.sh— stops Ubuntu’s “Software Updater” mapping over the player, and pins the other apt timer.align-signage-updates.shmovedapt-daily-upgrade.timerto 03:00 but leftapt-daily.timer— the package-list refresh — at its stock 12-hour random delay, so discovery still fired at random times of day andupdate-notifierpopped a GUI prompt onto the wall when it found something. Unlike the others this also runs from the tech workstation over theDS1–DS4ssh aliases, not just docker-server.setup-signage-watchdog.sh— arms the hardware watchdog so a hung boot auto-recovers instead of needing an on-site power cycle.setup-signage-watchdog-metrics.sh— exports watchdog state to the node_exporter textfile collector so Grafana can alert when the watchdog is not armed. Added because that failure was silent:RuntimeWatchdogUSecreads30seven when no device exists.setup-signage-kiosk-metrics.sh— exports kiosk state to the node_exporter textfile collector, so Grafana alerts when the OptiSigns player is not running (node_optisigns_running), when the update-nag suppression regresses (node_update_notifier_suppressed), or when a prompt is on the wall right now (node_update_manager_running). Added for the same reason as the watchdog metrics: each of these fails silently, and the detector would otherwise be someone noticing a blank wall or a window on it. The player check is the one with teeth — a box can ping, report SMART and keep its watchdog armed while showing nothing at all.check-signage-apt-sources.sh— audits/etc/apt/sources.list.d/ubuntu.sourcesfor the-updatespocket (FIX=1repairs it). DS1 was missing it entirely and sat seven months behind on point releases while reporting0 upgradable; the unattended-upgradesAllowed-Originsfix could not help, because that setting filters what apt already sees rather than adding a repository.setup-signage-screenshot.sh— installsspac-screenshot, the flash-free way to see what a sign is showing, on a box not built from the image. It runs the image’s own installer (signage-image/collectors/install-screenshot.sh), then takes one capture and checks it came through GNOME Shell, the faithful route. Never usegnome-screenshoton a live sign; it flashes the wall.
Add a new box to monitoring: install the host agent —
sudo apt install prometheus-node-exporter smartmontools (bare metal, :9100;
docker-server itself is set up the same way) — then add it to the signage job in
prometheus.yml and merge. The Grafana signage alerts (root-fs read-only, host
down, SMART failing) then cover it automatically. The watchdog and update-nag
alerts need setup-signage-watchdog-metrics.sh and setup-signage-kiosk-metrics.sh
run against the box as well — they key off textfile-collector metrics, not
anything node_exporter exposes on its own. Neither is part of provision.sh,
so a freshly imaged box is monitored for disk and reachability but silently
unmonitored for the watchdog, the player and the update nag until someone runs the
two scripts — which is exactly how a box provisioned on 2026-09-15 sat on a blank
wall with a dead player and every signage alert green.
Switch & router config backups
Pull the persisted config from every managed switch and the EdgeRouter into
Networking/ so git diff shows real changes. Scripts under
Automation/server/scripts/ (run from docker-server):
make switch-backups(fromAutomation/server/) — everything: the NETGEAR M4250s, the Cisco SG300/SG500, and the EdgeRouter.- Individually:
fetch-switch-configs.sh(M4250s),backup-cisco-configs.sh(Cisco),backup-router-config.sh(EdgeRouter).
Commit the resulting changes under Networking/ to keep the saved configs current.
A backup saves the switch first. The M4250 path runs write memory before it pulls
startup-config, so whatever is running is persisted and captured together, and it fails outright
if the switch does not confirm the save. It used to send copy running-config startup-config, which
this firmware rejects as invalid input; the failure was swallowed, so every backup silently captured
the last-saved config instead. Anything changed on a switch but never saved was neither persisted
nor backed up.
That makes running a backup a save, so only run it against a switch whose running config is
what should survive a reboot. Check first, without saving anything, by comparing
show running-config against show startup-config — an unsaved experiment on any one switch will
otherwise be made permanent by the next make switch-backups.
PoE on/off is left out of the backups. Companion’s system on/off buttons turn PoE on and off on
AV-CORE 0/19–0/28 and a dozen Control Room ports, and they do it through the port’s PoE
config. That’s the same setting as the no poe line: the switch’s PoE API has no runtime-only
off. So fetch-switch-configs.sh drops every no poe line; otherwise each backup would commit
whatever state the system was left in. The save still happens, so a switch backed up while the
system is off boots with those ports unpowered until the system is switched on. A port whose PoE
was turned off on purpose isn’t recorded in git either. Other PoE settings (poe power limit none)
are.
Where the backups live
Every destructive change this tooling makes is preceded by a backup on
docker-server under ~/volume-backups/ (Docker volume tarballs, retired stack
definitions, extracted secrets). Restore recipes are in the deploy README.
Some of this tooling is still landing via open pull requests — if a path above doesn’t exist on
Liveyet, check the open PRs.