Digital Signage

Digital Signage

The signage system is three layers: content PCs feed SVSi encoders, which are routed to SVSi decoders behind each display. Routing is driven by AMX — see Automation/amx/include/SVSi.axi.

Content PCs

These are the playback machines that source the signage encoders. They live on the house/office network, not on an AV VLAN.

Node IP MAC Feeds encoder Stream Location
DS1 192.168.0.115 00:23:24:f6:21:bb DigitalSignage1 (10.0.25.80) 11 Server room, AVoIP DS 0/23
DS2 192.168.1.157 00:23:24:f2:d8:76 DigitalSignage2 (10.0.25.81) 12 Server room, AVoIP DS 0/21
DS3 192.168.0.147 d8:cb:8a:88:0c:35 DigitalSignage3 (10.0.25.82) 13 Server room, AVoIP DS 0/22
DS4 192.168.0.103 d8:cb:8a:87:b2:5a DigitalSignage4 (10.0.25.83) 14 Server room, AVoIP DS 0/24

The PC-to-encoder link is HDMI, not network — the PCs sit on the office LAN while the encoders are on VLAN 25. DigitalSignage5 (10.0.25.84, stream 15) has no PC behind it. The whole rack — switch ports, encoders and power — is laid out in the network map.

All four drive their encoder from a DisplayPort output through the same passive DP→HDMI adaptor and slimline HDMI cable. Linux reports that link on HDMI-A-1, not DP-1: with a passive (DP++) adaptor the i915 driver runs the port in HDMI mode and marks its DP twin on the same DDI disconnected. That is expected, not a fault. A healthy box sees the encoder’s EDID on HDMI-A-1 with a valid header and checksum, and runs 1920x1080@60.

DS1/DS2 addresses recorded 2026-09-02; DS3 and DS4 moved to 192.168.0.147 and 192.168.0.103 when they were rebuilt on SSDs 2026-09-14/16. All four confirmed up in Prometheus 2026-09-16.

These are DHCP leases, not reservations, so an address can move if a box is re-racked onto a different port — which is exactly how the old 192.168.0.150 / 192.168.0.100 entries for DS3 and DS4 went stale in monitoring-targets.csv and Uptime Kuma. Both are corrected as of 2026-09-16; the Kuma monitors were re-addressed over socket.io with Automation/server/kuma-edit.js. The 2026-09-17 re-cabling moved all four onto new ports and every one kept its lease — but that is luck of lease timing, not a guarantee.

All four are on one flat 192.168.0.0/23, not two subnets — the office network is a /23 (255.255.254.0), so 192.168.0.0-192.168.1.255 is a single broadcast domain. DS2 at 192.168.1.157 is on the same network as the others despite the different third octet.

The office gateway is 192.168.0.1 — a UniFi Dream Machine Pro (MAC 68:d7:9a:59:8c:74, UniFi OS, SSH and HTTPS open). It is docker-server’s primary default route and the site’s only path to the internet.

It is the same physical box as 10.0.22.254 on the lighting VLAN — identical MAC, identical UniFi OS certificate (CN = unifi.local) — so treat any doc describing those as two devices as wrong. It also answers on 192.168.16.1, an internal UniFi network; 192.168.0.1 is the address worth knowing, because that is the one everything defaults to. The EdgeRouter (AV-ROUTER, 192.168.30.254) is a separate device and carries no default route at all.

The Dream Machine is not ours. It is managed by IT, and nothing in this repo has credentials for it or should acquire any. Everything from the EdgeRouter inward — the M4250s, the Ciscos, docker-server, the AV VLANs — is this team’s. Anything that needs changing on 192.168.0.1 is a request to IT, not a task. Worth knowing before you go looking for a login that does not exist.

Do not infer DHCP behaviour from Networking/Router/config/config.boot — that file was last updated 2024-04-11 and is the oldest network config in this repo by well over a year. The address it describes for the router, 10.0.0.1, does not respond to ping from inside the network at all. It has no shared-network-name block for 192.168.0.0/23 and its LAN1 block is disabled, but given its age that says nothing reliable about the live device. Capture a current export before using it to reason about anything.

Operating system

All four run Ubuntu 24.04.5 LTS (base-files 13ubuntu10.5) with 0 upgradable, on the GA kernel 6.8.0-139-generic (linux-generic). Confirmed across the fleet 2026-09-16. The HWE series that DS1 and DS2 were originally built on is gone — see Migrating an existing box from HWE to GA.

DS3 was rebuilt on 2026-09-16 — new Samsung SSD 870, re-imaged from spac-signage-24.04.3.iso. The failing ST500VT000 and its read-only-root I/O errors are gone.

Phased updates will hold a box back

A point release can lag with nothing wrong. Ubuntu rolls base-files out gradually, and until a box’s turn comes apt-cache policy shows (phased 10%) and apt-get -s full-upgrade reports deferred due to phasing. That is normal and self-resolving — it is not the same fault as a missing -updates source, which never resolves on its own.

If you do need them level immediately, force it with -o APT::Get::Always-Include-Phased-Updates=true — but force it on every box at once. Doing one leaves it ahead of the others, which is the same skew pointing the other way. Check what is actually being held first: inert packages (base-files, language packs, release-upgrade tooling, apt bindings) are safe to pull forward, anything kernel or graphics is not, and skipping the check is how you turn a cosmetic version difference into an outage.

Power-cycling an SVSi endpoint remotely

The signage encoders are not PoE-powered, with one temporary exception. All five sit in the SVSI sleds on AVoIP DS ports 0/13–0/17, and those ports read Searching at 0 mW, so poe reset does nothing to them — except DS3’s (0/15), which is on PoE for now, so its port can cycle it alone. The sleds and the four content PCs are powered through the RackLink on the same switch — RLNK-DS, 192.168.30.23 — which is the remote power path for the rack. The PCs come back by themselves after a power cut (After Power Loss = Power On), so a cycled outlet is a full recovery for them too; it is a hard power cut, though, so prefer a reboot over SSH while the box still answers.

This is the recovery path when an encoder stops seeing its source: the encoder can latch its HDMI input when the source drops during a reboot, and nothing on the PC clears it — not a connector re-probe, not rebooting the PC. Only a power cycle does. The symptom is the PC reporting 0 connected DRM outputs with 0-byte EDID while the encoder reports DVIINPUT:disconnected, usually alongside EDID has corrupt header and an escalating i915_hotplug_work_func count in the kernel log.

Outlet Feeds
1–4 DS1–DS4, one PC each
5 SVSI Cage — all five encoders
6–8 nothing assigned

Outlet 5 is all five encoders at once. There is no way to cycle one encoder on its own from here, and cycling the cage drops every signage stream until the encoders boot (~75 s). A single PC can be cycled on its own outlet. Use the outlet’s Cycle action in the unit’s web UI (http://192.168.30.23, admin); each outlet is set to a 3 s cycle. AMX does not drive this unit — racklink.axi has no .23 — so nothing automated will switch the rack.

A single decoder can always be cycled from the switch: the ones on AVoIP DS (0/1–0/11, 0/19) draw PoE. Find its port in the network map — the switch config carries only generic descriptions — then:

show mac-addr-table        # not `show mac-address-table`, which this firmware rejects
show poe port info 0/5     # confirm it is actually Delivering Power first
configure
interface 0/5
poe reset

Prefer poe reset over disabling and re-enabling PoE — it is a single atomic cycle, so a dropped session cannot leave the port dark. Allow ~75 s for the endpoint to boot and answer getStatus again, then confirm the port returned to Delivering Power rather than Disabled.

Two signals that will mislead you. VDET/HDET are unreliable alone — a healthy DS2 reads VDET:0 HDET:1 while a healthy DS4 reads VDET:1 HDET:0; use DVIINPUT:connected plus a plausible INPUTRES. And the sysfs edid file can read 0 bytes on a working link, so it is only meaningful alongside a connector that also reads disconnected.

Patching

unattended-upgrades tracks the full 24.04 stream on all four. Configured 2026-09-02 via /etc/apt/apt.conf.d/52unattended-upgrades-spac:

  • Allowed-Origins — the stock config allowed only ${distro_id}:${distro_codename} and -security. The drop-in appends -updates, which is the pocket carrying point-release fixes. Without it the boxes took security patches but never advanced 24.04.x, which is why DS2 and DS4 had accumulated 114 and 196 pending packages.
  • Automatic reboot at 03:30, which was previously unset — so reboot-required accumulated and kernel updates stayed dormant.
  • /etc/apt/apt.conf.d/20auto-upgrades was "0" on DS1 only, disabling unattended upgrades outright. That is the reason DS1 drifted behind the others. It was a cause, not the cause — see When a box cannot see the updates pocket below.
  • /etc/update-manager/release-upgrades is Prompt=never on all four, so nothing suggests a jump out of 24.04. unattended-upgrades never performs a release upgrade regardless.

When a box cannot see the updates pocket

A box can report 0 upgradable and still be months behind, if its apt sources omit the pocket the point releases come from. DS1 was in exactly that state — Suites: noble where the other three read Suites: noble noble-updates noble-backports.

Allowed-Origins cannot add a repository. It filters what unattended-upgrades will install out of what apt already indexes, so appending -updates to the drop-in does nothing on a host with no -updates source. Do not assume the unattended-upgrades config is the whole story.

Kernel parity is not evidence that a box is current. HWE kernel SRUs publish to -security as well as -updates. A box missing only -updates still takes every kernel bump and tracks the fleet exactly, while the pocket carrying base-files and the point-release rollup never arrives. That is what made this invisible for seven months.

To check a box properly, compare base-files across the fleet and confirm apt-cache policy lists a -updates source — /var/lib/apt/lists/ should hold files matching <codename>-updates_main (7 on a healthy box here, 0 on DS1). Audit and repair the whole fleet with Automation/server/scripts/check-signage-apt-sources.sh (FIX=1 rewrites the suites line and refreshes).

Suspect the Software & Updates GUI when you find a .save sibling next to ubuntu.sources — that is what it writes when someone unticks “Recommended updates”. Every box also keeps ubuntu.sources.curtin.orig from autoinstall, with the correct suites, which is the ground truth to diff against.

Name a backup so apt ignores it silently, or keep it out of the directory. apt has an explicit allowlist of extensions it skips without comment:

$ apt-config dump | grep Ignore-Files-Silently
Dir::Ignore-Files-Silently:: "\.bak$";     "\.save$";   "\.orig$";
Dir::Ignore-Files-Silently:: "\.disabled$";  "\.distUpgrade$";  "\.dpkg-[a-z]+$";  "\.ucf-[a-z]+$";

Those patterns are anchored. ubuntu.sources.save and ubuntu.sources.curtin.orig match and are silent; ubuntu.sources.bak.20260907140656 — left on DS1 by the 2026-09-07 repair — does not, because the timestamp comes after .bak, so apt flagged it as an unrecognised extension. Moved to /var/backups/spac-signage/ on 2026-09-10. .save and .curtin.orig were left in place; both are silent and both are diagnostically useful.

In practice that file was cosmetic, not a live problem: it produced N: Ignoring file … notices in a manual unattended-upgrade run, but none of apt-get update, apt-get full-upgrade, apt-cache policy or apt list surfaced it, and it never reached the journal from the nightly 03:00 run — 0 hits in the preceding 7 days. Worth fixing as hygiene, not worth chasing as a fault.

Local config lives in a 52- drop-in rather than edits to 50unattended-upgrades, which is a package conffile and can be replaced on upgrade. Keep backups out of /etc/apt/apt.conf.d/ — apt parses every file in that directory and warns on unrecognised extensions.

Pin both apt timers, not just the upgrade

There are two apt timers and aligning only one leaves the other loose:

Timer Does Stock schedule
apt-daily.timer refreshes the package lists *-*-* 6,18:00, RandomizedDelaySec=12h
apt-daily-upgrade.timer installs what u-u allows *-*-* 6:00, RandomizedDelaySec=60m

RandomizedDelaySec=12h on a twice-daily timer is a uniform random draw across the whole day. DS1 ran apt-daily.service 60 times in the 30 days to 2026-09-10, spread near-evenly over all 24 hours — about half during viewing hours. Only the upgrade timer had been pinned to 03:00, so installs were quiet while discovery still fired at random times of day, and discovery is what triggers the Software Updater popup.

Both are now pinned: refresh 02:45 (+10m), upgrade 03:00 (+10m), reboot 03:30. The gap between refresh and upgrade keeps them off each other’s apt lock — apt-daily.service takes 1–10 s on these boxes, so five minutes is ample.

What the loose timer actually cost, measured on 2026-09-10 at 18:30, before the pinning took effect. All three boxes were on the same libc6 (2.39-0ubuntu8.8) and the same sources:

Node Last apt-daily Upgradable
DS1 17:21 14
DS2 08:03 0
DS4 09:02 0

Identical boxes, three different views of reality — purely because of when each one’s random draw landed relative to when the packages were published. DS2 and DS4 were not up to date; they had simply not looked since morning. Worse, a box that discovers updates after its 03:00 window has passed then carries them, visible and uninstalled, for a full day — which is precisely the state that raises the update prompt. From the pinning onward all three refresh inside the same ten minutes and install in the same window.

Never let a GUI update prompt reach the wall

2026-09-10, 17:22 — DS1 raised Ubuntu’s “Software Updater” over the OptiSigns kiosk, and it sat on the nine atrium/Legato/balcony displays for ~40 minutes. Nothing was broken. This is stock Ubuntu desktop behaviour that had never been turned off:

17:21:50  apt-daily.service runs, refreshes the package lists
17:21:57  finds 14 new packages, writes /var/lib/update-notifier/updates-available
17:22:31  update-notifier (autostarted in the session) sees that file change
17:22:33  launches `update-manager --no-update --no-focus-on-map`

--no-focus-on-map does not mean invisible. It means the window does not steal focus — it still maps, on top of the fullscreen player. Do not read that flag as “harmless”.

It looked sporadic rather than daily because update-notifier throttles itself to one auto-launch a week (com.ubuntu.update-notifier regular-auto-launch-interval = 7). DS1’s launch-count was already 9 before anyone connected it to what was on the wall.

This was not caused by restoring DS1’s -updates source. That is the obvious suspect, and it is innocent: 13 of the 14 packages that triggered it are in noble-updates and noble-security, and DS1 always had -security. The popup predates the repair.

Fixed on DS1, DS2 and DS4 by Automation/server/scripts/silence-signage-update-nag.sh, and in the image build:

  • ~/.config/autostart/update-notifier.desktop with Hidden=true. This is the XDG autostart spec’s own mechanism — a same-named entry in the user’s autostart directory overrides /etc/xdg/autostart. Deliberately not an edit to /etc/xdg/autostart/update-notifier.desktop, which is a package conffile and would return on upgrade.
  • no-show-notifications=true, hide-reboot-notification=true, show-apport-crashes=false as belt and braces if something starts it anyway. The apport one is the same class of bug: a crash dialog over the signage.
  • Both apt timers pinned overnight (above), so discovery stops happening during viewing hours.

Do not purge update-notifier. It is a dependency of ubuntu-desktop-minimal, update-manager and apport-gtk, so removing it drags the desktop metapackage out with it. Suppressing the autostart gets the same result with nothing else moving.

The per-user half is tied to spacadmin. If a box is ever set to autologin as a different account, redo it for that account — the box will otherwise quietly go back to nagging. The collector below reads the account out of the autologin config rather than assuming spacadmin, so that case shows up as a broken guard rather than passing silently.

Monitoring that it stays suppressed

The suppression fails silently, so like the watchdog it is monitored rather than trusted. Installed by Automation/server/scripts/setup-signage-kiosk-metrics.sh: /usr/local/bin/spac-kiosk-metrics.sh plus spac-kiosk-metrics.timer, writing kiosk.prom into the same node_exporter textfile collector that carries smartmon.prom and watchdog.prom.

node_update_notifier_suppressed is 1 only when all of these hold, since any one alone can lie:

  1. update-notifier is not running
  2. ~/.config/autostart/update-notifier.desktop has Hidden=true
  3. no-show-notifications and hide-reboot-notification are true, show-apport-crashes is false

Condition 1 is the one that matters — a process that is already running does not care what the autostart file says, which is exactly the state a box sits in between the config being fixed and the session being restarted.

The components are exported alongside it so an alert says which guard slipped: node_update_notifier_running, node_update_notifier_autostart_hidden, node_update_notifier_gsettings_ok, and node_kiosk_info{kiosk_user="..."}.

node_update_manager_running is separate and answers a different question: is a prompt on the wall right now.

Rule Expression For Severity
Signage update prompt on screen max by (host) (node_update_manager_running{job="signage"}) > 0 0s critical
Signage update nag suppression broken min by (host) (node_update_notifier_suppressed{job="signage"}) < 1 15m warning

for: 0s on the first is deliberate — it is already visible to the congregation, so there is nothing to wait out. The second waits 15m because a box legitimately reads unsuppressed between the config being applied and the session restarting.

The cadence is 5 minutes, not the watchdog’s 15. A watchdog that failed to arm is a latent risk; a window on the wall is live damage, and the 2026-09-10 one lasted 40 minutes. A 15-minute poll would have been most of the incident.

pgrep -x update-manager will never match it. update-manager runs as /usr/bin/python3 /usr/bin/update-manager …, so its comm is python3. Any check — or any cleanup pkill — has to use pgrep -f '/usr/bin/update-manager'. update-notifier is a real binary and does match -x. Both verified on DS1, 2026-09-10.

Clearing a prompt that is already on the wall

The alert carries this, but it is worth stating where it can be found without an alert firing:

ssh DS1 "pkill -f '/usr/bin/update-[m]anager'"

The [m] is not a typo. Over SSH the pattern ends up in the remote shell’s own command line, so a plain pkill -f '/usr/bin/update-manager' matches itself, kills the session, and returns exit 255 — it looks like it failed when it actually worked. Wrapping one character in a character class makes the regex stop matching the literal text of the command that carries it. With the brackets the same command exits 0 and prints whatever you put after it. Both forms were run against a decoy on DS1 on 2026-09-10; the naive one returned 255, the bracketed one returned 0.

This only clears the window. It comes back unless the suppression is repaired.

“Signage watchdog metrics stale” was widened to cover kiosk.prom as well and renamed “Signage textfile metrics stale”. Note its grouping is by (host, file), not by (host) — with two files matched, max by (host) would let a freshly-written kiosk.prom mask a stale watchdog.prom.

Verified end-to-end on 2026-09-10, not just wired up. A decoy process wearing update-manager’s argv was started on DS1 — nothing drawn on screen, it exists only for pgrep -f to match:

   
18:26:45 collector writes node_update_manager_running 1
18:27:37 rule fires, host=DS1 only — DS2 and DS4 stayed Normal
18:28:50 decoy killed, collector writes 0
18:29:41 rule resolves

59 s from metric to firing, against a budget of ≤60 s scrape + ≤60 s evaluation + 30 s group_wait. Grafana logged the dispatch and zero errors or warnings, so the Slack webhook POST succeeded. The per-host grouping is the part worth having proven: one box faulting does not alert for the other two.

Hardware watchdog

A hung boot or kernel lockup resets the box after ~30 s instead of sitting dead until someone power-cycles it on site. Configured by Automation/server/scripts/setup-signage-watchdog.sh and by the image build.

Loading iTCO_wdt is the hard part, because the stock kernel package ships a blanket watchdog blacklist:

/lib/modprobe.d/blacklist_linux_6.8.0-139-generic.conf:blacklist iTCO_wdt

The filename tracks the running kernel package — it was blacklist_linux-hwe-7.0_7.0.0-31-generic.conf while the fleet was on HWE. Match it with a glob (/lib/modprobe.d/blacklist_linux*), not a literal path.

Three routes exist and only one works. Both of the obvious ones apply that blacklist:

Route Result
/etc/modules-load.d/ ❌ systemd-modules-load checks the blacklist and refuses — Module 'iTCO_wdt' is deny-listed (by kmod)
/etc/initramfs-tools/modules ❌ load_modules() runs modprobe -qb, and -b/--use-blacklist applies the blacklist even to an explicitly named module
Early systemd unit running a plain modprobe ✅

The initramfs route looks correct and is not — verified by reboot with iTCO_wdt present in the initramfs and listed in its conf/modules, dependency included, still not loading. A plain modprobe iTCO_wdt ignores the blacklist; that is the whole difference, and it is what spac-watchdog-module.service runs:

[Unit]
Description=Load the iTCO hardware watchdog module
DefaultDependencies=no
Before=sysinit.target shutdown.target
Conflicts=shutdown.target
ConditionPathExists=!/dev/watchdog

[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/sbin/modprobe iTCO_wdt

[Install]
WantedBy=sysinit.target

Before=sysinit.target matters: it arms the watchdog while the boot logo is still on screen, which is as far as DS3 got during its 19-hour hang. A watchdog that arms only once the system is fully up cannot recover the case it was bought for. The modules-load.d drop-in is left in place as well — inert today, but it starts working on its own if a future kernel package drops the blacklist.

systemctl show -p RuntimeWatchdogUSec is not evidence of anything. It reports 30s whether or not a device exists — it only says what systemd would do with a watchdog. Check the device instead, and verify after a reboot, because an armed watchdog in a running session says nothing about the next boot:

ls /dev/watchdog                                # must exist
cat /sys/class/watchdog/watchdog0/state         # active
journalctl -b -k | grep iTCO_wdt                # must appear seconds into boot, not hours
ls -l /proc/1/fd | grep watchdog                # systemd must hold it open

Monitoring that it stays armed

The arming failure above is silent, so it is monitored rather than trusted.

node_exporter has no watchdog collector — verified against the 1.7.0 binary, not assumed — so this goes through the textfile collector already in use for smartmon.prom. Installed by Automation/server/scripts/setup-signage-watchdog-metrics.sh: a script at /usr/local/bin/spac-watchdog-metrics.sh plus spac-watchdog-metrics.timer (OnBootSec=1min, OnUnitActiveSec=15min — matching smartmon’s cadence, and running just after boot because the reboot is exactly when arming silently fails).

node_watchdog_armed is 1 only when all three hold, since any one alone can lie:

  1. /dev/watchdog exists — the module actually loaded
  2. /sys/class/watchdog/watchdog0/state is active
  3. PID 1 holds an fd on the device — something is actually petting it

Condition 3 is the important one: a device can exist with nothing petting it.

Two Grafana rules in the existing signage group, both to slack-av-alerts:

Rule Expression For Severity
Signage hardware watchdog not armed min by (host) (node_watchdog_armed{job="signage"}) < 1 5m critical
Signage watchdog metrics stale time() - max by (host) (node_textfile_mtime_seconds{job="signage",file=~".*watchdog.prom"}) > 3600 10m warning

The second exists because node_exporter keeps serving a .prom file after its writer dies — a stopped timer would otherwise report armed 1 forever. Note node_textfile_mtime_seconds labels by full path, hence the regex.

Residual gap: if watchdog.prom is deleted outright, both rules go NoData → OK and say nothing. noDataState: OK is the house setting across the signage group, and “Signage PC unreachable” covers the common cause. Worth knowing rather than fixing.

Firmware and power settings

All four boxes are Lenovo ThinkCentres and expose their BIOS settings through thinklmi at /sys/class/firmware-attributes/thinklmi/attributes/, writable as root with no Admin password set. DS1/DS2 are M710q (Kaby Lake, i5-7500T); DS3/DS4 are M73 Tiny (Haswell, Pentium G3220T, machine type 10AX). The two generations do not expose the same attributes.

Setting Value Why
Wake on LAN Automatic Remote power-on; see the trade-off below
After Power Loss Power On A sign should come back by itself after an outage
Smart Power On Disabled USB wake from S5, not wanted
Wake Up on Alarm Disabled  
Boot Agent Disabled PXE; see Boot time
ICE Performance Modes Better Thermal Performance All four share one rack, so fan noise is irrelevant and thermal margin is free
Intel(R) Smart Connect Technology Disabled M73 only. Periodic ME-driven wake with no Linux support
Keyboardless Operation Enabled M73 only. Otherwise POST beeps twice on a box with no keyboard

A thinklmi write does not take effect until the next boot. It lands in NVRAM and reads back immediately, but the running firmware keeps its old value. Any same-boot “change it, power off, observe” test measures the previous value while appearing to test the new one — this is the single easiest way to draw a wrong conclusion about these boxes. Write, reboot, then test. A correct readback is not evidence the setting is live.

Ordered-list attributes (Primary Boot Sequence, Automatic Boot Sequence, Error Boot Sequence) are worse again: a write can read back correctly and still not survive a power cycle. Enumerations have been reliable.

Wake-on-LAN, and why these boxes will not stay powered off

A magic packet wakes any of them from S5 in ~25 s. Automation/server/scripts/wake-signage.sh sends one.

The cost is that firmware Wake on LAN also decides whether the NIC PHY gets +5VSB standby power, and an armed powered PHY self-triggers a wake entering S5 — so a box that has been told to power off comes back roughly 13–60 s later. Network traffic is not the trigger: it happens with the cable physically removed. It is binary, with no middle ground:

Wake on LAN PHY standby (link lights) magic packet stays off
Automatic on works ❌
Disabled off no ✅

ethtool is not in the circuit. The driver’s default is Wake-on: d — nothing armed — and the box wakes regardless; setting wol g changes nothing. Do not write a systemd unit to set it.

Remote power-on is judged worth more than the nuisance wakes on boxes meant to run continuously. To power one down for real: systemctl poweroff, then pull AC. A full AC cycle is the only thing that makes it stay down with WoL enabled.

CSM

DS1 and DS4 run with CSM disabled, which is preferable — it removes the legacy BBS entries from BootOrder, including an IBA GE Slot PXE entry that otherwise regenerates itself even after efibootmgr -b NNNN -B deletes it.

Two boxes cannot: DS2 boots in legacy mode (converting needs a rebuild, and is why efibootmgr returns nothing on it), and DS3 hangs on every boot with CSM off since its BIOS was flashed to FHKT87AUS — Linux starts, then stalls in early sysinit before networking. A fresh re-image hangs identically, so it is not a stale installation, and Lenovo’s own OS Optimized Defaults = Enabled does the same. FHKT63AUS does not have this problem. The mechanism is unidentified.

Disarm the watchdog before changing CSM on any box. With the 30 s iTCO watchdog armed a stalled boot resets and loops invisibly; disarmed, it stops on screen with the failing message visible. systemctl mask spac-watchdog-module.service fails — the unit lives in /etc/systemd/system, so the mask symlink would overwrite it. Use disable plus RuntimeWatchdogSec=0.

Boot time

Boot Agent = Disabled alone does not stop the PXE option ROM. DS4 sat at an 88-second firmware phase with it already disabled, because the network was still referenced in three other places: Primary Boot Sequence, Automatic Boot Sequence and Error Boot Sequence. Removing Network 1 from all three took it to 3.6 s. Disabling CSM removes the ROM entirely, which is the durable fix.

CSM is not a cause of slow POST — measured at 5.885 s with it off versus 5.456 s on.

BIOS updates

Do not flash the M73s. FHKT87AUS (Dec 2021) is the last release, LVFS carries nothing (fwupdmgr get-releases returns “No releases found”), and the update is a net negative:

  • CPU microcode is already superseded by the OS — intel-microcode loads 0x28 at boot over the BIOS’s 0x1d, on both boxes regardless of BIOS version.
  • The EC firmware (FHCT36A.BIN) is byte-identical between releases.
  • FHKT87AUS costs the ability to disable CSM.

What remains is BIOS/SMM-level fixes microcode does not cover — real, but low value for a kiosk on a private VLAN. Allow Flashing BIOS to a Previous Version is Yes with no Admin password, so a rollback is possible if wanted.

If a flash is ever needed, Lenovo’s catalog API returns 403 to plain curl but works with browser headers (User-Agent plus a pcsupport.lenovo.com Referer), taking productId as the machine type from /sys/class/dmi/id/product_sku. One machine type returns two BIOS families — a 10AX query returns both the M73 Tiny and the M73 Tower/SFF images — so read the sibling .txt readme before downloading; the filename does not encode the version. Every historical release is hosted, not just the current one. The ISOs are El Torito hard-disk (media type 4) emulation, so dd of the ISO does not boot; extract the payload using the MBR partition table inside the boot image.

Diagnosing an unexpected power-on

There is no wake-reason log on this platform. PM1_STS/GPE0_STS latch the cause but firmware clears them during POST; ACPI _SWS is absent from the DSDT and all SSDTs, so acpi-call-dkms buys nothing; and the SMBIOS event log has no wake-source descriptor. Differential testing — disable one source, reboot, observe — is the only method.

The SMBIOS event log (DMI type 15) does distinguish some failures. Decode /sys/firmware/dmi/entries/15-0/system_event_log/raw_event_log: 16-byte header, then type/length/6-byte BCD timestamp in UTC.

00 02 00 01 00 00 00 00   keyboard not functional (two POST beeps)
00 00 00 01 01 00 00 00   logged on DS3's CSM-off hang boots
(healthy boots log nothing at all)

Two beeps on a keyboardless box are the keyboard POST error, not a CMOS battery — check the RTC against system time before assuming otherwise.

Hostnames

All four had hostname inconsistencies from cloning, corrected 2026-09-02:

Node Was /etc/hosts 127.0.1.1 was
DS1 DS1 ✅ DS2 ❌
DS2 DS2 ✅ DS1 ❌
DS3 DS3 ✅ DS3 ✅
DS4 ds4 ❌ DS-01 ❌ (name of the source image)

cloud-init ships with preserve_hostname: false and the set_hostname/update_hostname modules active, so hostnamectl set-hostname alone can be reverted on boot. All four now carry /etc/cloud/cloud.cfg.d/99-preserve-hostname.cfg setting preserve_hostname: true.

Despite the DS-01 remnant, these are not bit-for-bit clones. Checked 2026-09-02: /etc/machine-id is unique on all four (and /var/lib/dbus/machine-id is correctly a symlink to it on each), and all four SSH host keys are distinct. The stale entry looks like a leftover manual edit rather than a duplicated identity.

Tailscale (removed)

Tailscale 1.102.3 was found installed on all four with tailscaled enabled at boot and every node logged out — no tunnel, empty /var/lib/tailscale, no reverse dependencies. Nothing in this repo explained why it was there.

Removed 2026-09-02 from all four: daemon disabled, tailscale and tailscale-archive-keyring purged, and /etc/apt/sources.list.d/tailscale.list deleted. That last step matters — purging the packages leaves the third-party apt source behind, and with unattended-upgrades active the hosts would otherwise keep polling Tailscale’s repo indefinitely. Copies of the removed source files are in /var/backups/spac-signage/.

Verified afterwards on each host: no binary, no packages, no systemd units, no tailscale0 interface, no apt source.

This is not a judgement on Tailscale generally — the tailnet is live and legitimately used infrastructure. The docker-server is on it and is the documented remote path to these machines (see Reaching the AV VLANs). Jumping through one hardened host beats four signage players each holding their own tailnet identity, which is why the removal stands even though the tailnet does not.

Network interface notes

  • DS1 and DS2 have onboard WiFi (wlp2s0), currently DOWN. DS3 and DS4 have no wireless interface. This matches the two hardware generations implied by the MAC OUIs.
  • All four route via 192.168.0.1 on enp0s31f6, addressed by DHCP.
  • DS2 carries eight IPv6 ULA addresses (fda0:163e:ec8:4221::/64, SLAAC privacy addresses) on its primary interface; DS1 has no IPv6 at all. Worth understanding why the two differ.

Playback software

All four run OptiSigns, started by a per-user systemd unit under an autologin spacadmin session (graphical.target, GDM, idle-delay=0).

  Path Launcher
All four /opt/optisigns/optisigns.AppImage ~/.config/systemd/user/optisigns.service
DS1 only ~/bin/optisigns-remote-agent ~/.config/systemd/user/optisigns-remote-agent.service

DS1 runs the remote agent in addition to the AppImage, not instead of it. Both units are enabled there. Nothing in this repo records why only DS1 has it; presumably remote management of the zone driving the nine atrium/Legato/balcony displays. Decide whether the other three should have it too, or DS1 should not.

Layout

Playback runs from /opt/optisigns/, owned by spacadmin rather than root so OptiSigns can update its own AppImage in place. Unit files are mode 644; systemd warns on executable unit files.

The systemd user unit is the only launcher — there is deliberately no GNOME autostart .desktop alongside it. Not true. All three healthy boxes carry ~/.config/autostart/OptiSigns Digital Signage.desktop as well, which OptiSigns writes itself when “launch on startup” is enabled in the player. So two things start playback, not one. Only one process is ever running — the AppImage takes an Electron single-instance lock, so whichever loses the race exits — but the arrangement is still worth collapsing to one launcher. Decide which, rather than leaving it to a lock to sort out. Noticed 2026-09-10 while suppressing the update prompt in the same directory.

Backups of anything this tooling has replaced are in /var/backups/spac-signage/ on each host.

Monitoring that the player is running

The player is the entire point of the appliance and, until 2026-09-15, nothing watched it. A box provisioned that day came up with optisigns.service dead and sat on a blank wall while all seven signage alerts stayed green: it pinged, node_exporter answered, SMART was clean, the watchdog was armed, the textfile metrics were fresh.

The hardware watchdog cannot cover this and never could. It fires when systemd stops petting /dev/watchdog — a kernel or systemd lockup. A box whose only job has failed goes on petting it quite happily.

node_optisigns_running rides in the same kiosk.prom collector as the update-nag metrics (Automation/server/scripts/setup-signage-kiosk-metrics.sh, 5-minute timer), with both halves exported alongside it: node_optisigns_unit_active and node_optisigns_process_running.

Rule Expression For Severity
Signage player not running min by (host) (node_optisigns_running{job="signage"}) < 1 15m critical

The headline metric is the OR of its two components, not the AND. Playback has two launchers and they race on boot; the loser exits on the AppImage’s single-instance lock, and the loser can be the systemd unit. That leaves a correct picture on the wall and an inactive optisigns.service — cosmetic, and alerting on the unit alone would page for it.

for: 15m, not 0s, even though a blank wall is live damage. The collector runs every 5 minutes, so anything shorter is theatre; 15m is three consecutive collector runs. It also has to ride out the 03:30 reboot, where the collector’s OnBootSec=1min run can sample the box before the session has finished starting the player. Latching the expression over a window — the trick that stops “Signage disk SMART failing” resolving itself across a reboot — is wrong here for exactly that reason: it would turn one transient 0 into a guaranteed nightly page. A failed disk does not heal; a restarted player does.

Two ways to get the collector wrong, both found while writing it and verified on DS1 on 2026-09-15:

  • systemctl --user does not work from root. The unit belongs to the autologin account, so a root-run collector asking systemctl --user is-active optisigns gets Failed to connect to bus: No medium found — it went looking for root’s own session bus. It has to be the kiosk user’s, and the runtime dir has to be named explicitly: sudo -u spacadmin env XDG_RUNTIME_DIR=/run/user/1000 systemctl --user is-active optisigns. (systemctl --machine=spacadmin@ --user is-active optisigns also returns active on systemd 255, with systemd-container not installed.)
  • pgrep -f optisigns reports a dead player as alive on DS1, because it matches the remote agent — /home/spacadmin/bin/optisigns-remote-agent run — which has nothing to do with playback. The check matches optisigns\.AppImage instead, which catches the player under either launcher and nothing else.

Only active counts as running. A player that cannot start sits in a Restart= loop and reads activating, not failed, for as long as anyone cares to watch it.

Seeing what is on the wall

A running player is not the same as a useful picture — an unpaired player is active and shows only its pairing code. spac-screenshot, on all four boxes, captures the display without flashing it:

ssh DS1 spac-screenshot /tmp/wall.png
scp DS1:/tmp/wall.png .

It asks GNOME Shell for the picture with the flash turned off. gnome-screenshot always flashes — it hard-codes the shell’s flash parameter to true — and the shell only answers callers holding an allowlisted D-Bus name. That check is only on the name, so spac-screenshot takes org.gnome.Screenshot for itself and passes flash=false. What comes back is what the compositor draws, so it is faithful on every box whatever the display’s pixel format, and it needs no privileges.

If the kiosk session is not running, it falls back to reading the scanout buffer (kmsgrab) and says so. That picture is only trustworthy on DS4: DS1 and DS2 scan out render-compressed buffers that come back as noise and fragments of earlier slides, and DS3’s 10-bit desktop cannot be read at all. Never use gnome-screenshot on a live sign; spac-screenshot --allow-flash uses it only for a screen that is out of service. Automation/server/scripts/setup-signage-screenshot.sh installs or updates it, and the details are in the signage image README.

Configuration drift

Audited 2026-09-02. Nothing below is currently faulting; these are inconsistencies worth normalising.

Setting DS1 DS2 DS3 DS4
Disk 233G 234G 233G 112G
sleep-inactive-ac-type suspend nothing nothing nothing
sleep-inactive-ac-timeout 0 3600 3600 3600
  • DS1’s power profile differs. Its idle action is suspend where the others are nothing. It does not currently suspend, because its timeout is 0 (never) — but it is a single setting away from a signage player that sleeps, and the other three are not. Set it to nothing.
  • screensaver lock-enabled is true on all four. With idle-delay=0 the session never idles, so no lock appears in practice; if that ever changes, a locked screen displays instead of signage. Worth disabling on a kiosk.
  • DS4 carries snaps the others do not — vlc, gnome-3-28-1804 and core18 removed 2026-09-02. DS4’s snap list now matches the other three exactly.
  • DS4 autologin casing — AutomaticLoginEnable=True normalised to true 2026-09-02. (The commented user1 lines alongside it are stock GDM template text, present on all four, and were left alone.)

Clean across all four: no failed systemd units, NTP synchronised, timezone America/Edmonton, disk usage under 10%, spacadmin the only human account and only member of sudo, one authorized key each, no user crontabs.

Not verified: ufw state and the effective sshd config both require root to read, so the firewall and SSH hardening posture on these hosts is unaudited.

Access

Login account is spacadmin on all four. SSH aliases DS1-DS4 are configured in ~/.ssh/config on the tech workstation, so ssh DS1 resolves to spacadmin@192.168.0.115 and so on. Key-based auth uses the id_ed25519 key; password auth is still enabled on the hosts as a fallback.

SVSi encoders (signage sources)

Defined at Automation/amx/include/SVSi.axi:30-34, addressed at :159-163. VLAN 25.

The AMX identifier and the label the device answers with are not the same string — the device labels are what you will see in N-Able and the SVSi web UI. Both are listed here; NAME values read live via getStatus.

AMX identifier Device NAME IP Stream Input
DigitalSignage1 DS1 - General Slides 10.0.25.80 11 DS1, 1920x1080
DigitalSignage2 DS2 - Kids Min 10.0.25.81 12 DS2, 1920x1080
DigitalSignage3 DS3 - SCS 10.0.25.82 13 nothing (0x0) — DS3 sees no link; see Open questions
DigitalSignage4 Digital Signage 4 10.0.25.83 14 DS4, 1920x1080
DigitalSignage5 Digital Signage 5 10.0.25.84 15 nothing (0x0)

All five run 24/7 — in the SVSI cage, networked on AVoIP DS 0/13–0/17 — and all five answer getStatus, so a non-answering encoder here is a fault, not a power state.

DigitalSignage5 has nothing plugged in (DVIINPUT:disconnected, INPUTRES:0x0), but it still streams ~370 Mbps of blank video into its port. Nothing subscribes, so it goes no further than the switch.

All five bypass their scaler (SCALERBYPASS:yes), so each streams exactly what its PC sends — 1920×1080 — and every signage decoder reports INPUTRES:1920x1080. With the scaler engaged, an encoder streams at its output mode instead, whatever the input, so a signage decoder reporting 1280x720 means its encoder’s scaler has been switched back on. The setting is the Scaler checkbox under the web UI’s advanced settings; unchecked is bypass. Scaling doesn’t save bandwidth either: DS1 measured 658 Mbps at native 1080p against 879 scaled to 720p, both with MPC off. With MPC on, like every encoder that has been checked, it streams 612 Mbps at native 1080p.

The web UI still takes the factory login (admin / password) — confirmed on DS1’s and DS3’s encoders, and probably true fleet-wide. Anyone on the AVoIP VLAN can reconfigure them with it.

Only encoders 1 and 2 are routed by the AMX logic. set_digital_signage_to_slides() (Automation/amx/include/SVSi.axi:140-153) splits the building into two content zones:

Encoder Source Stream Routed by AMX to
DigitalSignage1 DS1 11 AtriumLegato, AtriumAuditoriumL, AtriumAuditoriumR, AtriumDoorsSouth, LegatoOfficeSide, LegatoAtriumSide, BalconyTV1, BalconyTV2, NursingMothersRoom
DigitalSignage2 DS2 12 KidsMinCheckInDesk, KidsMinElevator, KidsMinHallway
DigitalSignage3 DS3 13 nothing — DS SCS (10.0.25.79) is set on the decoder itself, and is on 11 while DS3 has no link
DigitalSignage4 DS4 14 nothing
DigitalSignage5 (none) 15 nothing — no source connected

Live routing vs. the AMX code

Queried directly on 2026-09-02 via getStatus on TCP 50002, from the docker-server’s avoip interface (see Reaching the AV VLANs):

Decoder IP Live stream AMX expects  
AtriumAuditoriumL 10.0.25.60 11 11 ✅
AtriumAuditoriumR 10.0.25.61 11 11 ✅
AtriumDoorsSouth 10.0.25.62 11 11 ✅
AtriumLegato 10.0.25.63 11 11 ✅
LegatoAtriumSide 10.0.25.64 11 11 ✅
LegatoOfficeSide 10.0.25.65 11 11 ✅
KidsMinCheckInDesk 10.0.25.67 12 12 ✅
KidsMinHallway 10.0.25.68 12 12 ⚠️ display disconnected
KidsMinElevator 10.0.25.66 14 → 12 12 ✅ back in step (re-checked 2026-09-16)

Two things this exposes:

KidsMinElevator has been manually switched to DS4’s stream. Resolved, the way the warning predicted. On 2026-09-02 it was running stream 14, not the 12 that set_digital_signage_to_slides() assigns it — someone had repointed it outside AMX. On 2026-09-16 it reads 12 again, so the slides routine ran and silently undid the change, exactly as flagged. If DS4’s stream is wanted on that display, it has to be folded into SVSi.axi or the decoder taken out of AMX’s control; a hand edit in the web UI will not survive.

KidsMinHallway reports DVISTATUS: disconnected — it is subscribed to stream 12 correctly, but nothing is attached to its HDMI output. Display is off, unplugged, or failed.

DS3 drives the SCS display — when it has a link. The DS SCS decoder (10.0.25.79) is set on the decoder itself, which saves it, and nothing in SVSi.axi references .79, so no AMX routine moves it. It belongs on stream 13, but DS3 currently sees no link to its encoder, so it is on 11 (DS1 - General Slides) until that’s fixed; send set:13 to it on TCP 50002 to put it back. Nothing subscribes to DS4’s stream 14.

DigitalSignage1 is also the “Slides” source on both Legato TV panel buttons (Automation/amx/include/Legato.axi:243, :248).

Stream/slides asymmetries

set_digital_signage_to_stream() (:125-138) is not a clean inverse of the slides function:

  • DigitalSignageLegatoOfficeSide is deliberately excluded from stream — it stays on slides. There is a comment in the source saying so.
  • DigitalSignageKidsMinHallway is in the slides function but not the stream function, with no comment explaining why. Check whether that is intentional or an omission.
  • DigitalSignageLegacyDistribution is in the stream function but not the slides function, and has no IP (see below).

Nothing puts the signage back on slides by itself. set_digital_signage_to_slides() runs only from the video panel’s TV Stream Stop button and its power-off confirmation (Video.axi). If a Sunday ends any other way, the displays stay on Ross 01 (stream 26). Ross 01 is off outside services, so those screens show no picture all week. A decoder in that state reads STREAM:26 and NEEDVSTRM:1.

Companion has its own pair of buttons: page 6, row 3, column 1 sends the displays to Ross 01, and column 2 sends them back to General_Feed/Kids_Feed (11/12). Column 2 leaves out Balcony TV 1 and 2, which column 1 moves, so after a Companion round trip the balcony TVs stay on stream 26. The AMX slides function includes them.

SVSi decoders (displays)

Defined at Automation/amx/include/SVSi.axi:59-68, addressed at :189-197. VLAN 25.

Decoder IP
DigitalSignageAtriumAuditoriumL 10.0.25.60
DigitalSignageAtriumAuditoriumR 10.0.25.61
DigitalSignageAtriumDoorsSouth 10.0.25.62
DigitalSignageAtriumLegato 10.0.25.63
DigitalSignageLegatoAtriumSide 10.0.25.64
DigitalSignageLegatoOfficeSide 10.0.25.65
DigitalSignageKidsMinElevator 10.0.25.66
DigitalSignageKidsMinCheckInDesk 10.0.25.67
DigitalSignageKidsMinHallway 10.0.25.68
DigitalSignageLegacyDistribution unassigned — see below

DigitalSignageLegacyDistribution is declared at :66 and driven by set_digital_signage_to_stream() at :134, but has no svsi_decoder_init() call. Its ipAddress is an empty string, so svsi_send_command() opens a UDP client to no host and the node silently never switches.

Signage displays

The screens themselves. The four below are NEC large-format displays, read over their RS-232 ports with a laptop at the display; the rest haven’t been surveyed yet. No control system is wired to these ports. The SCS display is also on the network, at 10.0.25.154; see Over the LAN.

Location Model Serial Input in use Signal Fed by Switch port
Legato – Atrium Side NEC V421 blank HDMI 1080p60 DS Legato - Atrium Side (10.0.25.64) AVoIP DS 0/11
Legato – Office Side NEC LCD4215 9Z008486NA DVI 1080p60 DS Legato - Office Side (10.0.25.65) AVoIP DS 0/10
Outside the office NEC LCD3215 01010709NA DVI 1366×768 @ 60 not an SVSi decoder —
SCS NEC V321 09000136NA DVI 1080p60 DS SCS (10.0.25.79) AVoIP DS 0/6

“Serial” is what the display reports over RS-232; the V421 reports an empty one. “Signal” is the horizontal/vertical frequency the display measured on that input.

State of each, as left:

  • Clocks are set to Mountain time with the daylight-saving flag off, so they will read an hour fast after DST ends unless reset. Set with nec-display.py clock set.
  • Off timer (022B) is 0 on all four.
  • Schedules: the V421 and V321 have all seven programs disabled through 02E6 (Disable Schedule). That can’t be read back: the display acknowledges any value, including a program number that doesn’t exist, and the schedule read and write commands aren’t implemented. The LCD4215 and LCD3215 have no schedule access over serial at all. Checking or clearing a schedule on any of them needs the on-screen menu, and so a remote — serial remote-key emulation doesn’t work on any of the four.
  • No-signal power save (00E1) is on for the V421 and off for the V321, so the SCS screen stays lit without a source. The LCDs don’t have the setting.
  • All four are unmuted. The V321 is at volume 33.

Controlling them over RS-232

All four run NEC’s large-format display protocol (not the projector one the Legato projector uses) at 9600 8N1, monitor ID 1. Automation/server/scripts/nec-display.py covers status, power, input, volume, mute, backlight, the clock and any setting by its VCP code; nec-survey.py identifies a display and, for a model it hasn’t seen, scans every setting it supports into nec-display-scans/. Both default to the Mac’s USB-serial adapter; pass --port for another.

  V421 V321 LCD4215 LCD3215
Inputs that switch VGA, RGB/HV, DVI, Video 1/2, S-Video, Component, HDMI same as V421 same, less Video 2 and HDMI same as LCD4215
Answers settings queries while off no — power, model, serial and self-diagnosis only no yes, everything yes, everything
Unimplemented C2 command echoed back as if accepted echoed back refused (BE) refused (BE)
Remote-key emulation (C210) echoed, ignored echoed, ignored acknowledged, ignored acknowledged, ignored
Schedule over serial disable only (02E6), unverifiable same as V421 none none
No-signal power save setting 00E1 00E1 none none
VCP codes supported 87 98 69 69

An acknowledgement proves nothing on these displays. Input codes a model doesn’t have (04 on all four; 06 and 09–0B on the LCDs) and out-of-range values come back as success and change nothing. nec-display.py reads every change back and reports one that didn’t take. On the V421 and V321, an echo of a C2 command means “not implemented”, and the tool treats it as an error.

Other things that will catch you out:

  • The V421 drops into power save within 0.5 s of switching to an input with no signal, and a few seconds later stops answering everything except power commands. power on alone doesn’t wake it; switch back to an input with a source.
  • The V321 reports power save (state 2) while showing a picture, so its power state can’t tell you whether the screen is lit. Whether it answers a settings query is the useful test.
  • The LCD4215’s power command reply is malformed — no end-of-message byte or check code, and a wrong length field. The tool accepts it.
  • The first command after plugging the cable into an LCD4215 got a single stray byte back. Retry once.
  • The measured-signal command keeps reporting the last input’s timing after an input change on the LCDs and the V321, so it can’t show whether the new input has a signal.
  • Power off takes ~2 s to reach standby, during which even the power query can go unanswered. Power on takes 3–7 s.

Over the LAN

The SCS display’s LAN port is cabled to the second jack of the DS SCS decoder (.79), which bridges it onto VLAN 25 at AVoIP DS 0/6. It takes DHCP and has a router reservation, so it’s at 10.0.25.154 (MAC 00:25:5c:48:f0:0e). The other three displays’ LAN ports aren’t connected.

The whole serial protocol works over TCP 7142, message for message. A full nec-survey.py scan over the LAN found the same 98 VCP codes as over RS-232, the same identity, clock and diagnostic replies, and the same echo for C2 commands the V321 doesn’t implement. Settings writes work and read back. Both tools take the LAN as a pyserial URL, run from docker-server:

nec-display.py --port socket://10.0.25.154:7142 status
nec-survey.py --port socket://10.0.25.154:7142

Control is over RS-232 or the LAN, not both. The display’s menu picks one, and it’s set to LAN, so a laptop on its RS-232 port gets no answer until the menu is switched back. VCP 103E read 1 in the RS-232 scan and 2 over the LAN, so it’s probably this setting — unconfirmed, and not worth testing remotely: writing it over the LAN would cut the LAN off.

It stays on the network in standby, so it can be powered on remotely. Powered off over the LAN and watched for 8 minutes, it kept answering ping, port 80 and TCP 7142; the power query read standby, and settings reads got no reply at all, as over RS-232. nec-display.py power on over the LAN brought it back on DVI, answering settings reads, within 9 s.

Uptime Kuma watches it. The display_check service in the docker-server stack (Automation/server/scripts/display-check.py, with its display list in display-check.json) queries it every 60 s and pushes to the SCS Display push monitor (Signage group): up when it’s on and on DVI, down with the reason otherwise (“off (standby)”, “on, but input is hdmi (want dvi)”, no connection). The same service watches the three Sharp TVs; see the network map. “On” means it answers a settings read, not power state 1: the V321 has reported power save (2) while showing a picture.

If the checker stops, the pushes stop and Kuma marks the monitor down anyway. While a push monitor is down, Kuma also logs its own “No heartbeat in the time window” beat every couple of minutes between the checker’s pushes. That’s Kuma 1.23’s push-monitor behaviour, not a missed push; the checker’s reason is in the beats around it.

It accepts one connection at a time: a new connection straight after the last one closed is refused, so hold one open or leave a few seconds between them. The Kuma checker holds it for about a second a minute and retries a refused connection, but a tool run can still collide with it.

Port 80 serves a single “MONITOR NETWORK SETTINGS” page (firmware 1.00) that sets DHCP, the address, mask, gateway and DNS, and nothing else. Its form is a GET to jump.html (dp=1 for DHCP), which leads to reset.html?flg=flg and a restart of the network interface. The web server can hang after a connection is left half-open: it refused connections for over 10 minutes before recovering on its own, while TCP 7142 kept working.

It’s silent until spoken to: set static, it sent nothing in a 3-minute capture. Out of the box it’s on NEC’s default 192.168.0.10, found with an ARP sweep sent to its MAC from an address in each candidate subnet.

Reaching the AV VLANs

The docker-server is the single entry point for signage work. It sits on the office 192.168.0.0/23 (ens9, 192.168.0.200) and is trunked onto the AV VLANs as sub-interfaces of enp3s0f0 — video (10.0.23.200), lighting (10.0.22.200), avoip (10.0.25.200), danteprimary (192.168.10.200), control (192.168.30.200) and SPACmgmt (192.168.100.200). So it reaches both the signage PCs and the SVSi endpoints, neither of which is routable from an ordinary office machine.

comms is the exception, and the fault is host-side, not on the switches. The sub-interface exists, is UP and is correctly tagged VLAN 24, but carries no IPv4 address, so there is no connected route for 10.0.24.0/24 and traffic falls out the default route via 192.168.0.1, which cannot reach it.

The VLAN is being delivered. docker-server hangs off M21 port 0/40 — a 1 Gb copper trunk that spells out vlan participation include 2,5,21-26 and vlan tagging 2,5,21-26, so 24 is on it — and M21 reaches the core over 0/48 ↔ CORE 0/44, where 24 is tagged as well. The proof it is actually flowing rather than merely configured: every switch has learned docker-server’s MAC 78:7b:8a:bf:f3:ec on VLAN 24 in its forwarding database, which can only happen if it is receiving VLAN-24-tagged frames from it.

The address is already reserved for it. The EdgeRouter runs DHCP for the subnet — shared-network-name Comms, subnet 10.0.24.0/24, default-router 10.0.24.254, confirmed live on the router rather than read out of the stale config.boot — and it holds a reservation:

static-mapping COMMS-DOCKER-SERVER {
    ip-address 10.0.24.200
    mac-address 78:7b:8a:bf:f3:ec
}

The host was not asking, and now it does. comms carried only a DHCP6 Client DUID where avoip and control each carry a DHCP4 Client ID and collect their own .200 reservations, so it ran DHCPv6 — which EdgeOS does not serve — and never sent a DHCPv4 DISCOVER. Fixed 2026-09-17 by Automation/server/scripts/setup-docker-server-comms-lldp.sh, which drops a netplan file in beside the installer config rather than editing it. It takes the lease but refuses its routes — the lease carries default-router 10.0.24.254, and at the time this box already had three default routes, one of them (10.0.25.249, via avoip) pointing at an address that answered neither ping nor ARP. A fourth was the same trap. It also sets optional: true, clearing the Required For Online: yes that made an interface which can never come online a latent systemd-networkd-wait-online delay.

That third route is gone — see the third default route below.

The other half: a one-sided trunk on AVoIP-DS 0/28

Fixing the host was not enough, and the reason is worth keeping. With the host asking, the EdgeRouter answered — captured on its eth0.24, offering exactly the reserved address:

78:7b:8a:bf:f3:ec > ff:ff:ff:ff:ff:ff   0.0.0.0.68 > 255.255.255.255.67: BOOTP/DHCP, Request
d0:21:f9:e2:2d:78 > 78:7b:8a:bf:f3:ec   10.0.24.254.67 > 10.0.24.200.68: BOOTP/DHCP, Reply

The reply never arrived, because the offer is unicast and its return path crossed a port that was not in VLAN 24. 0/28 is the AV-CORE uplink, and CORE had VLAN 24 tagged on its side (port 0/46) while AVoIP-DS did not — a trunk configured at one end only:

AVoIP-DS port Goes to VLAN membership 24?
0/26 EdgeRouter vlan participation include 2,21,24-25 ✅
0/27 SPAC-SWT01-CORE switchport trunk allowed vlan 1-2,5,21-25 ✅
0/28 AV-CORE 2,5,21-23,25 → 2,5,21-25 fixed 2026-09-17

So the broadcast DISCOVER flooded through and reached the router, and the unicast reply was dropped on egress at 0/28. The switch still learned the MAC on that port, which is why the FDB claimed a VLAN 24 entry for a port that could not forward VLAN 24 — ingress accepted, egress denied. That asymmetry is what made the earlier evidence look self-contradictory.

It was an omission, not a design choice: every other VLAN on the switch (1, 2, 5, 21, 22, 23, 25) already rode both uplinks, so 24 was the lone exception and adding it introduced no topology that was not already the steady state.

comms now reads 10.0.24.200 (DHCP4 via 10.0.24.254), routable (configured), with a connected route for the subnet and no fourth default route. Note the range consolidated on save — 2,5,21-23,25 became 2,5,21-25 — so do not grep switch configs for a literal 24.

The third default route, via an address that does not exist

Fixed 2026-09-16, on both ends. docker-server carried three default routes:

default via 192.168.0.1     dev ens9    metric 100    <- the real one
default via 192.168.30.254  dev control metric 200    <- real, EdgeRouter failover
default via 10.0.25.249     dev avoip   metric 1000   <- does not exist

The third answered neither ping nor ARP — ip neigh reported it INCOMPLETE. Metric ordering masked it, which is exactly what made it worth removing: any metric shuffle promotes it, and the box then looks offline while still resolving DNS, because the resolvers sit on directly-connected subnets. That is a nasty symptom to work backwards from.

It did not come from this host. The lease said so:

Address: 10.0.25.200 (DHCP4 via 10.0.25.254)
Gateway: 10.0.25.249

The lease is served by 10.0.25.254 — which exists and answers — and that server was advertising 10.0.25.249. One wrong line in the EdgeRouter’s AVoIP scope, and the same line behind the 44 SVSI endpoints reporting GW:10.0.25.249. Both halves are fixed:

  • Router — default-router 10.0.25.249 → 10.0.25.254, matching every other scope. See the AVoIP network map.
  • Host — Automation/server/scripts/fix-docker-server-avoip-default-route.sh drops in 61-avoip-no-default-route.yaml with use-routes: false. Worth doing even with the router correct: avoip is a multicast media VLAN and this box should not take a default route out of it either way. The connected 10.0.25.0/24 route comes from the address, not from DHCP, so the SVSi endpoints stay reachable — verified against 10.0.25.201.

docker-server now carries two default routes, both real: ens9 at metric 100, and control at metric 200 as a failover through the EdgeRouter. Since the EdgeRouter has no WAN lease of its own, treat that failover as reaching the AV VLANs, not the internet.

Never trust a ping to a VLAN this box has no route to

With no connected route, 10.0.24.x traffic leaves via the office uplink — and Telus answers out of its own RFC1918 space, so the ping succeeds and tells you nothing:

$ ping -c4 10.0.24.1          # "COMMS-BRIDGEX" per the router's static mappings
4 packets transmitted, 4 received, 0% packet loss   ← not the comms bridge

$ traceroute -n 10.0.24.1
 1  192.168.0.1
 2  204.191.255.1             ← already off-site
 3  204.191.152.183

This is the same trap as the 10.0.0.x addresses recorded elsewhere in this repo, except those time out honestly and this one returns a reply. Check ip route get before believing any result on a VLAN whose interface has no address.

Why it answers at all. Nothing has a route to Telus’s private space — 192.168.0.1 simply has a default route, and private addresses are private by convention, not by any mechanism in IP forwarding. The path to 10.0.24.1 is byte-for-byte the path to 8.8.8.8, and the gateway NATs the source on the way out, which is the only reason a reply can return. Telus then finds a live host at its own 10.0.24.1 and delivers it. Confirm who owns an upstream hop with Team Cymru rather than guessing from the range — dig +short 1.255.191.204.origin.asn.cymru.com TXT returns 852 | 204.191.0.0/16 | CA | arin, i.e. AS852 TELUS.

192.168.100.0/24 exists twice

The same subnet is in use on both sides of the boundary, as two separate L2 domains with no connectivity between them:

  192.168.100.1 .200, .201–.205, .208
From the AV management VLAN (docker-server .200) nothing — no ARP entry docker-server and the six M4250s
From the office LAN (DS3) the UDM, ttl=64, CN = unifi.local, ports 22/80/443/8443 no reply

A full sweep of the /24 from docker-server finds exactly seven hosts and no others: itself at .200, and AV-CORE .201, M21-SWITCH .202, CONTROL-ROOM-SWITCH .203, AVoIP-DS-SWITCH .204, ATRIUM-SWITCH .205, STAGE-SWITCH .208. Note the gap — .206 and .207 are not on this VLAN. Those octets belong to the two Cisco switches, which are reachable at 192.168.30.206 and 10.0.25.207 on entirely different VLANs; writing the range as .201–.208 implies eight management addresses when there are six.

So 192.168.100.1 is the Dream Machine on one of its networks, not a gateway on ours — which is why it answers undecremented from the office and is absent from a full sweep of the VLAN it appears to belong to. Asking “is 192.168.100.x up” gets a different answer depending which side you ask from, and both answers look authoritative.

Two switches point at gateways that do not exist. Core has ip default-gateway 192.168.100.1, which resolves to nothing on VLAN 2; Atrium has ip default-gateway 10.0.0.1, an address this repo already records as dead. The other four have no default gateway at all. This is benign today — everything the switches talk to (NTP, SNMP to docker-server on .200) is on their own subnet, which is why four of them run happily without one — but none of the six can actually route off the management VLAN, so do not plan anything that assumes they can.

It is also latent: if the two 192.168.100.0/24 domains were ever bridged or routed together, both would break.

What to ask IT for

The fix belongs on the Dream Machine, which IT owns, so this is a request rather than something to schedule. It is small and self-contained:

On the UniFi gateway at 192.168.0.1, add two static routes with a black-hole / discard target: 10.0.0.0/8 and 172.16.0.0/12.

These ranges are private and must not be forwarded to TELUS (RFC 1918 §3). Today they are: an address we have no route for leaves on the default route, TELUS finds a live host at its own copy of it, and answers — so a ping to an internal address that does not exist comes back successful. That has already produced false “up” readings in our monitoring.

This cannot break internal routing. A blackhole is the least-specific match, so every real subnet — 10.0.22.0/24, 10.0.23.0/24, 10.0.24.0/24, 10.0.25.0/24 and anything else with a route — keeps winning on longest-prefix. Only traffic that was already going nowhere is affected, and it now fails immediately instead of silently succeeding against a stranger’s host.

Please use a route, not a firewall rule. The gateway routes between several internal networks that are themselves RFC1918, so a rule matching private destinations would break that; a route only ever catches what has no better match.

Please do not include 192.168.0.0/16. The office LAN is inside it, and so is 192.168.100.0/24, which your own gateway uses — we independently use that same range for AV switch management on a separate VLAN, so a blackhole there would be ambiguous at best.

Worth asking them to confirm nothing legitimately reaches private space over the WAN (a site-to-site VPN, say) — if something does it needs a more-specific route, which would win anyway, but it is cheaper to ask than to discover.

Before the address landed, the five Comms monitors in Uptime Kuma (LX Headset, Projection Headset, Main Camera Headset, Switcher Rackmount Station, Shading Headset) had never once recorded an up-beat — they were reporting the missing route, not the headsets.

They are still down, and now that means something. With comms holding 10.0.24.200, the router on the VLAN answers and the gear does not:

10.0.24.254   UP                      <- EdgeRouter on the comms VLAN
10.0.24.1     down   (ARP FAILED)     <- BridgeX
10.0.24.2     down   (ARP FAILED)     <- LX
10.0.24.3     down   (ARP FAILED)     <- Projection
10.0.24.4     down   (ARP FAILED)     <- Main camera
10.0.24.7     down   (ARP FAILED)     <- Switcher
10.0.24.8     down   (ARP FAILED)     <- Shading

That is the whole point of fixing the route: the answer went from unknowable to none of the Green-Go gear is powered up. ARP FAILED against a VLAN we now demonstrably reach is a real negative, not a missing path. The stations are service-time gear like the signage encoders, so expect them dark outside a service — a red Comms monitor on a Tuesday is the expected reading, not an incident.

It is also on the tailnet as docker-server.tail8bb556.ts.net (100.99.33.108) and offers an exit node, which makes the same path work from off-site.

SSH config on the tech workstation:

Host DS1 DS2 DS3 DS4 ds1 ds2 ds3 ds4
  User spacadmin
  IdentityFile ~/.ssh/id_ed25519
  IdentitiesOnly yes
  ProxyJump docker-server

Host docker-server ds-jump
  HostName docker-server.tail8bb556.ts.net
  User spac-admin
  IdentityFile ~/.ssh/id_ed25519
  IdentitiesOnly yes

ssh DS1 then works on-site and remote alike. Authentication to the signage PC still happens from the workstation’s own key — the jump host only forwards the connection, so no key of yours needs to live on it.

ssh -o ProxyJump=none DS1 connects directly and skips the hop. That works from the office LAN and from the SPAC Staff SSID. SPAC Guest does not reach these hosts.

The wireless and office networks are not described in this repo, and the guest/staff separation is enforced by gear that is undocumented here. Note also that Networking/Router/config/config.boot is stale (see Content PCs), so its absence of wireless entries is not evidence of anything.

SVSi endpoint status is read from the docker-server with getStatus over TCP 50002; the reply is a colon-delimited key/value blob. STREAM gives the current subscription, DVISTATUS whether a display is attached.

Receiving a stream on docker-server

docker-server can take in any SVSi stream over avoip (10.0.25.200) without root: an ordinary UDP socket bound to the group (239.255.37.N:50002) plus a multicast join. That is how captures and live previews are made. A 1080p60 stream is ~65,000 datagrams a second (~90 MB/s), so the box’s receive buffers are sized for it, by Automation/server/scripts/setup-docker-server-rx-buffers.sh:

  • net.core.rmem_max is 8 MB (/etc/sysctl.d/60-spactech-rx-buffers.conf), so a socket that asks for a 16 MB receive buffer gets one: ~120 ms of stream. At Ubuntu’s default of 212992 the kernel grants 416 KB and charges ~2.1 KB per 1454-byte datagram, which is ~200 datagrams, about 3 ms. Sockets that don’t ask keep the 208 KB default.
  • The enp3s0f0 RX ring is 511 descriptors, the BCM57766’s maximum, set by spactech-rx-ring.service whenever the NIC appears. The default, 200, is also about 3 ms at this rate.

Why: this box stalls for milliseconds, stream or no stream. With nothing streaming, a thread pinned to any of its four logical CPUs (two cores, two hyperthreads each) that sleeps 0.5 ms at a time wakes more than 3 ms late every few seconds, worst 6–14 ms: a 250 Hz tick and voluntary preemption, on two cores shared with Companion, Prometheus, Grafana and Uptime Kuma. A 3 ms buffer cannot ride that out. The socket buffer now absorbs every stall seen so far (the longest drained in testing was 24 ms of stream). What still gets through is the ring overflowing when the kernel’s own packet processing waits more than ~8 ms for a CPU: about one frame in 5,000 arrives with missing lines.

To see losses: a socket’s drop count is the last column of /proc/net/udp (all sockets together: RcvbufErrors in /proc/net/snmp), and ring overflows are rx_discards in ethtool -S enp3s0f0.

Open questions

  • DS3 sees no link to its encoder. All four of its outputs read disconnected and the encoder (.82) reports INPUTRES:0x0. Neither a PoE cycle of the encoder nor a reboot of DS3 cleared it, which rules out the latched-input fault in Power-cycling an SVSi endpoint remotely; check the cable and both ports. Once it is back, DS3 has also never been paired in OptiSigns since its rebuild, so it will show its pairing screen (code 43HQ9L) until it is paired at app.optisigns.com, and DS SCS needs putting back on stream 13.
  • What feeds the display outside the office? It isn’t on an SVSi decoder, and it receives 1366×768.
  • Are any schedules programmed on the four NEC displays? Serial can’t show them; it needs a remote and the on-screen menu. See Signage displays.
  • Does DS4 have a job? Nothing subscribes to stream 14.
  • DigitalSignage5 (stream 15, 10.0.25.84) runs 24/7 in the cage with no source PC. A spare?
  • KidsMinHallway reports its display disconnected — but it also reads DVIOFF:on, meaning the decoder’s own HDMI output is switched off. Check that setting before a physical check.
  • DigitalSignageLegacyDistribution needs an IP, or removal from set_digital_signage_to_stream().
  • Why does DS3 scan out 10-bit (XR30) when DS4, on the same image and GPU generation, scans out 8-bit (XR24)? It only matters to spac-screenshot’s fallback now. Not the EDID — all four PCs receive EDID 1.3 with no bit depth and no HDMI deep-colour flags — and neither box has a monitors.xml or any mutter experimental feature set.
  • Add the PC models to Tech Inventory.xlsx: DS1/DS2 are ThinkCentre M710q, DS3/DS4 ThinkCentre M73 Tiny (machine type 10AX).
  • Why does DS2 have IPv6 ULA addresses when DS1 does not?
  • Should DS1’s OptiSigns remote agent be present on the other three, or removed from DS1?
  • Playback is started twice on every box — the systemd user unit and an OptiSigns-written GNOME autostart entry. Pick one and remove the other; see Layout.
  • Grafana has anonymous access with the Admin role (GF_AUTH_ANONYMOUS_ENABLED=true, GF_AUTH_ANONYMOUS_ORG_ROLE=Admin). Anyone who can reach :3000 can edit dashboards, alert rules and datasources without logging in. That is how backup-config.sh works without a token. Fine on an internal VLAN, worth a decision rather than an accident.
  • ufw and sshd config are unaudited (need root). Check when next running something with sudo.
  • The office and wireless networks (including what separates SPAC Guest from SPAC Staff) are not described anywhere in this repo. Worth documenting alongside the AV VLANs.