Skip to content

A/B OS image: a slot that hangs, and a Debian major across /etc (#486) - #580

Merged
onel merged 7 commits into
devfrom
feat/486-hang-and-major
Oct 6, 2026
Merged

onel merged 7 commits into
devfrom
feat/486-hang-and-major

Conversation

@onel

@onel onel commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

A slice of #486. It settles points 3 and 5 of NEXT.md # A/B OS image. Read docs/progress/ab-hang-and-major.md first.

Point 3: a slot that hangs

  • Hetzner Cloud has a hardware watchdog. Checked on real cx23 VMs (the smallest box), on one Intel and two AMD hosts. The VM is QEMU q35, and its ICH9 chipset has a TCO watchdog. Debian's -cloud kernel has no driver for it, but the image runs the generic linux-image-amd64, whose iTCO_wdt drives it, and systemd feeds it (RuntimeWatchdogSec=60s, already in the image).
  • It fires. With PID 1 frozen (ptrace, no sysrq) the VM reset after 118 s, with no shutdown. Unfed with a 30 s timeout, it reset after about 60 s.
  • No image change. Every boot of the lane now checks systemd feeds the watchdog. The lane is q35 too, so BUILD.md's "none in the QEMU lane" was wrong.
  • Two gaps, written in BUILD.md # 1b # As built: the reset comes after about twice the timeout (about 2 min), and a hang in GRUB or the initramfs is not covered (the watchdog is not started earlier, so a long e2fsck cannot reset a box again and again). No hang proof in CI, by decision: the Hetzner probe is the proof.
  • The probe cost about €0.035. Every server and key it made was deleted and the deletion confirmed.

Point 5: a Debian major across the /etc overlay

DECISIONS.md 2026-10-05, BUILD.md # 1b rule 4.

  • When the slot's Debian major differs from the one recorded on the state partition, in either direction, state-setup rebuilds the /etc upper layer before it mounts /etc:
    • the files on /usr/lib/moose/etc-keep.list stay;
    • passwd, group, shadow and gshadow are merged;
    • daemon.json and login.defs are taken again from the slot when the remap stays;
    • everything else goes to an attic on the state partition.
  • A failed tidy-up never stops a boot. CI showed why: a bug panicked slot B, and slot A, running the same step, panicked too.
  • host-agent turns sshd back on at start when the drop-in names an account (sshaccess.EnsureOnAtStart). It never turns sshd off.
  • The image's accounts become a generated sysusers.d file. Build check 6 fails when an id differs from dev/os-lock/cloud-accounts.lock, which is edited by hand.
  • Two new CI checks:
    • The os-update boot fakes a major (Debian 12 to 13) and checks the tidy-up.
    • Every boot fails on an upper-layer file that no rule covers.
  • It must ship in a Debian 13 release before the first Debian 14 release.

Costs

  • Disk: none for point 3. For point 5, the attic is a few hundred KB per major crossed, under 0.001% of the smallest box.
  • Smallest box: I used the 40 GB cx23, because that is what BUILD.md and the cloud size table say. It is not 64 GB.
  • CI: no new boot. os-update stayed at 187 to 208 s.

Proof

CI / Cloud image run 37378140993 on code head f7108c7: the full list, all 14 jobs green, 14.6 min wall. The commit after it changes only the progress entry. make test-nopam is green.

Known gaps

Co-Authored-By: Claude noreply@anthropic.com

onel and others added 3 commits October 5, 2026 22:46
Point 3 (a slot that hangs): Hetzner Cloud VMs are QEMU q35 with the ICH9
TCO watchdog, which the image's generic kernel drives. Proved on real cx23
VMs: a frozen PID 1 reset the VM after 118 s. No image change; every boot of
the lane now checks systemd feeds the watchdog.

Point 5 (a Debian major): when the slot's major differs from the one the
state partition records, state-setup tidies the /etc upper layer before it
mounts it: the keep list stays, the account files are merged, the pinned
files are taken again from the slot when the remap stays, the rest moves to
an attic. A failure never stops a boot. host-agent turns sshd back on at
start when the drop-in names an account. The image's accounts become a
generated sysusers.d file, with ids checked against
dev/os-lock/cloud-accounts.lock. The os-update boot fakes a major, and every
boot fails on an upper-layer file no rule covers.

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Claude <noreply@anthropic.com>
…ap; sshd re-enabled only after a tidy-up

Review round 1 on #580.

Co-Authored-By: Claude <noreply@anthropic.com>
@greptile-apps

greptile-apps Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] OS image build and boot infrastructure for Debian major upgrades.

The PR appears safe to merge based on this review.

Findings

  1. P1 SSH stays off after interruption ▶

Reviews (3) · Last reviewed commit: "fixup: a repeated swap keeps the sshd ma..."

Comment thread internal/hostagent/sshaccess/sshaccess.go
Comment thread dev/cloud/cloud-assertions.sh
onel and others added 2 commits October 5, 2026 23:23
Co-Authored-By: Claude <noreply@anthropic.com>
…ert direction

Greptile on #580: an admin's sshd off stays off after a major, and the
revert direction (newer record, older slot) is now crossed in os-revert.

Co-Authored-By: Claude <noreply@anthropic.com>
Comment on lines +249 to +252
if [ -n "$sshd_on" ]; then
touch "$TIDIED_MARKER" 2>/dev/null || log "tidy-up: could not leave $TIDIED_MARKER for host-agent"
else
rm -f "$TIDIED_MARKER"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 SSH stays off after interruption

If power is lost after the tidy-up swaps the /etc upper layer but before it records the new Debian major, the next boot repeats the swap. The first swap removes the SSH enable link and leaves a marker for host-agent. The second sees no link and deletes that marker. host-agent then leaves SSH off even though the owner had enabled it. Keep the pending marker when a swap is retried.

Knowledge Base Used: Immutable image layout and boot

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid. Fixed in 6fd8006: a swap no longer removes an existing marker. It only creates one when the old layer had ssh.service enabled. A repeat swap after a power cut (before the record) now keeps the marker from the first swap, so host-agent still turns sshd back on once. A stale marker cannot turn sshd on against an admin's choice, because host-agent consumes it at its first start after the boot that wrote it.

@onel
onel merged commit 0b31a36 into dev Oct 6, 2026
19 checks passed
@onel
onel deleted the feat/486-hang-and-major branch October 6, 2026 11:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant