host-agent applies OS updates: install to the other slot, switch in the window, revert on its own (#563) - #574
Conversation
…oot, revert (#563) The update-target answer gains an OS list; host-agent installs the next release into the other slot with RAUC after checking the bundle's sha256 against the answer, switches inside the window, and marks the new slot good once the brain is healthy. A trial that fails reboots to the old slot; an image timer reboots a slot whose host-agent never started. Two new boots prove both paths under UEFI and BIOS.
|
| if err := os.Remove(TrialMarker(a.dir(), slot)); err != nil { | ||
| slog.Error("os update: could not remove the trial marker", "err", err, "slot", slot) | ||
| } |
There was a problem hiding this comment.
There was a problem hiding this comment.
Fixed in b638594. The trial records the good outcome before it removes the marker. Removal is retried for about 13 min, under the 15 min timer. A later boot that finds the good outcome for that switch only removes the marker. The image timer cannot use the grubenv to tell good from on-trial, because GRUB under UEFI does not save TRY (#575). Tests: TestStaleMarkerOnAGoodSlot, TestTrialStaysUpWhenMarkGoodFails.
There was a problem hiding this comment.
Covered by b638594 and 02ca9f4. The trial records the good outcome before it removes the marker. Removal is retried for about 13 min, under the 15 min timer, and the trial's deadline now ends 2 min before the timer. A later boot that finds the good outcome for that switch only removes the marker and does not start a new trial. Tests: TestStaleMarkerOnAGoodSlot, TestTrialStaysUpWhenMarkGoodFails, TestTrialDeadlineCountsFromBoot.
There was a problem hiding this comment.
Dismissing what is left. The remaining case needs the trial to fail to delete one small file on /state for about 13 min in a row, after it already wrote the good outcome there. A /state that can write but not delete for that long is a broken disk, and the box has bigger problems than an OS revert. Even then the result is a revert to the known-good OS and a retry the next night, not a broken box. The safety net cannot read GRUB's TRY under UEFI until #575 is fixed. #575 is where the timer gets a second signal.
There was a problem hiding this comment.
Understood. Since the good outcome is persisted before marker cleanup, the remaining failure requires an unusually long-lived /state delete failure. Even then, the safety behavior returns to the known-good slot and retries on a later night; the missing UEFI TRY signal remains tracked under #575.
Nothing installs or switches while the booted slot is on trial, and the box switches its OS at most once a night. A bad os part, a list that leaves a minor out, and a downgrade without a readable floor are refused for stream A only. A failed reboot undoes the switch; the image timer leaves a slot already marked good alone; rauc status is retried at start; an unreadable record is never written over; the download is capped. A local run builds the test bundle itself.
#563) Debian's rauc-service marks every booted slot good at the end of boot, so a slot whose host-agent never started was marked good and the box could not fall back. The os-revert boot caught it.
GRUB under UEFI does not save the booted slot's TRY flag (CI run 37045179497), so a grubenv check made the safety net never fire under UEFI. host-agent now retries removing the marker after mark-good instead. The rauc-mark-good mask is dropped: Debian ships no such unit.
Second review: a trial that cannot mark its slot reboots at most once, the image timer reboots a slot at most once, a switch counts only once activated, leftover downloads are removed, the trial timeout is capped under the image timer, and a stale marker on a good slot is only removed.
Third round: the trial ends 13 min after boot at the latest, 2 min before the image timer; a new switch clears the slot's old notes; a slot that gives up its trial leaves trial-failed-<slot>, so a power cut between mark-active and the activated flag still records the revert.
| if timeout < minTrialWait { | ||
| timeout = minTrialWait | ||
| } |
There was a problem hiding this comment.
If host-agent starts the trial after about 12 minutes 50 seconds of boot time, this limit gives the brain just ten seconds to answer. The brain may still be starting, so host-agent marks a healthy new slot bad and reboots to the old OS. A late start needs a way to avoid treating a slow brain as a failed update.
There was a problem hiding this comment.
Dismissed, by design. The 15 min image timer is the outer limit for any trial boot, and the trial must decide before it. A box whose host-agent only starts the trial about 13 min after boot is already far outside a normal boot (seconds to about a minute in the boot proofs). Treating it as a failed update and going back to the known-good OS is the safe choice. Extending the trial past the timer would let a slot that never gets healthy keep running. The next night retries the update.
There was a problem hiding this comment.
Understood. Given the 15-minute image timer is the hard outer bound, clamping the host-agent trial to decide before it is intentional: a very late host-agent start leaves only the remaining safety window, and treating that as a failed trial avoids letting an unverified slot run past the timer. Dismissed.
Closes #563. A slice of #486.
What this does
host-agent now applies OS updates on the hosted profile (
UPDATES.md# 1):oslist. host-agent picks the next release: it never skips a minor, goes at most one minor back, and never goes below the running control plane's floor.DECISIONS.md2026-10-02).rauc installwrites the other slot ahead of the window. The boot order does not move (activate-installed=false).rauc status mark-active other. It checks that the grubenv readsOK=1 TRY=0for the new slot (the Hosted image in the A/B layout: 1 GiB squashfs slots, a state partition, GRUB on both firmwares (#561) #570 gap), then reboots.moose-os-trial.timer, 15 min) reboots the box. Only the first boot after a switch is on trial.os_versionandos_slot. The update-target read reports stream A's decision. Admins get one notification per outcome.The appliance build has no OS applier yet (#564) and reports stream A as
unsupported.The private-side change this needs (for the maintainer)
No production box moves its OS until the control plane sends the OS part. This PR does not touch the private side. The exact wire change for
GET /api/updates/target:os: a JSON array, oldest first, of{"version": "X.Y.Z", "bundle_url": "<url>", "bundle_sha256": "<64 lowercase hex>"}.osout when there is no OS target.bundle_url.https://github.com/onmoose/os/releases/download/vX.Y.Z/moose-vX.Y.Z-amd64.raucb. The box refuses any URL outside that prefix.bundle_sha256. The hex digest in the release'smoose-vX.Y.Z-amd64.raucb.sha256asset, read once when the release is recorded. Never resolved per request, and never "latest".ospart makes the box refuse stream A only. Stream B is not affected.Tested
make checkandmake check-webare green. There are new tests for the pick rules, the loop ordering, the transaction against a fake two-slot RAUC, the report, the brain pass-through, the notification and the outcome check.os-update: a wrong digest is refused, the install happens ahead of the window, the box switches, boots slot B and marks it good. The owner, session, password hash, SSH host key, machine-id, data and app are intact.os-revert: host-agent is kept off slot B, so the image's safety net reboots the box, and it goes back to slot A on its own.02ca9f4; later commits change only docs). It is green, with all 14 boot jobs under UEFI and legacy BIOS. All runs are listed indocs/progress/host-agent-os-update.md.Numbers
rauc install3 to 6 s, switch to marked good 17 s, reboot included. The guest reads from a local file server, so a real download adds the network time.Found on the way
TRYflag (GRUB under UEFI does not save the slot's TRY flag, so a panicking OS slot is never skipped #575, filed). Under BIOS it does. This PR does not depend on it: host-agent and the image's trial timer both mark a failed slot bad, and the revert boots prove that under both firmwares. Still open: under UEFI, a new slot that panics before userspace boots again and again. GRUB under UEFI does not save the slot's TRY flag, so a panicking OS slot is never skipped #575 tracks it.Notes
CLAUDE.mdgains the log fieldsos,slotanddigest, and nothing else./healthzonly, no report back to the cloud, and the test bundle is not a release bundle.