Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/progress/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -316,3 +316,4 @@ Oldest first; append new entries to the bottom.
| [uefi-grub-try.md](uefi-grub-try.md) — Closes #575, a slice of #486, after [host-agent-os-update.md](host-agent-os-update.md). **GRUB's try flag now survives under UEFI in the boot lane.** The cause was the harness, not GRUB or the image: it ran OVMF with no writable VARS store (Ubuntu 24.04 renamed the files), so OVMF saved its variables to an `NvVars` file on the ESP at every boot and lost GRUB's `save_env` write (traced in runs 37061482623, 37063254219 and 37064755329: GRUB wrote and read back `A_TRY=1`, the initramfs read `A_TRY=0`). The harness now needs a CODE and VARS pair. Every boot checks GRUB saved `<booted slot>_TRY=1` and the ESP has no `NvVars`, under both firmwares, and `os-revert` gains a last stage: slot B made active again crashes its kernel in the initramfs (a test-only hook), and GRUB must skip it. The trial timer stays on the marker, not the grubenv. No image change, no disk cost. **Gaps:** QEMU only, a real UEFI provider box not checked for `NvVars`; the appliance medium lane has the same OVMF fallback (#564) | done |
| [bake-released-control-plane.md](bake-released-control-plane.md) — Part of #566 (kept open until the released path is green), a slice of #486, after [control-plane-version-line.md](control-plane-version-line.md). **An OS-only release bakes the last released control plane:** the brain and UI of `v<CONTROL_PLANE_VERSION>`, resolved to digests once and pulled from ghcr by digest (`make control-plane-released`, `dev/release/ghcr-resolve.sh`). A control-plane release, or one that bumps both files, bakes the pair it builds and releases; a PR, a lock bump or a dispatch bakes a build of its commit (`-f control_plane=released` checks the release path ahead). The image records the pair in `/usr/lib/moose/control-plane.env`, and every boot checks the running brain and UI against it. The released run (37345392679) baked 0.15.0 and passed the six original gate boots under both firmwares; `os-update` and `os-revert` fail with it because the 0.15.0 brain predates #563, so the next OS release must also release the control plane. |
| [ab-hang-and-major.md](ab-hang-and-major.md) — A slice of #486, after [hosted-ab-layout.md](hosted-ab-layout.md) and [host-agent-os-update.md](host-agent-os-update.md). Settles points 3 and 5 of `NEXT.md` # A/B OS image. **A slot that hangs:** real Hetzner cx23 VMs (Intel and AMD hosts) are QEMU `q35` with the ICH9 TCO watchdog; the image's generic kernel drives it (`iTCO_wdt`), and with PID 1 frozen the VM reset after 118 s. No image change; every boot checks systemd feeds it. Gaps: the reset comes after about twice the 60 s timeout, and a hang in GRUB or the initramfs is not covered. **A Debian major:** when the slot's major differs from the one the state partition records, in either direction, the initramfs tidies the `/etc` upper layer: the keep list (`etc-keep.list`) stays, the account files are merged, the pinned files are taken again from the slot when the remap stays, the rest goes to an attic; a failure never stops a boot. host-agent turns sshd back on once after a tidy-up when the drop-in names an account. The image's accounts are a generated `sysusers.d` file, with ids checked against `dev/os-lock/cloud-accounts.lock`. The `os-update` boot fakes a major, and every boot fails on an upper-layer file no rule covers. It must ship in a Debian 13 release before the first 14 release | done |
| [os-switch-undo-marker.md](os-switch-undo-marker.md) — A fix to [host-agent-os-update.md](host-agent-os-update.md) (#563), part of #486, found by Greptile on the 0.16.0 release PR (#583). **A failed switch undo keeps the trial marker:** `doSwitch` removed the new slot's trial marker even when it could not put the booted slot first again, so the new slot could boot next with no trial. The marker now goes only after the booted slot is first again; `Boot` already removes a marker for a switch that never took effect. Unit test only | done |
22 changes: 22 additions & 0 deletions docs/progress/os-switch-undo-marker.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# A failed switch undo keeps the trial marker

- **Status:** done
- **Date:** 2026-10-06
- **Specs touched:** `docs/specs/UPDATES.md`

A fix to [host-agent-os-update.md](host-agent-os-update.md) (#563), part of #486. Greptile found it on the release PR for 0.16.0 (#583).

## What was done

- **The bug.** When an `os-switch` job stops after `rauc status mark-active other` (for example, the grubenv check finds `TRY=1`, or the record cannot be saved), it undoes the switch: it puts the booted slot first again and removes the new slot's trial marker. It removed the marker even when putting the booted slot first failed. Then the new slot could still boot next, with no trial marker, so no safety net and no revert, and host-agent would mark it good.
- **The fix** (`internal/hostagent/osupdate/osupdate.go`, `doSwitch`): the marker is removed only after the booted slot is first again. If that fails, the marker stays and host-agent logs it. If the old slot boots after all, `Boot` already removes a marker for a switch that never took effect (`TestMarkerWithoutActivationIsNoRevert`).
- **Test:** `TestFailedUndoKeepsTheTrialMarker` makes the grubenv check fail and the undo's mark-active fail, and wants the marker kept. It fails without the fix.
- `UPDATES.md` # 1 says it in the switch step.

## What's next

Nothing new. The release of 0.16.0 (#583) goes on after this lands.

## Known gaps

- Only unit-tested: making two RAUC calls fail in a row in a booted image needs a test hook that the boot lane does not have, and it is not worth one for this path.
2 changes: 1 addition & 1 deletion docs/specs/UPDATES.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,7 +64,7 @@ The OS underneath us: kernel, libc, OpenSSL, firmware, Docker itself, and `host-

- **The answer's OS part** is an optional `os` list, oldest first, each entry `{version, bundle_url, bundle_sha256}` (# 8.4 has the wire). The box refuses an entry with no 64-hex sha256, a version that is not a plain `X.Y.Z`, a list out of order, or a `bundle_url` that does not start with the expected prefix (`MOOSE_UPDATE_OS_URL_PREFIX`, default `https://github.com/onmoose/os/releases/download/`). It then picks as # The update transaction says: the next minor's entry, the target in its own or the previous minor, or a refusal. It also refuses a list whose next step is not the minor right after the box's own (a list that leaves a minor out), and a release below the running control plane's floor: the brain writes `minimumAgentVersion` to `/var/lib/moose/state/minimum-host-agent` at start. **A box that cannot read that file refuses every move to an older release.** An `os` part of the wrong JSON shape is refused like a bad entry. **Each stream is judged on its own**: a bad OS part holds back stream A only, and a bad control-plane part holds back stream B only. The running OS release is `host-agent`'s own version, which is the version of the slot it ships in.
- **Install, ahead of the window** (an `os-install` job, under the same lock as `system-update`). The bundle is downloaded to `/var/lib/moose/os-update/` on the state partition and its sha256 checked against the answer before RAUC sees it. Then `rauc install`, which checks the signature and writes the other slot. `system.conf` has `activate-installed=false`, so the boot order does not move. The file is deleted afterwards. A failed attempt is not repeated the same night for the same release and digest.
- **Switch, inside the window** (an `os-switch` job), and only when stream B does not hold the window (it did not just start an update, and no job runs). `rauc status mark-active other` puts the new slot first with `OK=1 TRY=0`; RAUC's GRUB backend resets the slot's try flag there, and `host-agent` checks the grubenv says so before it reboots. It writes a trial marker for the new slot (`/var/lib/moose/os-update/trial-<slot>`) and the record (`state.json`), then reboots. A reboot that is not accepted undoes the switch. **One OS switch per night**, so a box several minors behind takes one step per window. **Nothing installs or switches while the booted slot is on trial**, because the other slot is then the way back.
- **Switch, inside the window** (an `os-switch` job), and only when stream B does not hold the window (it did not just start an update, and no job runs). `rauc status mark-active other` puts the new slot first with `OK=1 TRY=0`; RAUC's GRUB backend resets the slot's try flag there, and `host-agent` checks the grubenv says so before it reboots. It writes a trial marker for the new slot (`/var/lib/moose/os-update/trial-<slot>`) and the record (`state.json`), then reboots. A reboot that is not accepted undoes the switch. An undo removes the trial marker only after the booted slot is first again; if that fails, the marker stays, so a new slot that still boots next boots on trial. **One OS switch per night**, so a box several minors behind takes one step per window. **Nothing installs or switches while the booted slot is on trial**, because the other slot is then the way back.
- **The trial.** On the new slot, `host-agent` waits for the brain's `/healthz` for up to `MOOSE_OS_TRIAL_TIMEOUT` (10 min, at most 14), and never past 13 min after boot: the image's timer counts 15 min from boot, and the 2 min between are for the final RAUC mark, so `host-agent` always decides first. Healthy: `mark-good`, the marker goes, the outcome is `good`. Not healthy: `mark-bad booted` and a reboot, and GRUB boots the old slot. **A `host-agent` that never starts** cannot do either, so the image carries `moose-os-trial.timer` (15 min after boot): if the booted slot still has its marker, it marks the slot bad, leaves a note (`safety-net-<slot>`) and reboots. The timer goes by the marker only, never by the grubenv's `TRY` flag: GRUB writes that flag through the firmware, and a firmware can lose the write (#575), so a hung slot that looked good would never be rebooted. On the old slot, `host-agent` finds the new slot's marker, records the outcome `reverted`, marks that slot bad, logs whether the safety net made the revert, and stays. The same release is tried again the next night. Only a switch that took effect counts: the record's `activated` flag is written after `mark-active` succeeds, and a slot that gives up its trial leaves a note (`trial-failed-<slot>` from `host-agent`, `safety-net-<slot>` from the timer). So a power cut before `mark-active` never reads as a revert, and a power cut right after it still records one. A new switch clears the old notes of its slot. When RAUC cannot mark the slot, the trial does not loop: a slot it cannot mark bad is rebooted once, a slot it cannot mark good is not rebooted, and the image's timer reboots a slot at most once; after that the box stays up for a person. A download left by a kill or a power cut is deleted at the next start.
- **Report.** `GET /api/v1/system/version` carries `os_version` and `os_slot`; `GET /api/v1/system/update-target` carries an `os` object beside stream B's fields (`BRAIN_HOST_PROTOCOL.md`). The brain raises one admin notification per outcome (`NOTIFICATIONS.md` # Updates). Reporting back to the cloud (# 8.4 step 5) still waits for real box authentication.
- **What it costs a box** (CI, `../progress/host-agent-os-update.md`): one bundle download per update (438.5 MB for a real release), the same space on the state partition while it installs (1.1% of a 40 GB disk), and one reboot.
Expand Down
7 changes: 6 additions & 1 deletion internal/hostagent/osupdate/osupdate.go
Original file line number Diff line number Diff line change
Expand Up @@ -774,8 +774,13 @@ func (a *Applier) doSwitch(ctx context.Context, rel updatetarget.OSRelease, nigh
return err
}
undo := func(cause error) error {
// The trial marker goes only once the booted slot is first again. If
// that fails, the new slot may still boot next, and it must boot on
// trial. A marker left while the old slot boots is removed by Boot as a
// switch that never took effect.
if err := a.RAUC.Mark(context.Background(), "active", "booted"); err != nil {
slog.Error("os update: could not put the booted slot first again", "err", err)
slog.Error("os update: could not put the booted slot first again; the trial marker stays", "slot", target, "err", err)
return cause
}
os.Remove(TrialMarker(a.dir(), target))
return cause
Expand Down
27 changes: 27 additions & 0 deletions internal/hostagent/osupdate/osupdate_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -220,6 +220,33 @@ func TestStaleTryStopsTheSwitch(t *testing.T) {
}
}

// failingUndo fails only the mark-active that puts the booted slot first again.
type failingUndo struct{ *fakeRAUC }

func (f failingUndo) Mark(ctx context.Context, state, which string) error {
if state == "active" && which == "booted" {
return errors.New("rauc: d-bus timeout")
}
return f.fakeRAUC.Mark(ctx, state, which)
}

// When the switch stops after mark-active and the booted slot cannot be put
// first again, the new slot may boot next: its trial marker must stay, so it
// boots on trial (Greptile on #583).
func TestFailedUndoKeepsTheTrialMarker(t *testing.T) {
h := newHarness(t, "A")
h.a.Apply(h.rel, false, h.night)
h.rauc.staleTry = true
h.a.RAUC = failingUndo{h.rauc}
h.a.Apply(h.rel, true, h.night)
if h.jobs.errs[1] == nil || h.reboots.Load() != 0 {
t.Fatalf("a TRY=1 after mark-active must stop the reboot: %v, reboots %d", h.jobs.errs[1], h.reboots.Load())
}
if _, err := os.Stat(TrialMarker(h.a.dir(), "B")); err != nil {
Comment thread
greptile-apps[bot] marked this conversation as resolved.
t.Fatalf("the trial marker must stay while slot B may boot next: %v", err)
}
}

func TestBusyLockWaits(t *testing.T) {
h := newHarness(t, "A")
h.jobs.busy = true
Expand Down
Loading