Skip to content

Release 0.16.0 - #583

Merged
onel merged 59 commits into
mainfrom
release/0.16.0
Oct 6, 2026
Merged

onel merged 59 commits into
mainfrom
release/0.16.0

Conversation

@onel

@onel onel commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Release 0.16.0, on both lines: the OS (VERSION 0.16.0) and the control plane (CONTROL_PLANE_VERSION 0.16.0), bumped in #582. The headline: the hosted image is now an A/B image, and a box can update its own OS from a signed bundle, and go back on its own when the new OS fails. This release also carries a Debian security fix (libpng, #571).

What is new

Testing

Upgrading

  • No box changes its OS until the cloud control plane names a bundle in its update target. That part is not live yet. So this release does not move any existing box's OS.
  • New hosted boxes made from this image can take OS updates once the update target names them. Boxes made from 0.15.0 or earlier have the old disk layout. They keep their OS, as before.
  • Merging this PR tags v0.16.0 and control-plane-v0.16.0, creates both Releases, runs the full boot proof, signs the bundle, attaches the image and the bundle, and pushes the brain and UI images to ghcr.

onel added 30 commits October 1, 2026 09:44
Records the RAUC and systemd-sysupdate proofs from CI run 36836466471,
answers the #485 questions, prices the UEFI-only hosted fallback, and
recommends RAUC with GRUB on both firmwares.
Pin login.defs too (SUB_UID_COUNT 0 from #530). Say the data partition is
on the OS drive, so the recovery key stays on the encrypted OS drive. The
hosted root grows to the whole disk today, so state the real trade-off of a
fixed 4 GiB slot, including /var outside the bind mounts.
docs: A/B update engine spike write-up (#485)
Stream A becomes an A/B OS image on both profiles, built with RAUC and
GRUB, with host-agent inside it. Appliance OS updates apply and reboot in
the window. A moose release is the OS release, and the control plane gets
its own version line. Specs only; the build is split into #559 to #564.
The manifest floor never holds back an OS update. Tier-2 packages are baked
into the image. An OS downgrade goes back one release at most. A lock bump
ships with the next VERSION release, not on merge. Mark the reboot-required
detector and the pending-reboot notification as retiring.
Greptile round 2. A box that skipped a release could revert to one that
cannot read its state, so formats change only in minor releases and a box
steps through each minor. An OS downgrade also respects the control plane's
floor. Tier-2 packages keep the run state their own spec gives them.
Greptile round 3. A box that never skips a minor must learn the newest
patch of each minor on the way, so the hosted answer and the manifest's
os field carry that list with each bundle digest.
Greptile round 4: say the last entry is the target, and cover the same-minor
and one-minor-back cases, so a deliberate downgrade can be selected.
Greptile round 5: 2.0 follows 1.9, so stepping and the one-back downgrade
work across a major boundary.
docs: design the A/B OS update and split the version lines (#486)
moose now has two version lines, one per update stream (DECISIONS.md
2026-10-01). VERSION stays the moose (OS) release and keeps stamping
host-agent, so the brain's minimumAgentVersion floor still compares
against the moose version with no wire change. The brain and both
control-plane images belong to the new CONTROL_PLANE_VERSION line, which
starts at 0.15.0 so nothing a box reports jumps.

The brain's --version now reads "moose control plane X.Y.Z (g<sha>)", so
its number is never read as the OS release. host-agent keeps
"moose X.Y.Z (g<sha>)", which the cloud boot proof greps. Both images
gain OCI version and revision labels.

Part of #559.
release.yml now decides each line on its own: a VERSION bump tags vX.Y.Z
and attaches the disk image, a CONTROL_PLANE_VERSION bump tags
control-plane-vX.Y.Z and pushes the brain and UI images to ghcr. A merge
that bumps both cuts both in one cloud-image run, and one that bumps
neither stays a green no-op.

The ghcr image tag keeps the vX.Y.Z shape, with the control-plane
number, because the private control plane resolves digests by that tag.
That shape is shared with every release before the split, so the
control-plane line refuses a version whose image tag is already
published rather than overwrite released images with new bytes.

A control-plane-only release still runs the full boot proof and skips
only the disk-image attach, so a control plane is only published after
it booted. The control-plane GitHub Release is created with
--latest=false so "Latest" stays on the OS release and its disk image.

The decision moved into dev/release/decide.sh so it can be tested off
main; decide_test.go runs it against throwaway git repos.

Part of #559.
BUILD.md # Versioning drops its "designed, not built" marker and says
which file stamps which binary, how each line is cut, and why the ghcr
image tag keeps the vX.Y.Z shape. contributing.md # Release model now
says how to cut each line, and records the one-time
control-plane-v0.15.0 hand tag that the overwrite guard needs before the
next release merge. DECISIONS.md 2026-10-01 gains a one-line note that
control-plane-vX.Y.Z is the git tag, not the image tag.

The same-commit bake of the control plane into OS images is recorded as
a known gap, with #566 as the follow-up.

Closes #559.
A run that tagged one line and then failed on the other tag or a
GitHub Release could not be re-run: decide.sh saw its own tag and hit
the bumped-to-an-already-tagged-version error. A tag at the commit
being released now means "already cut here, carry on": release.yml
does not tag again, creates the Release only on a definite 404, and
the publish steps add only what is missing. A tag at another commit
stays a hard error.
Either image already published under the control-plane version is a
reason to stop, so ghcr-tag-exists.sh now takes several repositories
and release.yml checks both in one call. An unclear answer from any of
them, with none published, is still "unknown", never "missing".
Re-publishing only the control plane through dispatch used to
re-upload the OS disk image with --clobber over a published Release.
Dispatch now also takes publish_os and publish_control_plane, and keeps
publish (both lines, default true) so the documented build-only command
still works. The attach uploads only files the Release lacks, without
--clobber, and the ghcr push skips an image whose tag is already there.
A rebuild is not byte-identical, so replacing a published file is left
to a person who deletes it first.
…guards

DECISIONS.md 2026-10-01 still said the workflows read one VERSION above
the new as-built note. It now says "Built in #559" once. BUILD.md,
contributing.md and the progress entry describe the resume path, the
two-image check and the per-line dispatch inputs.
A resumed attach uploaded only missing assets, so a run that died
between the .raw.xz and the .sha256 paired an old image with a checksum
of a new build. dev/release/attach-image.sh now leaves both, uploads
neither-or-both, and refuses a half pair with the delete-asset command
a person runs first. Nothing published is deleted.

A resumed ghcr push skipped an image whose version tag existed, so a
push that died before latest left latest stale, and brain and ui could
disagree. For such an image, dev/release/ghcr-retag.sh copies the
version tag's manifest to latest byte for byte through the registry
API and checks the digest, with no rebuild or re-push, and only when
this run publishes that version. Both scripts are tested with stubs.
Re-running an older tagged control-plane release passed the commit
guard, and the latest repair then pointed brain and ui latest back at
older images. A dispatch of an old version could do the same through
the normal push. Both paths now move latest only when this version is
the newest control-plane-v* tag by semver (dev/release/is-newest.sh).
An unreadable tag list leaves latest alone with a warning and still
publishes the version tag.
Control plane gets its own version line; moose vX.Y.Z is the OS release
* tmp: probe mkosi Snapshot= for trixie in CI (#560, removed before the PR)

* tmp: remove the snapshot probe workflow (#560)

* Add the OS package lock: Debian snapshot, Docker pins and the oslock tool (#560)

With no apt on the box, every package version is chosen at build time, so the
build must pin them. dev/os-lock/ holds the snapshot.debian.org timestamp, the
exact Docker versions (Docker's repo is not in the snapshot), and the shared
staging that turns them into apt sources and pins. oslock normalizes mkosi's
manifest, checks it against the committed list, and names changes for the bump
PR. cloud-packages.lock itself comes from the first CI build.

* Build the hosted images from the OS package lock (#560)

Both the lean image and the boot-proof image now install from the locked
snapshot and pins and must resolve to dev/os-lock/cloud-packages.lock. Two
separate builds of one commit checking the same list is the proof that the lock
holds. trixie-updates is added back from the snapshot, because mkosi drops it in
snapshot mode.

* Add the daily OS lock bump and run full boots on lock PRs (#560)

The bump opens or updates one PR when OS_LOCK_BOT_TOKEN exists, because a PR
made with GITHUB_TOKEN can neither be opened here (#484) nor start CI. Without
the token it pushes the branch, boots it by calling ci-cloud-image.yml with the
new ref input, and leaves a summary and a tracking issue with an open-a-PR
link. A PR touching dev/os-lock/ boots the full list, and the mkosi setup is
shared so the bump resolves exactly what CI checks.

* Commit the resolved package list of the hosted image (#560)

Taken from the first CI build at the locked snapshot and pins (run
36873363298, artifact cloud-packages-lock). From now on both cloud builds must
resolve to exactly this list.

* Document the OS package lock and the OS patch release process (#560)

BUILD.md # 1b gets the as-built lock. The maintainer decided how a lock-only
release is cut (from main as hotfix/X.Y.Z, within 7 days for a trixie-security
change), so contributing.md, DECISIONS.md and NEXT.md point 6 record it, along
with how the bump PR is opened and what the bot token needs.

* tmp: run the OS lock bump on this branch (#560, removed before the PR)

* Group a called cloud-image run by the ref it builds (#560)

The bump calls ci-cloud-image.yml from dev with ref=bot/os-lock. Grouped by the
caller's ref, the call and an unrelated run on dev would cancel each other.

* tmp: let the bump test commit when nothing really moved (#560)

* tmp: remove the bump test workflow (#560)

* Say the bump's token path is checked by the first run after merge (#560)

The secret now exists, but a bump from an unmerged branch would carry its
changes into the PR, so the first real dispatch after merge is the check.

* Record how the OS package lock was verified (#560)

* fixup: put the resolved list in the boot-image canary (#560)

A re-run after only cloud-packages.lock changed exited early and skipped the
lock check.

* fixup: run CI on lock bump PRs into hotfix/** and on lock-only PRs (#560)

A bump PR into hotfix/X.Y.Z matched no pull_request filter, so it got no lock
check and no boots. CI / Go also missed lock-only PRs, so the lock consistency
test never ran on a bump.

* fixup: keep the bot branch, split the token from the build, fall back on a rejected token (#560)

The bump rebuilt bot/os-lock from the base on every run: the no-token path
deleted it (closing a PR a person had opened) and the token path force-pushed
over human commits. Now it merges the base into the branch, adds a commit only
when the lock files changed, and never forces or deletes; a conflict stops the
run. The build job holds no secret; a separate job that builds nothing gets
only the lock files and PR text and does the push. An expired token is still a
set secret, so the push job checks it and falls back to the no-token path,
saying renew it, instead of failing red.

* fixup: document the kept bot branch, the token fallback and hotfix CI (#560)

* tmp: test the reworked OS lock bump on this branch (#560, removed before merge)

* tmp: second bump test run, with the bot branch already there (#560)

* tmp: remove the bump test workflow (#560)

* fixup: record how the reworked bump was verified (#560)

* fixup: fall back when the bot token can push but not open PRs (#560)

A token without Pull requests: write passed the probe and pushed, then the PR
call failed and the run ended with no PR, no boots and no tracking issue. Now a
failed PR call takes the no-token hand-over without pushing again, and says
which permission to add.

* tmp: test the bump's PR-call fallback (#560, removed before merge)

* tmp: remove the bump test workflow (#560)

* fixup: record the PR-call fallback test (#560)

* fixup: describe the push-but-no-PR token case on its own

Greptile: in that case the bot token pushed the branch, not GITHUB_TOKEN,
so an open bump PR's CI does start.
* docs: fix what the version split and the OS lock made stale

CLAUDE.md no longer lists apt as unwired, and names the per-line publish
inputs. BRAIN_HOST_PROTOCOL.md versioning is the minimumAgentVersion floor,
not lockstep. Up next gains the A/B OS update with #561 next.

* fixup: latest moves, and the job class is os-update everywhere

Greptile: say version tags stay fixed while latest moves forward, and rename
the apt job class in every example. host-agent self-update is a slot switch
and reboot, not an apt install.

* fixup: one outcome when jobs do not drain before the OS switch

Greptile: the self-update rule and the locked decision disagreed. Tonight's
attempt ends not applied and the next window tries again.
On the A/B layout (#561) the root is a fixed 4 GiB OS slot that holds only
the image. Docker, apps and the brain all write to the state partition,
which grows with the provider disk. Reporting / would show a 40 GB box as
about 4 GB, and its fullness is nothing the owner can act on. The hosted
reporter now measures /var/lib/moose, which is bind-mounted from the state
partition.
…firmwares (#561)

The image now carries the ESP, the BIOS boot partition and slot A (4 GiB,
read-only). systemd-repart, run in the initramfs on every boot, makes slot
B and the state partition at first boot and grows the state partition with
the disk, which replaces moose-grow-root. A local-bottom hook in Debian's
initramfs-tools runs /usr/lib/moose/state-setup from the slot: it mounts
the state partition, makes /etc an overlay on it, pins daemon.json, subuid,
subgid and login.defs on first boot, and bind-mounts the per-box dirs, each
copied once from the slot when it is first made.

GRUB boots both firmwares from one grub.cfg and grubenv (RAUC's GRUB
backend), with the kernel and initramfs inside the slot; systemd-boot is
gone. rauc and rauc-service carry the slot config for #562/#563. moose's
users and groups come from sysusers.d. Emergency and rescue reboot, and
panic=10 is on the command line, so a broken slot cannot hang.

The boot lane gives every box a 24G disk, runs every boot under UEFI and
again under legacy BIOS, checks the layout on every boot, and fails on any
failed unit or any write to the read-only slot.
The reviewed delta from CI run 36906863622: GRUB EFI for the UEFI path,
RAUC and its glib/libcurl stack for the slot config, and e2fsprogs for
the state partition. systemd-boot and its two siblings leave.
…561)

The maintainer rejected two 4 GiB ext4 slots. Each OS slot is now a 1 GiB
squashfs compressed with xz, the one compressed format GRUB 2.12 reads, so
the kernel and initramfs can stay inside the slot. The ESP is 128 MiB: the
postinst moves the kernel into the slot, so mkosi (Bootable=auto) copies no
kernel to the ESP and writes no menu entries.

dev/cloud/slotbudget fails the shipped build when the squashfs fills more
than 60% of the slot; the boot-proof image only reports, since it bakes
test-only images. /home now comes from the state partition's srv/moose/home,
and the databases keep their own bind mounts. RAUC slots are type=raw. The
lock takes rauc, grub-efi and e2fsprogs, and drops systemd-boot.
systemd-repart forces FAT32 on an ESP, and at 128 MiB mkfs.vfat's default
cluster size leaves too few clusters for FAT32. OVMF saw no filesystem on
the ESP and fell through to its shell (run 36925538434). -s1 gives about
258000 clusters.
onel and others added 25 commits October 1, 2026 23:34
…ecision (#561)

BUILD.md # 1b gets the 1 GiB squashfs-xz slots, a 128 MiB ESP, a disk budget
table with each partition's share of a 40 GB disk and the measured content
(run 36932849491), and the as-built notes and state inventory. DECISIONS.md
records the slot size change. ENVIRONMENT.md # Storage (hosted) says the
local disk is the main data store, with the databases on their own bind
mounts and /home under /srv/moose. The boot-proof doc, TESTING.md, NEXT.md,
architecture.md and the progress index follow.
When no slot was left to try, the fallback cleared the try flag of every
good slot, so a slot that failed could boot again later (Greptile P1).
It now falls back to the last good slot in ORDER and clears only that
slot's flag. A failed slot keeps TRY=1 until a new install resets it (#563).
state-setup took the first moose-state partition on any disk, so a second
disk with that label could be checked and mounted as /state (Greptile P1).
It now resolves the boot disk from the device behind / and runs repart and
the label lookup on that disk only, with no fallback. The boot lane checks
the disk it reported and that the state partition is on it.
lsblk takes PARTLABEL from the udev database, which the chroot's fresh /run
does not have, so the boot-disk lookup found no state partition (run
36940136412). The partitions now come from the boot disk's sysfs entries,
and each GPT name from blkid's low-level probe.
Hosted image in the A/B layout: 1 GiB squashfs slots, a state partition, GRUB on both firmwares (#561)
Moves the OS package lock forward (`BUILD.md` # 1b # The OS package lock). Made by `.github/workflows/os-lock-bump.yml`.

- Debian snapshot: `20261001T082322Z` to `20261002T082734Z`

**Security:** 1 of the 1 changed packages come from `trixie-security`. Release rule: cut an OS patch release with this change within 7 days (`docs/dev/contributing.md` # Release model).

| Package | From | To | Source |
|---|---|---|---|
| `libpng16-16t64` | `1.6.48-1+deb13u5` | `1.6.48-1+deb13u6` | trixie-security |

Merging this puts the change on `dev`, not on any box. It ships with the next OS release (`docs/dev/contributing.md` # Release model).
* ci: build the cloud image once, boot every group in parallel (#486)

One build job makes the production and boot-proof images and uploads
them. A matrix boots each boot group under each firmware at once, and a
publish job runs only after every boot passed. The boot-proof build
reuses the production tools tree, and runs that publish nothing use the
BuildKit gha layer cache for the brain, UI and hosted Caddy images.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs: describe the three-job cloud image workflow (#486)

Co-Authored-By: Claude <noreply@anthropic.com>

* ci: download the QEMU packages once and hand them to the boot jobs (#486)

A boot job spent 6 min fetching QEMU from a slow Ubuntu mirror. With ten
boot jobs at once the slowest one sets the wall time, so the build job
downloads the packages once and every boot job installs them locally,
with the mirror as a fallback.

Co-Authored-By: Claude <noreply@anthropic.com>

* ci: install the handed-over QEMU packages with dpkg (#486)

apt downloads a local .deb again when the mirror has the same version,
so every boot job fell back to the mirror. dpkg installs the set as is.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs: progress entry and timings for the cloud image speed-up (#486)

Docs and workflow comments only. The full list ran green on 6588081
(run 36999140120): 10.5 min, down from 27.4.

Co-Authored-By: Claude <noreply@anthropic.com>

* fixup: cache only the hosted Caddy build; read-only default permissions (#572)

The maintainer's call after the measured numbers: the brain and UI build
plain again, and only the pinned hosted Caddy build uses the gha layer
cache, still never on a run that publishes. Also from review: every
upload in build sets overwrite so a job re-run does not 409, the
workflow default is contents: read, the publish if names the boot
result, and no em dash in the copied kvm line.

Co-Authored-By: Claude <noreply@anthropic.com>

* fixup: timings for the Caddy-only cache, run 37001109373 (#572)

Docs and workflow comments only.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
OS lock (security): libpng16-16t64 1.6.48-1+deb13u5 to 1.6.48-1+deb13u6
* ci: build, check and sign a RAUC bundle per OS release (#562)

The hosted image bakes /etc/rauc/keyring.pem: the release root on a run
that publishes the OS line, a throwaway root on every other run. The build
job cuts slot A out of the image that ships, bundles it in the verity
format with the throwaway signer, and checks it against the keyring read
back out of the slot. The publish job alone re-signs it with the release
signer from the os-release environment and attaches it beside the image,
as one set of four files.

* docs: the signed OS bundle, its key custody and the maintainer's how-to (#562)

* fixup: release.yml tags no OS release without a valid release root CA (#562)

A merge that bumps VERSION before the maintainer's key setup used to tag
both lines and then fail the one cloud-image run, so the control-plane
images were never pushed either. The check now runs before anything is
tagged, and also refuses an expired or throwaway root. rauc-ca.sh removes
the files of a run that fails part way.

* fixup: sign the OS bundle in a job of its own, by digest (#562)

The signer now lives in a sign job with a plain os-release environment,
which does nothing else; the control-plane-only publish path is as before.
It signs only the bundle whose sha256 the build job reported, and publish
attaches only the bundle whose sha256 sign reported. A publishing run is
no longer cancelled by a later push.

* fixup: slotbudget -extract fails on an image cut short inside the slot (#562)

* fixup: docs for the sign job, the release rule and root replacement (#562)

* fixup: publishing runs get a concurrency group of their own (#562)

cancel-in-progress is read from the new run, so a build-only run on the
same ref could still cancel a release upload in flight.

* fixup: the boot-proof guide says publish pushes the images; BUILD.md names four jobs (#562)

* fixup: put the publish marker before the ref in the concurrency group (#562)
…he window, revert on its own (#563) (#574)

* host-agent applies OS updates: install, switch in the window, trial boot, revert (#563)

The update-target answer gains an OS list; host-agent installs the next
release into the other slot with RAUC after checking the bundle's sha256
against the answer, switches inside the window, and marks the new slot
good once the brain is healthy. A trial that fails reboots to the old
slot; an image timer reboots a slot whose host-agent never started.
Two new boots prove both paths under UEFI and BIOS.

* ci: fix the OS test bundle script's repo root (#563)

* test: keep the in-guest update source across the OS boots' reboots (#563)

* docs: the A/B OS update as built on hosted (#563)

* os update: shorter test safety net, and the revert says which path made it (#563)

* os update: review fixes (#563)

Nothing installs or switches while the booted slot is on trial, and the
box switches its OS at most once a night. A bad os part, a list that
leaves a minor out, and a downgrade without a readable floor are
refused for stream A only. A failed reboot undoes the switch; the image
timer leaves a slot already marked good alone; rauc status is retried at
start; an unreadable record is never written over; the download is
capped. A local run builds the test bundle itself.

* test: show the grubenv and the trial timer on the broken slot (#563)

* image: mask rauc-mark-good.service; host-agent alone marks a slot good (#563)

Debian's rauc-service marks every booted slot good at the end of boot,
so a slot whose host-agent never started was marked good and the box
could not fall back. The os-revert boot caught it.

* test: trace grubenv changes on the broken slot (#563)

* test: log the grubenv GRUB left for each boot (#563)

* os update: the safety net trusts the marker, not the grubenv (#563)

GRUB under UEFI does not save the booted slot's TRY flag (CI run
37045179497), so a grubenv check made the safety net never fire under
UEFI. host-agent now retries removing the marker after mark-good
instead. The rauc-mark-good mask is dropped: Debian ships no such unit.

* docs: GRUB under UEFI does not save the try flag (#575)

* docs: numbers and runs for #563

* docs: record the dismissed Greptile point (#563)

* fixup! os update: review fixes (#563)

Second review: a trial that cannot mark its slot reboots at most once,
the image timer reboots a slot at most once, a switch counts only once
activated, leftover downloads are removed, the trial timeout is capped
under the image timer, and a stale marker on a good slot is only
removed.

* fixup! docs: numbers and runs for #563

* fixup! os update: review fixes (#563)

Third round: the trial ends 13 min after boot at the latest, 2 min
before the image timer; a new switch clears the slot's old notes; a
slot that gives up its trial leaves trial-failed-<slot>, so a power cut
between mark-active and the activated flag still records the revert.

* fixup! docs: numbers and runs for #563
…re (#575) (#576)

* Give OVMF a writable VARS store so GRUB's try flag survives under UEFI (#575)

The boot lane attached the combined OVMF.fd read-only with no VARS store,
because Ubuntu 24.04 renamed the VARS template. OVMF then saved its variables
to an NvVars file on the ESP at every boot, and that lost GRUB's save_env
write of the booted slot's TRY flag.

- The harness picks OVMF as a CODE and VARS pair and stops without one.
- Every boot checks GRUB saved <booted slot>_TRY=1 and the ESP has no NvVars.
- os-revert gains a last stage: slot B, made active again, crashes its kernel
  in the initramfs (a test-only hook), and GRUB must skip it.
- moose-test-grubenv waits for the ESP mount; the A_OK check no longer races
  under pipefail.
- The trial timer stays on the marker; comments no longer say UEFI loses the
  flag.

Closes #575

* Progress entry: the full run of #575

* fixup: one OVMF pair list in dev/cloud/ovmf.sh, needed only when UEFI runs (#575)

* fixup: the NvVars check needs the ESP mounted (#575)

* fixup: the never-ending panic stage without the fix was not run; say would (#575)
…bian.13~trixie

Moves the OS package lock forward (`BUILD.md` # 1b # The OS package lock). Made by `.github/workflows/os-lock-bump.yml`.

- Debian snapshot: `20261002T082734Z` to `20261003T083111Z`
- Docker pins: moved to the newest version within each major (`dev/os-lock/third-party.lock`)

No change comes from `trixie-security`, so this ships with the next normal OS release.

| Package | From | To | Source |
|---|---|---|---|
| `docker-compose-plugin` | `5.5.1-1~debian.13~trixie` | `5.6.0-1~debian.13~trixie` |  |

Merging this puts the change on `dev`, not on any box. It ships with the next OS release (`docs/dev/contributing.md` # Release model).
OS lock: docker-compose-plugin 5.5.1-1~debian.13~trixie to 5.6.0-1~debian.13~trixie
host-agent logs the refusal just before it writes the state, so a single read
could still see 'installing'. It failed once on BIOS in run 37116684035.
os-update boot: poll os.state after a wrong digest
* Bake the last released control plane into OS-only releases

An OS-only release baked the brain and UI built from its own commit, so a
box could report the last control-plane number while running unreleased
code. Now the cloud-image build takes the brain and UI from
MOOSE_CONTROL_PLANE_SOURCE: an OS-only release pulls the last released pair
from ghcr by digest, a control-plane release and every run that publishes
nothing build them from the commit. The image records what it baked, and
every boot checks the running brain and UI against that record.

Refs #566

* Document what each release bakes (#566)

BUILD.md # Versioning now holds the as-built rule and the record, with the
release how-to, the boot-proof how-to and the architecture map to match.
The progress entry records the released run: the six original gate boots
pass with control plane 0.15.0, and the OS update boots need a control-plane
release that carries #563 first.

Refs #566

* Progress entry: the final full run (#566)

* fixup! Bake the last released control plane into OS-only releases

Review fixes: the PR trigger watches the bundle code, the local caches key
on CONTROL_PLANE_VERSION, the publish job checks the loaded images against
the recorded IDs, the hotfix steps check the released bake first, and #566
stays open until the released path is green.

Refs #566

* fixup! Document what each release bakes (#566)

The full run on the head with the review fixes.

Refs #566

* fixup! Bake the last released control plane into OS-only releases

Remove the stray root c.yml, drop a stale known gap, and tell the hotfix
reader to merge main after a control-plane release before re-checking.
#580)

* A/B OS image: a slot that hangs, and a Debian major across /etc (#486)

Point 3 (a slot that hangs): Hetzner Cloud VMs are QEMU q35 with the ICH9
TCO watchdog, which the image's generic kernel drives. Proved on real cx23
VMs: a frozen PID 1 reset the VM after 118 s. No image change; every boot of
the lane now checks systemd feeds the watchdog.

Point 5 (a Debian major): when the slot's major differs from the one the
state partition records, state-setup tidies the /etc upper layer before it
mounts it: the keep list stays, the account files are merged, the pinned
files are taken again from the slot when the remap stays, the rest moves to
an attic. A failure never stops a boot. host-agent turns sshd back on at
start when the drop-in names an account. The image's accounts become a
generated sysusers.d file, with ids checked against
dev/os-lock/cloud-accounts.lock. The os-update boot fakes a major, and every
boot fails on an upper-layer file no rule covers.

Co-Authored-By: Claude <noreply@anthropic.com>

* Progress entry: CI runs and numbers for #486 points 3 and 5

Co-Authored-By: Claude <noreply@anthropic.com>

* fixup: tidy-up never dies on os-release or the record; safe failed swap; sshd re-enabled only after a tidy-up

Review round 1 on #580.

Co-Authored-By: Claude <noreply@anthropic.com>

* Progress entry: runs after review round 1

Co-Authored-By: Claude <noreply@anthropic.com>

* fixup: tidy-up marker only when sshd was on; os-revert proves the revert direction

Greptile on #580: an admin's sshd off stays off after a major, and the
revert direction (newer record, older slot) is now crossed in os-revert.

Co-Authored-By: Claude <noreply@anthropic.com>

* Progress entry: runs after the Greptile fixes

Co-Authored-By: Claude <noreply@anthropic.com>

* fixup: a repeated swap keeps the sshd marker

---------

Co-authored-by: Claude <noreply@anthropic.com>
The root's key is offline with the maintainer. The signer it issued is in the
os-release environment. Per docs/dev/rauc-signing.md # 4.
Bump VERSION and CONTROL_PLANE_VERSION to 0.16.0
@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Critical risk] Introduces OS update mechanism and dual version streams.

No outstanding finding from this review blocks the PR from merging.

What we checked:

  • Old slot reports a false revert: When the saved switch was not activated and no trial failure was recorded, Boot removes the marker without reporting a revert.

Reviews (2) · Last reviewed commit: "os update: a failed switch undo keeps th..."

Comment thread internal/hostagent/osupdate/osupdate.go
Comment thread .github/workflows/ci-cloud-image.yml
* os update: a failed switch undo keeps the trial marker

Found by Greptile on the 0.16.0 release PR (#583). The marker is removed only
after the booted slot is first again, so a new slot that may still boot next
boots on trial.

* fixup: the test boots slot B and wants a trial
@onel
onel merged commit d9837cc into main Oct 6, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant