Repository navigation
Release 0.16.0 - #583
Merged
Merged
Release 0.16.0#583
Conversation
Records the RAUC and systemd-sysupdate proofs from CI run 36836466471, answers the #485 questions, prices the UEFI-only hosted fallback, and recommends RAUC with GRUB on both firmwares.
Pin login.defs too (SUB_UID_COUNT 0 from #530). Say the data partition is on the OS drive, so the recovery key stays on the encrypted OS drive. The hosted root grows to the whole disk today, so state the real trade-off of a fixed 4 GiB slot, including /var outside the bind mounts.
docs: A/B update engine spike write-up (#485)
The manifest floor never holds back an OS update. Tier-2 packages are baked into the image. An OS downgrade goes back one release at most. A lock bump ships with the next VERSION release, not on merge. Mark the reboot-required detector and the pending-reboot notification as retiring.
Greptile round 2. A box that skipped a release could revert to one that cannot read its state, so formats change only in minor releases and a box steps through each minor. An OS downgrade also respects the control plane's floor. Tier-2 packages keep the run state their own spec gives them.
Greptile round 3. A box that never skips a minor must learn the newest patch of each minor on the way, so the hosted answer and the manifest's os field carry that list with each bundle digest.
Greptile round 4: say the last entry is the target, and cover the same-minor and one-minor-back cases, so a deliberate downgrade can be selected.
Greptile round 5: 2.0 follows 1.9, so stepping and the one-back downgrade work across a major boundary.
docs: design the A/B OS update and split the version lines (#486)
moose now has two version lines, one per update stream (DECISIONS.md 2026-10-01). VERSION stays the moose (OS) release and keeps stamping host-agent, so the brain's minimumAgentVersion floor still compares against the moose version with no wire change. The brain and both control-plane images belong to the new CONTROL_PLANE_VERSION line, which starts at 0.15.0 so nothing a box reports jumps. The brain's --version now reads "moose control plane X.Y.Z (g<sha>)", so its number is never read as the OS release. host-agent keeps "moose X.Y.Z (g<sha>)", which the cloud boot proof greps. Both images gain OCI version and revision labels. Part of #559.
release.yml now decides each line on its own: a VERSION bump tags vX.Y.Z and attaches the disk image, a CONTROL_PLANE_VERSION bump tags control-plane-vX.Y.Z and pushes the brain and UI images to ghcr. A merge that bumps both cuts both in one cloud-image run, and one that bumps neither stays a green no-op. The ghcr image tag keeps the vX.Y.Z shape, with the control-plane number, because the private control plane resolves digests by that tag. That shape is shared with every release before the split, so the control-plane line refuses a version whose image tag is already published rather than overwrite released images with new bytes. A control-plane-only release still runs the full boot proof and skips only the disk-image attach, so a control plane is only published after it booted. The control-plane GitHub Release is created with --latest=false so "Latest" stays on the OS release and its disk image. The decision moved into dev/release/decide.sh so it can be tested off main; decide_test.go runs it against throwaway git repos. Part of #559.
BUILD.md # Versioning drops its "designed, not built" marker and says which file stamps which binary, how each line is cut, and why the ghcr image tag keeps the vX.Y.Z shape. contributing.md # Release model now says how to cut each line, and records the one-time control-plane-v0.15.0 hand tag that the overwrite guard needs before the next release merge. DECISIONS.md 2026-10-01 gains a one-line note that control-plane-vX.Y.Z is the git tag, not the image tag. The same-commit bake of the control plane into OS images is recorded as a known gap, with #566 as the follow-up. Closes #559.
A run that tagged one line and then failed on the other tag or a GitHub Release could not be re-run: decide.sh saw its own tag and hit the bumped-to-an-already-tagged-version error. A tag at the commit being released now means "already cut here, carry on": release.yml does not tag again, creates the Release only on a definite 404, and the publish steps add only what is missing. A tag at another commit stays a hard error.
Either image already published under the control-plane version is a reason to stop, so ghcr-tag-exists.sh now takes several repositories and release.yml checks both in one call. An unclear answer from any of them, with none published, is still "unknown", never "missing".
Re-publishing only the control plane through dispatch used to re-upload the OS disk image with --clobber over a published Release. Dispatch now also takes publish_os and publish_control_plane, and keeps publish (both lines, default true) so the documented build-only command still works. The attach uploads only files the Release lacks, without --clobber, and the ghcr push skips an image whose tag is already there. A rebuild is not byte-identical, so replacing a published file is left to a person who deletes it first.
…guards DECISIONS.md 2026-10-01 still said the workflows read one VERSION above the new as-built note. It now says "Built in #559" once. BUILD.md, contributing.md and the progress entry describe the resume path, the two-image check and the per-line dispatch inputs.
A resumed attach uploaded only missing assets, so a run that died between the .raw.xz and the .sha256 paired an old image with a checksum of a new build. dev/release/attach-image.sh now leaves both, uploads neither-or-both, and refuses a half pair with the delete-asset command a person runs first. Nothing published is deleted. A resumed ghcr push skipped an image whose version tag existed, so a push that died before latest left latest stale, and brain and ui could disagree. For such an image, dev/release/ghcr-retag.sh copies the version tag's manifest to latest byte for byte through the registry API and checks the digest, with no rebuild or re-push, and only when this run publishes that version. Both scripts are tested with stubs.
Re-running an older tagged control-plane release passed the commit guard, and the latest repair then pointed brain and ui latest back at older images. A dispatch of an old version could do the same through the normal push. Both paths now move latest only when this version is the newest control-plane-v* tag by semver (dev/release/is-newest.sh). An unreadable tag list leaves latest alone with a warning and still publishes the version tag.
Control plane gets its own version line; moose vX.Y.Z is the OS release
* tmp: probe mkosi Snapshot= for trixie in CI (#560, removed before the PR) * tmp: remove the snapshot probe workflow (#560) * Add the OS package lock: Debian snapshot, Docker pins and the oslock tool (#560) With no apt on the box, every package version is chosen at build time, so the build must pin them. dev/os-lock/ holds the snapshot.debian.org timestamp, the exact Docker versions (Docker's repo is not in the snapshot), and the shared staging that turns them into apt sources and pins. oslock normalizes mkosi's manifest, checks it against the committed list, and names changes for the bump PR. cloud-packages.lock itself comes from the first CI build. * Build the hosted images from the OS package lock (#560) Both the lean image and the boot-proof image now install from the locked snapshot and pins and must resolve to dev/os-lock/cloud-packages.lock. Two separate builds of one commit checking the same list is the proof that the lock holds. trixie-updates is added back from the snapshot, because mkosi drops it in snapshot mode. * Add the daily OS lock bump and run full boots on lock PRs (#560) The bump opens or updates one PR when OS_LOCK_BOT_TOKEN exists, because a PR made with GITHUB_TOKEN can neither be opened here (#484) nor start CI. Without the token it pushes the branch, boots it by calling ci-cloud-image.yml with the new ref input, and leaves a summary and a tracking issue with an open-a-PR link. A PR touching dev/os-lock/ boots the full list, and the mkosi setup is shared so the bump resolves exactly what CI checks. * Commit the resolved package list of the hosted image (#560) Taken from the first CI build at the locked snapshot and pins (run 36873363298, artifact cloud-packages-lock). From now on both cloud builds must resolve to exactly this list. * Document the OS package lock and the OS patch release process (#560) BUILD.md # 1b gets the as-built lock. The maintainer decided how a lock-only release is cut (from main as hotfix/X.Y.Z, within 7 days for a trixie-security change), so contributing.md, DECISIONS.md and NEXT.md point 6 record it, along with how the bump PR is opened and what the bot token needs. * tmp: run the OS lock bump on this branch (#560, removed before the PR) * Group a called cloud-image run by the ref it builds (#560) The bump calls ci-cloud-image.yml from dev with ref=bot/os-lock. Grouped by the caller's ref, the call and an unrelated run on dev would cancel each other. * tmp: let the bump test commit when nothing really moved (#560) * tmp: remove the bump test workflow (#560) * Say the bump's token path is checked by the first run after merge (#560) The secret now exists, but a bump from an unmerged branch would carry its changes into the PR, so the first real dispatch after merge is the check. * Record how the OS package lock was verified (#560) * fixup: put the resolved list in the boot-image canary (#560) A re-run after only cloud-packages.lock changed exited early and skipped the lock check. * fixup: run CI on lock bump PRs into hotfix/** and on lock-only PRs (#560) A bump PR into hotfix/X.Y.Z matched no pull_request filter, so it got no lock check and no boots. CI / Go also missed lock-only PRs, so the lock consistency test never ran on a bump. * fixup: keep the bot branch, split the token from the build, fall back on a rejected token (#560) The bump rebuilt bot/os-lock from the base on every run: the no-token path deleted it (closing a PR a person had opened) and the token path force-pushed over human commits. Now it merges the base into the branch, adds a commit only when the lock files changed, and never forces or deletes; a conflict stops the run. The build job holds no secret; a separate job that builds nothing gets only the lock files and PR text and does the push. An expired token is still a set secret, so the push job checks it and falls back to the no-token path, saying renew it, instead of failing red. * fixup: document the kept bot branch, the token fallback and hotfix CI (#560) * tmp: test the reworked OS lock bump on this branch (#560, removed before merge) * tmp: second bump test run, with the bot branch already there (#560) * tmp: remove the bump test workflow (#560) * fixup: record how the reworked bump was verified (#560) * fixup: fall back when the bot token can push but not open PRs (#560) A token without Pull requests: write passed the probe and pushed, then the PR call failed and the run ended with no PR, no boots and no tracking issue. Now a failed PR call takes the no-token hand-over without pushing again, and says which permission to add. * tmp: test the bump's PR-call fallback (#560, removed before merge) * tmp: remove the bump test workflow (#560) * fixup: record the PR-call fallback test (#560) * fixup: describe the push-but-no-PR token case on its own Greptile: in that case the bot token pushed the branch, not GITHUB_TOKEN, so an open bump PR's CI does start.
* docs: fix what the version split and the OS lock made stale CLAUDE.md no longer lists apt as unwired, and names the per-line publish inputs. BRAIN_HOST_PROTOCOL.md versioning is the minimumAgentVersion floor, not lockstep. Up next gains the A/B OS update with #561 next. * fixup: latest moves, and the job class is os-update everywhere Greptile: say version tags stay fixed while latest moves forward, and rename the apt job class in every example. host-agent self-update is a slot switch and reboot, not an apt install. * fixup: one outcome when jobs do not drain before the OS switch Greptile: the self-update rule and the locked decision disagreed. Tonight's attempt ends not applied and the next window tries again.
On the A/B layout (#561) the root is a fixed 4 GiB OS slot that holds only the image. Docker, apps and the brain all write to the state partition, which grows with the provider disk. Reporting / would show a 40 GB box as about 4 GB, and its fullness is nothing the owner can act on. The hosted reporter now measures /var/lib/moose, which is bind-mounted from the state partition.
…firmwares (#561) The image now carries the ESP, the BIOS boot partition and slot A (4 GiB, read-only). systemd-repart, run in the initramfs on every boot, makes slot B and the state partition at first boot and grows the state partition with the disk, which replaces moose-grow-root. A local-bottom hook in Debian's initramfs-tools runs /usr/lib/moose/state-setup from the slot: it mounts the state partition, makes /etc an overlay on it, pins daemon.json, subuid, subgid and login.defs on first boot, and bind-mounts the per-box dirs, each copied once from the slot when it is first made. GRUB boots both firmwares from one grub.cfg and grubenv (RAUC's GRUB backend), with the kernel and initramfs inside the slot; systemd-boot is gone. rauc and rauc-service carry the slot config for #562/#563. moose's users and groups come from sysusers.d. Emergency and rescue reboot, and panic=10 is on the command line, so a broken slot cannot hang. The boot lane gives every box a 24G disk, runs every boot under UEFI and again under legacy BIOS, checks the layout on every boot, and fails on any failed unit or any write to the read-only slot.
The reviewed delta from CI run 36906863622: GRUB EFI for the UEFI path, RAUC and its glib/libcurl stack for the slot config, and e2fsprogs for the state partition. systemd-boot and its two siblings leave.
…561) The maintainer rejected two 4 GiB ext4 slots. Each OS slot is now a 1 GiB squashfs compressed with xz, the one compressed format GRUB 2.12 reads, so the kernel and initramfs can stay inside the slot. The ESP is 128 MiB: the postinst moves the kernel into the slot, so mkosi (Bootable=auto) copies no kernel to the ESP and writes no menu entries. dev/cloud/slotbudget fails the shipped build when the squashfs fills more than 60% of the slot; the boot-proof image only reports, since it bakes test-only images. /home now comes from the state partition's srv/moose/home, and the databases keep their own bind mounts. RAUC slots are type=raw. The lock takes rauc, grub-efi and e2fsprogs, and drops systemd-boot.
systemd-repart forces FAT32 on an ESP, and at 128 MiB mkfs.vfat's default cluster size leaves too few clusters for FAT32. OVMF saw no filesystem on the ESP and fell through to its shell (run 36925538434). -s1 gives about 258000 clusters.
…ecision (#561) BUILD.md # 1b gets the 1 GiB squashfs-xz slots, a 128 MiB ESP, a disk budget table with each partition's share of a 40 GB disk and the measured content (run 36932849491), and the as-built notes and state inventory. DECISIONS.md records the slot size change. ENVIRONMENT.md # Storage (hosted) says the local disk is the main data store, with the databases on their own bind mounts and /home under /srv/moose. The boot-proof doc, TESTING.md, NEXT.md, architecture.md and the progress index follow.
When no slot was left to try, the fallback cleared the try flag of every good slot, so a slot that failed could boot again later (Greptile P1). It now falls back to the last good slot in ORDER and clears only that slot's flag. A failed slot keeps TRY=1 until a new install resets it (#563).
state-setup took the first moose-state partition on any disk, so a second disk with that label could be checked and mounted as /state (Greptile P1). It now resolves the boot disk from the device behind / and runs repart and the label lookup on that disk only, with no fallback. The boot lane checks the disk it reported and that the state partition is on it.
lsblk takes PARTLABEL from the udev database, which the chroot's fresh /run does not have, so the boot-disk lookup found no state partition (run 36940136412). The partitions now come from the boot disk's sysfs entries, and each GPT name from blkid's low-level probe.
Hosted image in the A/B layout: 1 GiB squashfs slots, a state partition, GRUB on both firmwares (#561)
Moves the OS package lock forward (`BUILD.md` # 1b # The OS package lock). Made by `.github/workflows/os-lock-bump.yml`. - Debian snapshot: `20261001T082322Z` to `20261002T082734Z` **Security:** 1 of the 1 changed packages come from `trixie-security`. Release rule: cut an OS patch release with this change within 7 days (`docs/dev/contributing.md` # Release model). | Package | From | To | Source | |---|---|---|---| | `libpng16-16t64` | `1.6.48-1+deb13u5` | `1.6.48-1+deb13u6` | trixie-security | Merging this puts the change on `dev`, not on any box. It ships with the next OS release (`docs/dev/contributing.md` # Release model).
* ci: build the cloud image once, boot every group in parallel (#486) One build job makes the production and boot-proof images and uploads them. A matrix boots each boot group under each firmware at once, and a publish job runs only after every boot passed. The boot-proof build reuses the production tools tree, and runs that publish nothing use the BuildKit gha layer cache for the brain, UI and hosted Caddy images. Co-Authored-By: Claude <noreply@anthropic.com> * docs: describe the three-job cloud image workflow (#486) Co-Authored-By: Claude <noreply@anthropic.com> * ci: download the QEMU packages once and hand them to the boot jobs (#486) A boot job spent 6 min fetching QEMU from a slow Ubuntu mirror. With ten boot jobs at once the slowest one sets the wall time, so the build job downloads the packages once and every boot job installs them locally, with the mirror as a fallback. Co-Authored-By: Claude <noreply@anthropic.com> * ci: install the handed-over QEMU packages with dpkg (#486) apt downloads a local .deb again when the mirror has the same version, so every boot job fell back to the mirror. dpkg installs the set as is. Co-Authored-By: Claude <noreply@anthropic.com> * docs: progress entry and timings for the cloud image speed-up (#486) Docs and workflow comments only. The full list ran green on 6588081 (run 36999140120): 10.5 min, down from 27.4. Co-Authored-By: Claude <noreply@anthropic.com> * fixup: cache only the hosted Caddy build; read-only default permissions (#572) The maintainer's call after the measured numbers: the brain and UI build plain again, and only the pinned hosted Caddy build uses the gha layer cache, still never on a run that publishes. Also from review: every upload in build sets overwrite so a job re-run does not 409, the workflow default is contents: read, the publish if names the boot result, and no em dash in the copied kvm line. Co-Authored-By: Claude <noreply@anthropic.com> * fixup: timings for the Caddy-only cache, run 37001109373 (#572) Docs and workflow comments only. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com>
OS lock (security): libpng16-16t64 1.6.48-1+deb13u5 to 1.6.48-1+deb13u6
* ci: build, check and sign a RAUC bundle per OS release (#562) The hosted image bakes /etc/rauc/keyring.pem: the release root on a run that publishes the OS line, a throwaway root on every other run. The build job cuts slot A out of the image that ships, bundles it in the verity format with the throwaway signer, and checks it against the keyring read back out of the slot. The publish job alone re-signs it with the release signer from the os-release environment and attaches it beside the image, as one set of four files. * docs: the signed OS bundle, its key custody and the maintainer's how-to (#562) * fixup: release.yml tags no OS release without a valid release root CA (#562) A merge that bumps VERSION before the maintainer's key setup used to tag both lines and then fail the one cloud-image run, so the control-plane images were never pushed either. The check now runs before anything is tagged, and also refuses an expired or throwaway root. rauc-ca.sh removes the files of a run that fails part way. * fixup: sign the OS bundle in a job of its own, by digest (#562) The signer now lives in a sign job with a plain os-release environment, which does nothing else; the control-plane-only publish path is as before. It signs only the bundle whose sha256 the build job reported, and publish attaches only the bundle whose sha256 sign reported. A publishing run is no longer cancelled by a later push. * fixup: slotbudget -extract fails on an image cut short inside the slot (#562) * fixup: docs for the sign job, the release rule and root replacement (#562) * fixup: publishing runs get a concurrency group of their own (#562) cancel-in-progress is read from the new run, so a build-only run on the same ref could still cancel a release upload in flight. * fixup: the boot-proof guide says publish pushes the images; BUILD.md names four jobs (#562) * fixup: put the publish marker before the ref in the concurrency group (#562)
…he window, revert on its own (#563) (#574) * host-agent applies OS updates: install, switch in the window, trial boot, revert (#563) The update-target answer gains an OS list; host-agent installs the next release into the other slot with RAUC after checking the bundle's sha256 against the answer, switches inside the window, and marks the new slot good once the brain is healthy. A trial that fails reboots to the old slot; an image timer reboots a slot whose host-agent never started. Two new boots prove both paths under UEFI and BIOS. * ci: fix the OS test bundle script's repo root (#563) * test: keep the in-guest update source across the OS boots' reboots (#563) * docs: the A/B OS update as built on hosted (#563) * os update: shorter test safety net, and the revert says which path made it (#563) * os update: review fixes (#563) Nothing installs or switches while the booted slot is on trial, and the box switches its OS at most once a night. A bad os part, a list that leaves a minor out, and a downgrade without a readable floor are refused for stream A only. A failed reboot undoes the switch; the image timer leaves a slot already marked good alone; rauc status is retried at start; an unreadable record is never written over; the download is capped. A local run builds the test bundle itself. * test: show the grubenv and the trial timer on the broken slot (#563) * image: mask rauc-mark-good.service; host-agent alone marks a slot good (#563) Debian's rauc-service marks every booted slot good at the end of boot, so a slot whose host-agent never started was marked good and the box could not fall back. The os-revert boot caught it. * test: trace grubenv changes on the broken slot (#563) * test: log the grubenv GRUB left for each boot (#563) * os update: the safety net trusts the marker, not the grubenv (#563) GRUB under UEFI does not save the booted slot's TRY flag (CI run 37045179497), so a grubenv check made the safety net never fire under UEFI. host-agent now retries removing the marker after mark-good instead. The rauc-mark-good mask is dropped: Debian ships no such unit. * docs: GRUB under UEFI does not save the try flag (#575) * docs: numbers and runs for #563 * docs: record the dismissed Greptile point (#563) * fixup! os update: review fixes (#563) Second review: a trial that cannot mark its slot reboots at most once, the image timer reboots a slot at most once, a switch counts only once activated, leftover downloads are removed, the trial timeout is capped under the image timer, and a stale marker on a good slot is only removed. * fixup! docs: numbers and runs for #563 * fixup! os update: review fixes (#563) Third round: the trial ends 13 min after boot at the latest, 2 min before the image timer; a new switch clears the slot's old notes; a slot that gives up its trial leaves trial-failed-<slot>, so a power cut between mark-active and the activated flag still records the revert. * fixup! docs: numbers and runs for #563
…re (#575) (#576) * Give OVMF a writable VARS store so GRUB's try flag survives under UEFI (#575) The boot lane attached the combined OVMF.fd read-only with no VARS store, because Ubuntu 24.04 renamed the VARS template. OVMF then saved its variables to an NvVars file on the ESP at every boot, and that lost GRUB's save_env write of the booted slot's TRY flag. - The harness picks OVMF as a CODE and VARS pair and stops without one. - Every boot checks GRUB saved <booted slot>_TRY=1 and the ESP has no NvVars. - os-revert gains a last stage: slot B, made active again, crashes its kernel in the initramfs (a test-only hook), and GRUB must skip it. - moose-test-grubenv waits for the ESP mount; the A_OK check no longer races under pipefail. - The trial timer stays on the marker; comments no longer say UEFI loses the flag. Closes #575 * Progress entry: the full run of #575 * fixup: one OVMF pair list in dev/cloud/ovmf.sh, needed only when UEFI runs (#575) * fixup: the NvVars check needs the ESP mounted (#575) * fixup: the never-ending panic stage without the fix was not run; say would (#575)
…bian.13~trixie Moves the OS package lock forward (`BUILD.md` # 1b # The OS package lock). Made by `.github/workflows/os-lock-bump.yml`. - Debian snapshot: `20261002T082734Z` to `20261003T083111Z` - Docker pins: moved to the newest version within each major (`dev/os-lock/third-party.lock`) No change comes from `trixie-security`, so this ships with the next normal OS release. | Package | From | To | Source | |---|---|---|---| | `docker-compose-plugin` | `5.5.1-1~debian.13~trixie` | `5.6.0-1~debian.13~trixie` | | Merging this puts the change on `dev`, not on any box. It ships with the next OS release (`docs/dev/contributing.md` # Release model).
OS lock: docker-compose-plugin 5.5.1-1~debian.13~trixie to 5.6.0-1~debian.13~trixie
host-agent logs the refusal just before it writes the state, so a single read could still see 'installing'. It failed once on BIOS in run 37116684035.
os-update boot: poll os.state after a wrong digest
* Bake the last released control plane into OS-only releases An OS-only release baked the brain and UI built from its own commit, so a box could report the last control-plane number while running unreleased code. Now the cloud-image build takes the brain and UI from MOOSE_CONTROL_PLANE_SOURCE: an OS-only release pulls the last released pair from ghcr by digest, a control-plane release and every run that publishes nothing build them from the commit. The image records what it baked, and every boot checks the running brain and UI against that record. Refs #566 * Document what each release bakes (#566) BUILD.md # Versioning now holds the as-built rule and the record, with the release how-to, the boot-proof how-to and the architecture map to match. The progress entry records the released run: the six original gate boots pass with control plane 0.15.0, and the OS update boots need a control-plane release that carries #563 first. Refs #566 * Progress entry: the final full run (#566) * fixup! Bake the last released control plane into OS-only releases Review fixes: the PR trigger watches the bundle code, the local caches key on CONTROL_PLANE_VERSION, the publish job checks the loaded images against the recorded IDs, the hotfix steps check the released bake first, and #566 stays open until the released path is green. Refs #566 * fixup! Document what each release bakes (#566) The full run on the head with the review fixes. Refs #566 * fixup! Bake the last released control plane into OS-only releases Remove the stray root c.yml, drop a stale known gap, and tell the hotfix reader to merge main after a control-plane release before re-checking.
#580) * A/B OS image: a slot that hangs, and a Debian major across /etc (#486) Point 3 (a slot that hangs): Hetzner Cloud VMs are QEMU q35 with the ICH9 TCO watchdog, which the image's generic kernel drives. Proved on real cx23 VMs: a frozen PID 1 reset the VM after 118 s. No image change; every boot of the lane now checks systemd feeds the watchdog. Point 5 (a Debian major): when the slot's major differs from the one the state partition records, state-setup tidies the /etc upper layer before it mounts it: the keep list stays, the account files are merged, the pinned files are taken again from the slot when the remap stays, the rest moves to an attic. A failure never stops a boot. host-agent turns sshd back on at start when the drop-in names an account. The image's accounts become a generated sysusers.d file, with ids checked against dev/os-lock/cloud-accounts.lock. The os-update boot fakes a major, and every boot fails on an upper-layer file no rule covers. Co-Authored-By: Claude <noreply@anthropic.com> * Progress entry: CI runs and numbers for #486 points 3 and 5 Co-Authored-By: Claude <noreply@anthropic.com> * fixup: tidy-up never dies on os-release or the record; safe failed swap; sshd re-enabled only after a tidy-up Review round 1 on #580. Co-Authored-By: Claude <noreply@anthropic.com> * Progress entry: runs after review round 1 Co-Authored-By: Claude <noreply@anthropic.com> * fixup: tidy-up marker only when sshd was on; os-revert proves the revert direction Greptile on #580: an admin's sshd off stays off after a major, and the revert direction (newer record, older slot) is now crossed in os-revert. Co-Authored-By: Claude <noreply@anthropic.com> * Progress entry: runs after the Greptile fixes Co-Authored-By: Claude <noreply@anthropic.com> * fixup: a repeated swap keeps the sshd marker --------- Co-authored-by: Claude <noreply@anthropic.com>
The root's key is offline with the maintainer. The signer it issued is in the os-release environment. Per docs/dev/rauc-signing.md # 4.
Commit the RAUC release root CA
Bump VERSION and CONTROL_PLANE_VERSION to 0.16.0
|
* os update: a failed switch undo keeps the trial marker Found by Greptile on the 0.16.0 release PR (#583). The marker is removed only after the booted slot is first again, so a new slot that may still boot next boots on trial. * fixup: the test boots slot B and wants a trial
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 0.16.0, on both lines: the OS (
VERSION0.16.0) and the control plane (CONTROL_PLANE_VERSION0.16.0), bumped in #582. The headline: the hosted image is now an A/B image, and a box can update its own OS from a signed bundle, and go back on its own when the new OS fails. This release also carries a Debian security fix (libpng, #571).What is new
/etcis an overlay over the slot, so a box keeps its users, keys and settings across OS updates.moose-v0.16.0-amd64.raucband its sha256 sit next to the disk image on this Release. The bundle is signed by a CI signer issued from an offline root CA, and a box only installs a bundle that chains to that root./etc. Moose's own files are kept, and any other hand edit moves to an attic on the state partition. This works for an upgrade and for a revert (Image-based A/B OS updates (stream A) with a persistent data partition, for both profiles #486).control-plane-v0.16.0. An OS-only release bakes the last released control plane (OS releases bake the last released control plane, pulled from ghcr by digest #566).Testing
CI / Cloud imagebuilds once and boots in parallel (ci: build the cloud image once, boot in parallel (#486) #572). A full run takes about 15 minutes. New boots cover an OS update, a revert after a failed trial, a kernel panic on the new slot, and a Debian-major tidy-up in both directions.Upgrading
v0.16.0andcontrol-plane-v0.16.0, creates both Releases, runs the full boot proof, signs the bundle, attaches the image and the bundle, and pushes the brain and UI images to ghcr.