Repository navigation
One backup reader and a spread start checkpoint by default; --start-fast when it would outlast db-timeout - #153
Open
paulocsanz wants to merge 1 commit into
Open
paulocsanz wants to merge 1 commit into
paulocsanz wants to merge 1 commit into
Conversation
…t; --start-fast when it would outlast db-timeout Backups competed with live queries for the same volume: process-max for backup was clamp(cpus/4, 1, 2) and start-fast=y forced an immediate checkpoint at every backup start. Default backup process-max to 1 at every vCPU size and start-fast=n, mirroring postgres-ssl #151. Explicit PGBACKREST_BACKUP_PROCESS_MAX and PGBACKREST_START_FAST keep winning. A spread backup-start checkpoint waits about checkpoint_completion_target x checkpoint_timeout however little is dirty (measured on PG 16: 275 MB dirty, checkpoint_timeout=60s -> 55 s; 0.5 s with fast), about twice that when a spread checkpoint is already running. pgBackRest bounds pg_backup_start with db-timeout (1800 s) and the stall watchdog sees no byte progress meanwhile, so a checkpoint_timeout of roughly 17 min or more would fail every backup before it copies anything. Before each backup the watcher now reads both settings and, when the worst case reaches the lower of db-timeout (PGBACKREST_DB_TIMEOUT, s/m/h suffixes) and WAL_BACKUP_STALL_SECONDS, adds --start-fast and logs why. An operator-set PGBACKREST_START_FAST is left alone; a failed settings query changes nothing. Mirrors postgres-ssl #155. Tests: process_max_defaults is pure and asserted at 1-256 vCPU; the rendered conf is asserted to carry start-fast=n (both fail on the previous defaults); start_fast_tests cover db-timeout parsing, the limit and the threshold. The e2e asserts start-fast config/env precedence on the initial full and adds a cluster with checkpoint_timeout=20min that must log the --start-fast decision, then one with PGBACKREST_START_FAST=n that must not.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Backups competed with live queries for the same volume:
process-maxfor backup wasclamp(cpus/4, 1, 2)andstart-fast=yforced an immediate checkpoint at every backup start. This defaults backupprocess-maxto 1 at every vCPU size andstart-fastton, mirroring railwayapp-templates/postgres-ssl#151. ExplicitPGBACKREST_BACKUP_PROCESS_MAXandPGBACKREST_START_FASTkeep winning; WAL shipping, archive-get, restore workers, cadence and retention are unchanged.The spread default needs a guard, shipped here so the HA image never carries the gap: a spread backup-start checkpoint waits about
checkpoint_completion_target × checkpoint_timeouthowever little is dirty (measured on PG 16: 275 MB dirty,checkpoint_timeout=60s→ 55 s; 0.5 s with fast), about twice that when a spread checkpoint is already running. pgBackRest boundspg_backup_startwithdb-timeout(1800 s) and the stall watchdog sees no byte progress meanwhile, so acheckpoint_timeoutof roughly 17 min or more would fail every backup before it copies anything. Before each backup the watcher now reads both settings and, when the worst case reaches the lower ofdb-timeout(PGBACKREST_DB_TIMEOUT, s/m/h suffixes accepted) andWAL_BACKUP_STALL_SECONDS, adds--start-fastand logs why. An operator-setPGBACKREST_START_FASTis left alone; a failed settings query changes nothing. Same logic as railwayapp-templates/postgres-ssl#155.Verification:
process_max_defaultsis pure and asserted at 1–256 vCPU, and the rendered conf is asserted to carrystart-fast=n; both tests fail with the previous defaults.start_fast_testscover db-timeout parsing, the limit and the threshold.cargo test --lockedpasses (384 + 20 + 6),cargo fmt --checkclean. The e2e asserts start-fast config/env precedence on the initial full and adds a cluster withcheckpoint_timeout=20minthat must log the--start-fastdecision, then one withPGBACKREST_START_FAST=nthat must not. e2e not run locally.Roll: the image publishes on merge (
postgres-patroni/**is in the build-and-push paths) and on the daily 00:00 UTC rebuild; clusters pick it up on their next deploy. Reaches every postgres-ha cluster with WAL archiving enabled. Do not merge with e2e pending.