Skip to content

One backup reader and a spread start checkpoint by default; --start-fast when it would outlast db-timeout - #153

Open
paulocsanz wants to merge 1 commit into
mainfrom
pitr/backup-io-defaults
Open

paulocsanz wants to merge 1 commit into
mainfrom
pitr/backup-io-defaults

Conversation

@paulocsanz

Copy link
Copy Markdown
Collaborator

Backups competed with live queries for the same volume: process-max for backup was clamp(cpus/4, 1, 2) and start-fast=y forced an immediate checkpoint at every backup start. This defaults backup process-max to 1 at every vCPU size and start-fast to n, mirroring railwayapp-templates/postgres-ssl#151. Explicit PGBACKREST_BACKUP_PROCESS_MAX and PGBACKREST_START_FAST keep winning; WAL shipping, archive-get, restore workers, cadence and retention are unchanged.

The spread default needs a guard, shipped here so the HA image never carries the gap: a spread backup-start checkpoint waits about checkpoint_completion_target × checkpoint_timeout however little is dirty (measured on PG 16: 275 MB dirty, checkpoint_timeout=60s → 55 s; 0.5 s with fast), about twice that when a spread checkpoint is already running. pgBackRest bounds pg_backup_start with db-timeout (1800 s) and the stall watchdog sees no byte progress meanwhile, so a checkpoint_timeout of roughly 17 min or more would fail every backup before it copies anything. Before each backup the watcher now reads both settings and, when the worst case reaches the lower of db-timeout (PGBACKREST_DB_TIMEOUT, s/m/h suffixes accepted) and WAL_BACKUP_STALL_SECONDS, adds --start-fast and logs why. An operator-set PGBACKREST_START_FAST is left alone; a failed settings query changes nothing. Same logic as railwayapp-templates/postgres-ssl#155.

Verification: process_max_defaults is pure and asserted at 1–256 vCPU, and the rendered conf is asserted to carry start-fast=n; both tests fail with the previous defaults. start_fast_tests cover db-timeout parsing, the limit and the threshold. cargo test --locked passes (384 + 20 + 6), cargo fmt --check clean. The e2e asserts start-fast config/env precedence on the initial full and adds a cluster with checkpoint_timeout=20min that must log the --start-fast decision, then one with PGBACKREST_START_FAST=n that must not. e2e not run locally.

Roll: the image publishes on merge (postgres-patroni/** is in the build-and-push paths) and on the daily 00:00 UTC rebuild; clusters pick it up on their next deploy. Reaches every postgres-ha cluster with WAL archiving enabled. Do not merge with e2e pending.

…t; --start-fast when it would outlast db-timeout

Backups competed with live queries for the same volume: process-max for
backup was clamp(cpus/4, 1, 2) and start-fast=y forced an immediate
checkpoint at every backup start. Default backup process-max to 1 at every
vCPU size and start-fast=n, mirroring postgres-ssl #151. Explicit
PGBACKREST_BACKUP_PROCESS_MAX and PGBACKREST_START_FAST keep winning.

A spread backup-start checkpoint waits about checkpoint_completion_target x
checkpoint_timeout however little is dirty (measured on PG 16: 275 MB dirty,
checkpoint_timeout=60s -> 55 s; 0.5 s with fast), about twice that when a
spread checkpoint is already running. pgBackRest bounds pg_backup_start with
db-timeout (1800 s) and the stall watchdog sees no byte progress meanwhile,
so a checkpoint_timeout of roughly 17 min or more would fail every backup
before it copies anything. Before each backup the watcher now reads both
settings and, when the worst case reaches the lower of db-timeout
(PGBACKREST_DB_TIMEOUT, s/m/h suffixes) and WAL_BACKUP_STALL_SECONDS, adds
--start-fast and logs why. An operator-set PGBACKREST_START_FAST is left
alone; a failed settings query changes nothing. Mirrors postgres-ssl #155.

Tests: process_max_defaults is pure and asserted at 1-256 vCPU; the rendered
conf is asserted to carry start-fast=n (both fail on the previous defaults);
start_fast_tests cover db-timeout parsing, the limit and the threshold. The
e2e asserts start-fast config/env precedence on the initial full and adds a
cluster with checkpoint_timeout=20min that must log the --start-fast
decision, then one with PGBACKREST_START_FAST=n that must not.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant