Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions charts/paradedb/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -213,6 +213,12 @@ Alternatively, you can manually import the dashboard from the `monitoring` direc
### Metrics Configuration

Additionally, we recommend enabling the `kube-state-metrics` CRD monitoring and adding the CNPG metrics. The configuration can be found in `monitoring/metrics-clusters_postgresql_cnpg_io.yaml`.
Enable list/watch permissions for both `clusters` and `scheduledbackups` in the
`postgresql.cnpg.io` API group. Backup alerts use Cluster backup-status timestamps
and ScheduledBackup's `nextScheduleTime`; apply the updated metrics configuration
alongside the chart upgrade. A stale-backup alert requires a successful-backup
timestamp; a cluster that has never produced a backup and has no reported failure
is not covered by that rule.

## Examples

Expand Down
5 changes: 4 additions & 1 deletion charts/paradedb/docs/runbooks/CNPGBackupStale.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,10 @@ The `CNPGBackupStale` alert is triggered when a CloudNativePG cluster's most rec

Backups are scheduled nightly, so 26 hours is one missed run plus two hours of grace. The alert does not distinguish why the backup is old: it fires whether backups have been failing, whether the ScheduledBackup stopped being reconciled, or whether backups were switched off and nobody noticed.

This is the only backup alert that fires when backups stop happening silently. `CNPGBackupFailed` needs a failure to report, and a backup that is never attempted never fails.
This alert uses the Cluster's `lastSuccessfulBackupByMethod` status exported by
kube-state-metrics, including plugin backups. It requires a previous successful
backup timestamp. `CNPGBackupFailed` reports failures and `CNPGScheduledBackupStalled`
reports a schedule that stopped advancing; neither replaces this recovery-age check.

## Impact

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -191,3 +191,19 @@ spec:
backup_capabilities: [backupCapabilities]
restore_job_hook_capabilities: [restoreJobHookCapabilities]
status: [status]

- groupVersionKind:
group: postgresql.cnpg.io
version: "v1"
kind: "ScheduledBackup"
labelsFromPath:
namespace: [metadata, namespace]
cluster: [spec, cluster, name]
scheduled_backup: [metadata, name]
metrics:
- name: "next_schedule_time"
help: "Timestamp of the next scheduled backup"
each:
type: Gauge
gauge:
path: [status, nextScheduleTime]
21 changes: 21 additions & 0 deletions charts/paradedb/prometheus_rules/cluster-backup_failed.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
{{- $alert := "CNPGBackupFailed" -}}
{{- if not (has $alert .excludeRules) -}}
alert: {{ $alert }}
annotations:
summary: ParadeDB CNPG Cluster's most recent backup failed
description: |-
CloudNativePG Cluster "{{ .namespace }}/{{ .cluster }}" has a failed backup newer than its last successful backup, or has never completed a successful backup. Check the Backup objects and object-store credentials.
runbook_url: https://github.com/paradedb/charts/blob/main/charts/paradedb/docs/runbooks/CNPGBackupFailed.md
expr: |
(max(kube_customresource_last_failed_backup{customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",namespace="{{ .namespace }}",cluster="{{ .cluster }}"}) unless max(kube_customresource_last_successful_backup{customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",namespace="{{ .namespace }}",cluster="{{ .cluster }}"}))
or
(max(kube_customresource_last_failed_backup{customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",namespace="{{ .namespace }}",cluster="{{ .cluster }}"}) > max(kube_customresource_last_successful_backup{customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",namespace="{{ .namespace }}",cluster="{{ .cluster }}"}))
for: 15m
labels:
severity: warning
namespace: {{ .namespace }}
cnpg_cluster: {{ .cluster }}
{{- range $key, $val := .additionalLabels }}
{{ $key }}: {{ $val | toString | quote }}
{{- end }}
{{- end -}}
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
{{- $alert := "CNPGScheduledBackupStalled" -}}
{{- if not (has $alert .excludeRules) -}}
alert: {{ $alert }}
annotations:
summary: ParadeDB CNPG ScheduledBackup is not advancing its schedule
description: |-
CloudNativePG Cluster "{{ .namespace }}/{{ .cluster }}" scheduled backup "{{ .labels.scheduled_backup }}" was due {{ .value }} seconds ago and has not been rescheduled. Check operator reconciliation.
runbook_url: https://github.com/paradedb/charts/blob/main/charts/paradedb/docs/runbooks/CNPGScheduledBackupStalled.md
expr: |
time() - max by (namespace, cluster, scheduled_backup) (
kube_customresource_next_schedule_time{customresource_group="postgresql.cnpg.io",customresource_kind="ScheduledBackup",namespace="{{ .namespace }}",cluster="{{ .cluster }}"}
) > 3600
for: 15m
labels:
severity: warning
namespace: {{ .namespace }}
cnpg_cluster: {{ .cluster }}
{{- range $key, $val := .additionalLabels }}
{{ $key }}: {{ $val | toString | quote }}
{{- end }}
{{- end -}}
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ annotations:
This fires on the outcome regardless of cause, so it covers backups that are failing, a schedule that has stopped being reconciled, and backups that were switched off and forgotten.
runbook_url: https://github.com/paradedb/charts/blob/main/charts/paradedb/docs/runbooks/CNPGBackupStale.md
expr: |
time() - max(cnpg_collector_last_available_backup_timestamp{namespace="{{ .namespace }}",pod=~"{{ .podSelector }}"}) > 26 * 3600
time() - max(kube_customresource_last_successful_backup{customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",namespace="{{ .namespace }}",cluster="{{ .cluster }}"}) > 26 * 3600
for: 15m
labels:
severity: critical
Expand Down
2 changes: 1 addition & 1 deletion charts/paradedb/templates/prometheus-rule.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ spec:
{{- $_ := set $dict "valueHuman" "{{ $value | humanize }}" -}}
{{- $_ := set $dict "valuePercent" "{{ $value | humanizePercentage }}" -}}
{{- $_ := set $dict "cluster" (include "cluster.fullname" .) -}}
{{- $_ := set $dict "labels" (dict "job" "{{ $labels.job }}" "node" "{{ $labels.node }}" "pod" "{{ $labels.pod }}" "deployment" "{{ $labels.deployment }}" "subname" "{{ $labels.subname }}" "lag_type" "{{ $labels.lag_type }}" "stop_reason" "{{ $labels.stop_reason }}" "slot_name" "{{ $labels.slot_name }}" "slot_type" "{{ $labels.slot_type }}" "persistentvolumeclaim" "{{ $labels.persistentvolumeclaim }}" "datname" "{{ $labels.datname }}" "age_kind" "{{ $labels.age_kind }}" "current_primary" "{{ $labels.current_primary }}" "target_primary" "{{ $labels.target_primary }}") -}}
{{- $_ := set $dict "labels" (dict "scheduled_backup" "{{ $labels.scheduled_backup }}" "job" "{{ $labels.job }}" "node" "{{ $labels.node }}" "pod" "{{ $labels.pod }}" "deployment" "{{ $labels.deployment }}" "subname" "{{ $labels.subname }}" "lag_type" "{{ $labels.lag_type }}" "stop_reason" "{{ $labels.stop_reason }}" "slot_name" "{{ $labels.slot_name }}" "slot_type" "{{ $labels.slot_type }}" "persistentvolumeclaim" "{{ $labels.persistentvolumeclaim }}" "datname" "{{ $labels.datname }}" "age_kind" "{{ $labels.age_kind }}" "current_primary" "{{ $labels.current_primary }}" "target_primary" "{{ $labels.target_primary }}") -}}
{{- $_ := set $dict "podSelector" (printf "%s-([1-9][0-9]*)$" (include "cluster.fullname" .)) -}}
{{- $_ := set $dict "pvcSelector" (printf "%s-([1-9][0-9]*)" (include "cluster.fullname" .)) -}}
{{- $_ := set $dict "Values" .Values -}}
Expand Down
27 changes: 27 additions & 0 deletions charts/paradedb/test/alert-parity-backups/chainsaw-test.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
apiVersion: chainsaw.kyverno.io/v1alpha1
kind: Test
metadata:
name: alert-parity-backups
spec:
timeouts:
exec: 5m
steps:
- name: Evaluate rendered alert rules
try:
- script:
content: |
set -eu
test_dir="$(mktemp -d)"
trap 'rm -rf "$test_dir"' EXIT
helm template parity ../../ --namespace alert-test \
--set cluster.monitoring.enabled=true \
--set cluster.monitoring.prometheusRule.enabled=true \
--show-only templates/prometheus-rule.yaml \
| sed -n '/^spec:/,$p' | tail -n +2 > "$test_dir/rules.generated.yaml"
cp rules.test.yaml "$test_dir/"
if [ -z "${VMALERT_TOOL:-}" ]; then
curl -fsSL https://github.com/VictoriaMetrics/VictoriaMetrics/releases/download/v1.148.0/vmutils-linux-amd64-v1.148.0.tar.gz \
| tar xz -C "$test_dir" vmalert-tool-prod
VMALERT_TOOL="$test_dir/vmalert-tool-prod"
fi
(cd "$test_dir" && "$VMALERT_TOOL" unittest --disableAlertgroupLabel --files=rules.test.yaml)
128 changes: 128 additions & 0 deletions charts/paradedb/test/alert-parity-backups/rules.test.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,128 @@
rule_files:
- rules.generated.yaml
evaluation_interval: 1m
tests:
- name: last successful=300, failed=600
interval: 1m
input_series:
- series: kube_customresource_last_failed_backup{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",cluster="parity-paradedb"}
values: 946685400+0x40
- series: kube_customresource_last_successful_backup{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",cluster="parity-paradedb",method="plugin"}
values: 946685100+0x40
alert_rule_test:
- eval_time: 20m
alertname: CNPGBackupFailed
exp_alerts:
- exp_labels:
severity: warning
namespace: alert-test
cnpg_cluster: parity-paradedb
exp_annotations:
summary: ParadeDB CNPG Cluster's most recent backup failed
description: CloudNativePG Cluster "alert-test/parity-paradedb" has a failed backup newer than its last successful backup, or has never completed
a successful backup. Check the Backup objects and object-store credentials.
runbook_url: https://github.com/paradedb/charts/blob/main/charts/paradedb/docs/runbooks/CNPGBackupFailed.md
- name: last successful=900, failed=600
interval: 1m
input_series:
- series: kube_customresource_last_failed_backup{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",cluster="parity-paradedb"}
values: 946685400+0x40
- series: kube_customresource_last_successful_backup{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",cluster="parity-paradedb",method="plugin"}
values: 946685700+0x40
alert_rule_test:
- eval_time: 20m
alertname: CNPGBackupFailed
exp_alerts: []
- name: last successful=None, failed=600
interval: 1m
input_series:
- series: kube_customresource_last_failed_backup{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",cluster="parity-paradedb"}
values: 946685400+0x40
alert_rule_test:
- eval_time: 20m
alertname: CNPGBackupFailed
exp_alerts:
- exp_labels:
severity: warning
namespace: alert-test
cnpg_cluster: parity-paradedb
exp_annotations:
summary: ParadeDB CNPG Cluster's most recent backup failed
description: CloudNativePG Cluster "alert-test/parity-paradedb" has a failed backup newer than its last successful backup, or has never completed
a successful backup. Check the Backup objects and object-store credentials.
runbook_url: https://github.com/paradedb/charts/blob/main/charts/paradedb/docs/runbooks/CNPGBackupFailed.md
- name: stale plugin backup without collector backup metrics
interval: 1m
input_series:
- series: kube_customresource_last_successful_backup{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",cluster="parity-paradedb",method="plugin"}
values: 946584800+0x40
alert_rule_test:
- eval_time: 20m
alertname: CNPGBackupStale
exp_alerts:
- exp_labels:
severity: critical
namespace: alert-test
cnpg_cluster: parity-paradedb
exp_annotations:
summary: ParadeDB CNPG Cluster has no recent successful backup
description: |-
CloudNativePG Cluster "alert-test/parity-paradedb" last completed a successful backup 101200 seconds ago.

Backups are expected nightly, so 26 hours is one missed run plus two hours of grace. Everything written since that backup is outside the recovery window.

This fires on the outcome regardless of cause, so it covers backups that are failing, a schedule that has stopped being reconciled, and backups that were switched off and forgotten.
runbook_url: https://github.com/paradedb/charts/blob/main/charts/paradedb/docs/runbooks/CNPGBackupStale.md
- name: recent backup through any method
interval: 1m
input_series:
- series: kube_customresource_last_successful_backup{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",cluster="parity-paradedb",method="plugin"}
values: 946685400+0x40
- series: kube_customresource_last_successful_backup{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="Cluster",cluster="parity-paradedb",method="volumeSnapshot"}
values: 946584800+0x40
alert_rule_test:
- eval_time: 20m
alertname: CNPGBackupStale
exp_alerts: []
- name: scheduled backup due=-7200
interval: 1m
input_series:
- series: kube_customresource_next_schedule_time{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="ScheduledBackup",cluster="parity-paradedb",scheduled_backup="daily"}
values: 946677600+0x40
alert_rule_test:
- eval_time: 20m
alertname: CNPGScheduledBackupStalled
exp_alerts:
- exp_labels:
severity: warning
namespace: alert-test
cnpg_cluster: parity-paradedb
cluster: parity-paradedb
scheduled_backup: daily
exp_annotations:
summary: ParadeDB CNPG ScheduledBackup is not advancing its schedule
description: CloudNativePG Cluster "alert-test/parity-paradedb" scheduled backup "daily" was due 8400 seconds ago and has not been rescheduled.
Check operator reconciliation.
runbook_url: https://github.com/paradedb/charts/blob/main/charts/paradedb/docs/runbooks/CNPGScheduledBackupStalled.md
- name: scheduled backup due=7200
interval: 1m
input_series:
- series: kube_customresource_next_schedule_time{namespace="alert-test",customresource_group="postgresql.cnpg.io",customresource_kind="ScheduledBackup",cluster="parity-paradedb",scheduled_backup="daily"}
values: 946692000+0x40
alert_rule_test:
- eval_time: 20m
alertname: CNPGScheduledBackupStalled
exp_alerts: []
- name: no backup history and no failure
interval: 1m
input_series: []
alert_rule_test:
- eval_time: 20m
alertname: CNPGBackupFailed
exp_alerts: []
- eval_time: 20m
alertname: CNPGBackupStale
exp_alerts: []
- eval_time: 20m
alertname: CNPGScheduledBackupStalled
exp_alerts: []
Loading