Skip to content

Auto-Heal System

The Problem: Split-Brain After Power Loss

When a node loses power and then reboots, its PostgreSQL instance starts up in whatever state it was in before the outage. This creates a dangerous scenario:

  • If the node was a standby: It may have stale data that is behind the current primary. It needs to re-synchronize with the current primary before it can safely serve as a standby again.
  • If the node was the primary before failover: It still believes it is the primary, but another node has since been promoted. Two nodes both acting as primary create a split-brain condition -- writes go to both nodes, and the data diverges irrecoverably.

The node must automatically detect the current state of the cluster and rejoin as a standby without manual intervention. This is what the pg-autoheal system does.

How pg-autoheal Works

The pg-autoheal.sh script runs once at boot via a systemd oneshot service. It examines the local node's state, the cluster's state, and makes a decision: either exit (no action needed) or rebuild the local PostgreSQL as a standby.

Decision Flowchart

Boot
  |
  v
Wait 15 seconds (let Keepalived determine VIP ownership)
  |
  v
Do I hold the VIP? --YES--> Exit (I am the primary, nothing to do)
  |
  NO
  |
  v
Is PostgreSQL already a standby? --YES--> Exit (already correct state)
  |
  NO
  |
  v
Wait for VIP to be reachable (max 120 seconds)
  |
  v
Apply hostname-based stagger delay (C1=0s, C2=30s, C3=60s)
  |
  v
Re-check: Did I get promoted during the stagger? --YES--> Exit
  |
  NO
  |
  v
Stop local PostgreSQL
  |
  v
Delete local data directory
  |
  v
Run pg_basebackup from VIP (with up to 3 retries)
  |
  v
Start PostgreSQL as standby
  |
  v
Verify: Is PostgreSQL in recovery mode? --YES--> Success

Step-by-Step Explanation

1. Boot delay (15 seconds)

The script starts with a 15-second pause (implemented via ExecStartPre=/bin/sleep 15 in the systemd unit). This delay allows Keepalived to complete its VRRP election and assign the VIP to the appropriate node before pg-autoheal begins examining cluster state.

2. VIP ownership check

The script checks whether the local node holds the VIP (10.0.0.50). If it does, this node is the current primary and no action is needed. The script exits immediately.

3. Standby state check

If the node does not hold the VIP, the script checks whether PostgreSQL is already running as a standby (by looking for the standby.signal file in the data directory). If it is already a standby, the node is in the correct state and no action is needed.

4. VIP reachability wait

The script waits for the VIP to be reachable (via ping), with a maximum timeout of 120 seconds. This ensures that at least one other node is up and serving as the primary before attempting to rebuild. If the VIP is not reachable within 120 seconds, the script logs an error and exits.

5. Stagger delay

After confirming the VIP is reachable, the script applies a hostname-based stagger delay (see Stagger Delays below). This prevents multiple nodes from running pg_basebackup simultaneously.

6. Post-stagger re-check

After the stagger delay, the script re-checks whether conditions have changed. During the stagger period, Keepalived may have promoted this node to MASTER (if the previous primary failed during the stagger). If the node now holds the VIP or has become a standby through other means, the script exits without taking action.

7. Rebuild as standby

If all guards pass, the script performs the actual rebuild:

  1. Stops the local PostgreSQL service
  2. Deletes the entire data directory (/var/lib/postgresql/16/main)
  3. Runs pg_basebackup from the VIP to create a fresh copy of the primary's data
  4. Starts PostgreSQL, which comes up in standby mode (the pg_basebackup -R flag creates the standby.signal file and configures primary_conninfo automatically)

8. Verification

After PostgreSQL starts, the script verifies that the node is in recovery mode (pg_is_in_recovery() returns true). If verification fails, the script logs an error.

Safety Guards

The pg-autoheal script is designed to be safe under all conditions. It includes multiple guards that prevent it from taking harmful actions:

Guard Protection
Never wipes the primary First check: if this node holds the VIP, exit immediately. The script will never delete data on the active primary.
Idempotent If PostgreSQL is already a correctly configured standby, the script does nothing. Running it multiple times has no side effects.
Waits for cluster Does not attempt to rebuild until the VIP is reachable, ensuring at least one primary exists in the cluster.
Post-stagger re-check Re-examines conditions after the stagger delay, catching cases where the node's role changed during the wait.
Retries with backoff If pg_basebackup fails, retries up to 3 times with increasing wait periods between attempts.
Full logging Every decision point and action is logged to /var/log/pg-autoheal.log for post-incident analysis.

Why the script deletes the data directory

The pg_basebackup utility requires an empty target directory. Rather than attempting to incrementally repair a potentially corrupted data directory, the script deletes it entirely and creates a fresh copy from the primary. This approach is more reliable because it eliminates any possibility of data corruption from the previous state carrying over into the new standby.

Stagger Delays

The Thundering Herd Problem

During testing, a critical failure mode was discovered: when all three nodes boot simultaneously after a full power cycle, they all attempt pg_basebackup from the primary at the same time. The primary's WAL sender processes cannot handle three simultaneous full-database copies, causing all three to fail.

Worse, because the script deletes the data directory before running pg_basebackup, a failed basebackup leaves the node with an empty data directory and no way to recover automatically. This was a data safety issue that required a design solution.

Hostname-Based Stagger

The solution is deterministic stagger delays based on the node's hostname:

Node Hostname Pattern Stagger Delay
k3s-c1-node1 Contains c1 0 seconds
k3s-c2-node1 Contains c2 30 seconds
k3s-c3-node1 Contains c3 60 seconds

The 30-second intervals ensure that only one node runs pg_basebackup at a time. A full basebackup of the Astra database completes in well under 30 seconds on the LAN, so each node finishes before the next one starts.

Why hostname-based and not random

Random delays could still result in collisions (two nodes choosing similar random values). Deterministic delays based on hostname guarantee that nodes never overlap, and the behavior is predictable and reproducible for debugging.

Post-Stagger Guard

After the stagger delay completes, the script re-checks two conditions:

  1. Did I get promoted? During the stagger, Keepalived may have promoted this node to MASTER (if the previous primary failed). If the node now holds the VIP, it should not rebuild as a standby.
  2. Did I become a standby? Another process (or a previous boot cycle's pg-autoheal) may have already configured this node as a standby during the stagger period.

Only if both conditions are negative does the script proceed with the rebuild.

Retry Logic

If pg_basebackup fails, the script retries with increasing wait periods:

Attempt Wait Before Retry
1st retry 30 seconds
2nd retry 60 seconds
3rd retry 90 seconds

The maximum number of retries is 3 (MAX_RETRIES=3). If all retries fail, the script exits with an error and the node remains without a PostgreSQL data directory. Manual intervention is required in this case (see the manual rebuild procedure in Cascade Failover).

Possible causes of pg_basebackup failure include:

  • The primary is temporarily overloaded (another node's basebackup is still running)
  • Network issues between the node and the VIP
  • The primary's max_wal_senders limit is reached

The increasing wait times give transient conditions time to resolve before retrying.

Systemd Service

The pg-autoheal script runs as a systemd oneshot service:

Property Value
Service name pg-autoheal.service
Type oneshot (runs once at boot, then exits)
Boot delay 15 seconds (ExecStartPre=/bin/sleep 15)
Script path /opt/hunch-status-dashboard/pg-autoheal.sh
Log file /var/log/pg-autoheal.log

The service is configured to run after the network is up and after systemd basic targets are reached. The 15-second ExecStartPre delay ensures Keepalived has time to complete its VRRP election before pg-autoheal examines the cluster state.

Oneshot services run once and exit

Unlike long-running services (like PostgreSQL or Keepalived), a oneshot service executes its script once and then terminates. The pg-autoheal.service does its work during boot and does not continue running. If a rebuild is needed later (not at boot), the script can be run manually.

Logging

All pg-autoheal activity is logged to /var/log/pg-autoheal.log. The log records:

  • Script start time and hostname
  • Each guard check result (VIP ownership, standby status, VIP reachability)
  • Stagger delay duration and post-stagger re-check results
  • pg_basebackup start, completion, or failure
  • Retry attempts and wait times
  • Final verification result

Example log output for a successful rebuild:

[2026-04-15 10:32:15] pg-autoheal starting on k3s-c2-node1
[2026-04-15 10:32:15] VIP check: I do not hold 10.0.0.50
[2026-04-15 10:32:15] Standby check: standby.signal not found, not a standby
[2026-04-15 10:32:16] VIP 10.0.0.50 is reachable
[2026-04-15 10:32:16] Applying stagger delay: 30 seconds (c2)
[2026-04-15 10:32:46] Post-stagger check: still need to rebuild
[2026-04-15 10:32:46] Stopping PostgreSQL
[2026-04-15 10:32:47] Removing data directory
[2026-04-15 10:32:47] Running pg_basebackup from 10.0.0.50
[2026-04-15 10:33:02] pg_basebackup completed successfully
[2026-04-15 10:33:02] Starting PostgreSQL
[2026-04-15 10:33:04] Verification: pg_is_in_recovery() = t (standby mode confirmed)
[2026-04-15 10:33:04] pg-autoheal complete: successfully rebuilt as standby

Full Recovery Scenario

The following timeline shows the complete auto-heal process after a full power cycle (all three nodes lose power and then reboot simultaneously):

Time Event
0s All three nodes begin booting
~30s All nodes are up. Keepalived begins VRRP election.
~32s C1 (priority 100) wins VRRP election, receives VIP. C1's PostgreSQL becomes primary.
~45s pg-autoheal starts on all nodes (after 15s boot delay).
~45s C1's pg-autoheal detects it holds the VIP. Exits (no action).
~45s C2 and C3's pg-autoheal detect they do not hold the VIP and are not standbys.
~46s C2 applies 30s stagger. C3 applies 60s stagger.
~76s C2's stagger completes. Re-checks pass. Begins pg_basebackup.
~91s C2's pg_basebackup completes. C2 starts as standby. Replication from C1 begins.
~106s C3's stagger completes. Re-checks pass. Begins pg_basebackup.
~121s C3's pg_basebackup completes. C3 starts as standby. Replication from C1 begins.
~121s Full three-way redundancy restored. System is fully operational.

Total recovery time: approximately 2 minutes

From a complete power loss to full three-way redundancy with streaming replication active across all nodes, the system recovers in approximately 2 minutes with zero manual intervention.

Deployment Files

Path Purpose
/opt/hunch-status-dashboard/pg-autoheal.sh The auto-heal script
/etc/systemd/system/pg-autoheal.service Systemd unit file
/var/log/pg-autoheal.log Log file