Auto-Heal System¶
The Problem: Split-Brain After Power Loss¶
When a node loses power and then reboots, its PostgreSQL instance starts up in whatever state it was in before the outage. This creates a dangerous scenario:
- If the node was a standby: It may have stale data that is behind the current primary. It needs to re-synchronize with the current primary before it can safely serve as a standby again.
- If the node was the primary before failover: It still believes it is the primary, but another node has since been promoted. Two nodes both acting as primary create a split-brain condition -- writes go to both nodes, and the data diverges irrecoverably.
The node must automatically detect the current state of the cluster and rejoin as a standby without manual intervention. This is what the pg-autoheal system does.
How pg-autoheal Works¶
The pg-autoheal.sh script runs once at boot via a systemd oneshot service. It examines the local node's state, the cluster's state, and makes a decision: either exit (no action needed) or rebuild the local PostgreSQL as a standby.
Decision Flowchart¶
Boot
|
v
Wait 15 seconds (let Keepalived determine VIP ownership)
|
v
Do I hold the VIP? --YES--> Exit (I am the primary, nothing to do)
|
NO
|
v
Is PostgreSQL already a standby? --YES--> Exit (already correct state)
|
NO
|
v
Wait for VIP to be reachable (max 120 seconds)
|
v
Apply hostname-based stagger delay (C1=0s, C2=30s, C3=60s)
|
v
Re-check: Did I get promoted during the stagger? --YES--> Exit
|
NO
|
v
Stop local PostgreSQL
|
v
Delete local data directory
|
v
Run pg_basebackup from VIP (with up to 3 retries)
|
v
Start PostgreSQL as standby
|
v
Verify: Is PostgreSQL in recovery mode? --YES--> Success
Step-by-Step Explanation¶
1. Boot delay (15 seconds)
The script starts with a 15-second pause (implemented via ExecStartPre=/bin/sleep 15 in the systemd unit). This delay allows Keepalived to complete its VRRP election and assign the VIP to the appropriate node before pg-autoheal begins examining cluster state.
2. VIP ownership check
The script checks whether the local node holds the VIP (10.0.0.50). If it does, this node is the current primary and no action is needed. The script exits immediately.
3. Standby state check
If the node does not hold the VIP, the script checks whether PostgreSQL is already running as a standby (by looking for the standby.signal file in the data directory). If it is already a standby, the node is in the correct state and no action is needed.
4. VIP reachability wait
The script waits for the VIP to be reachable (via ping), with a maximum timeout of 120 seconds. This ensures that at least one other node is up and serving as the primary before attempting to rebuild. If the VIP is not reachable within 120 seconds, the script logs an error and exits.
5. Stagger delay
After confirming the VIP is reachable, the script applies a hostname-based stagger delay (see Stagger Delays below). This prevents multiple nodes from running pg_basebackup simultaneously.
6. Post-stagger re-check
After the stagger delay, the script re-checks whether conditions have changed. During the stagger period, Keepalived may have promoted this node to MASTER (if the previous primary failed during the stagger). If the node now holds the VIP or has become a standby through other means, the script exits without taking action.
7. Rebuild as standby
If all guards pass, the script performs the actual rebuild:
- Stops the local PostgreSQL service
- Deletes the entire data directory (
/var/lib/postgresql/16/main) - Runs
pg_basebackupfrom the VIP to create a fresh copy of the primary's data - Starts PostgreSQL, which comes up in standby mode (the
pg_basebackup -Rflag creates thestandby.signalfile and configuresprimary_conninfoautomatically)
8. Verification
After PostgreSQL starts, the script verifies that the node is in recovery mode (pg_is_in_recovery() returns true). If verification fails, the script logs an error.
Safety Guards¶
The pg-autoheal script is designed to be safe under all conditions. It includes multiple guards that prevent it from taking harmful actions:
| Guard | Protection |
|---|---|
| Never wipes the primary | First check: if this node holds the VIP, exit immediately. The script will never delete data on the active primary. |
| Idempotent | If PostgreSQL is already a correctly configured standby, the script does nothing. Running it multiple times has no side effects. |
| Waits for cluster | Does not attempt to rebuild until the VIP is reachable, ensuring at least one primary exists in the cluster. |
| Post-stagger re-check | Re-examines conditions after the stagger delay, catching cases where the node's role changed during the wait. |
| Retries with backoff | If pg_basebackup fails, retries up to 3 times with increasing wait periods between attempts. |
| Full logging | Every decision point and action is logged to /var/log/pg-autoheal.log for post-incident analysis. |
Why the script deletes the data directory
The pg_basebackup utility requires an empty target directory. Rather than attempting to incrementally repair a potentially corrupted data directory, the script deletes it entirely and creates a fresh copy from the primary. This approach is more reliable because it eliminates any possibility of data corruption from the previous state carrying over into the new standby.
Stagger Delays¶
The Thundering Herd Problem¶
During testing, a critical failure mode was discovered: when all three nodes boot simultaneously after a full power cycle, they all attempt pg_basebackup from the primary at the same time. The primary's WAL sender processes cannot handle three simultaneous full-database copies, causing all three to fail.
Worse, because the script deletes the data directory before running pg_basebackup, a failed basebackup leaves the node with an empty data directory and no way to recover automatically. This was a data safety issue that required a design solution.
Hostname-Based Stagger¶
The solution is deterministic stagger delays based on the node's hostname:
| Node | Hostname Pattern | Stagger Delay |
|---|---|---|
| k3s-c1-node1 | Contains c1 |
0 seconds |
| k3s-c2-node1 | Contains c2 |
30 seconds |
| k3s-c3-node1 | Contains c3 |
60 seconds |
The 30-second intervals ensure that only one node runs pg_basebackup at a time. A full basebackup of the Astra database completes in well under 30 seconds on the LAN, so each node finishes before the next one starts.
Why hostname-based and not random
Random delays could still result in collisions (two nodes choosing similar random values). Deterministic delays based on hostname guarantee that nodes never overlap, and the behavior is predictable and reproducible for debugging.
Post-Stagger Guard¶
After the stagger delay completes, the script re-checks two conditions:
- Did I get promoted? During the stagger, Keepalived may have promoted this node to MASTER (if the previous primary failed). If the node now holds the VIP, it should not rebuild as a standby.
- Did I become a standby? Another process (or a previous boot cycle's pg-autoheal) may have already configured this node as a standby during the stagger period.
Only if both conditions are negative does the script proceed with the rebuild.
Retry Logic¶
If pg_basebackup fails, the script retries with increasing wait periods:
| Attempt | Wait Before Retry |
|---|---|
| 1st retry | 30 seconds |
| 2nd retry | 60 seconds |
| 3rd retry | 90 seconds |
The maximum number of retries is 3 (MAX_RETRIES=3). If all retries fail, the script exits with an error and the node remains without a PostgreSQL data directory. Manual intervention is required in this case (see the manual rebuild procedure in Cascade Failover).
Possible causes of pg_basebackup failure include:
- The primary is temporarily overloaded (another node's basebackup is still running)
- Network issues between the node and the VIP
- The primary's
max_wal_senderslimit is reached
The increasing wait times give transient conditions time to resolve before retrying.
Systemd Service¶
The pg-autoheal script runs as a systemd oneshot service:
| Property | Value |
|---|---|
| Service name | pg-autoheal.service |
| Type | oneshot (runs once at boot, then exits) |
| Boot delay | 15 seconds (ExecStartPre=/bin/sleep 15) |
| Script path | /opt/hunch-status-dashboard/pg-autoheal.sh |
| Log file | /var/log/pg-autoheal.log |
The service is configured to run after the network is up and after systemd basic targets are reached. The 15-second ExecStartPre delay ensures Keepalived has time to complete its VRRP election before pg-autoheal examines the cluster state.
Oneshot services run once and exit
Unlike long-running services (like PostgreSQL or Keepalived), a oneshot service executes its script once and then terminates. The pg-autoheal.service does its work during boot and does not continue running. If a rebuild is needed later (not at boot), the script can be run manually.
Logging¶
All pg-autoheal activity is logged to /var/log/pg-autoheal.log. The log records:
- Script start time and hostname
- Each guard check result (VIP ownership, standby status, VIP reachability)
- Stagger delay duration and post-stagger re-check results
pg_basebackupstart, completion, or failure- Retry attempts and wait times
- Final verification result
Example log output for a successful rebuild:
[2026-04-15 10:32:15] pg-autoheal starting on k3s-c2-node1
[2026-04-15 10:32:15] VIP check: I do not hold 10.0.0.50
[2026-04-15 10:32:15] Standby check: standby.signal not found, not a standby
[2026-04-15 10:32:16] VIP 10.0.0.50 is reachable
[2026-04-15 10:32:16] Applying stagger delay: 30 seconds (c2)
[2026-04-15 10:32:46] Post-stagger check: still need to rebuild
[2026-04-15 10:32:46] Stopping PostgreSQL
[2026-04-15 10:32:47] Removing data directory
[2026-04-15 10:32:47] Running pg_basebackup from 10.0.0.50
[2026-04-15 10:33:02] pg_basebackup completed successfully
[2026-04-15 10:33:02] Starting PostgreSQL
[2026-04-15 10:33:04] Verification: pg_is_in_recovery() = t (standby mode confirmed)
[2026-04-15 10:33:04] pg-autoheal complete: successfully rebuilt as standby
Full Recovery Scenario¶
The following timeline shows the complete auto-heal process after a full power cycle (all three nodes lose power and then reboot simultaneously):
| Time | Event |
|---|---|
| 0s | All three nodes begin booting |
| ~30s | All nodes are up. Keepalived begins VRRP election. |
| ~32s | C1 (priority 100) wins VRRP election, receives VIP. C1's PostgreSQL becomes primary. |
| ~45s | pg-autoheal starts on all nodes (after 15s boot delay). |
| ~45s | C1's pg-autoheal detects it holds the VIP. Exits (no action). |
| ~45s | C2 and C3's pg-autoheal detect they do not hold the VIP and are not standbys. |
| ~46s | C2 applies 30s stagger. C3 applies 60s stagger. |
| ~76s | C2's stagger completes. Re-checks pass. Begins pg_basebackup. |
| ~91s | C2's pg_basebackup completes. C2 starts as standby. Replication from C1 begins. |
| ~106s | C3's stagger completes. Re-checks pass. Begins pg_basebackup. |
| ~121s | C3's pg_basebackup completes. C3 starts as standby. Replication from C1 begins. |
| ~121s | Full three-way redundancy restored. System is fully operational. |
Total recovery time: approximately 2 minutes
From a complete power loss to full three-way redundancy with streaming replication active across all nodes, the system recovers in approximately 2 minutes with zero manual intervention.
Deployment Files¶
| Path | Purpose |
|---|---|
/opt/hunch-status-dashboard/pg-autoheal.sh |
The auto-heal script |
/etc/systemd/system/pg-autoheal.service |
Systemd unit file |
/var/log/pg-autoheal.log |
Log file |