Cascade Failover¶
Streaming Replication Topology¶
PostgreSQL 16 runs bare metal on each of the three node1 servers. At any given time, one node is the PRIMARY (writable) and the other two are streaming STANDBYs (read-only). Data flows from the primary to the standbys in real time via PostgreSQL's built-in streaming replication protocol.
Normal State:
+---------------------+
| k3s-c1-node1 |
| PRIMARY (RW) |
| Priority: 100 |
+---------+-----------+
|
Async Streaming Replication
|
+--------------+--------------+
| |
+---------+----------+ +---------+----------+
| k3s-c2-node1 | | k3s-c3-node1 |
| STANDBY (RO) | | STANDBY (RO) |
| Priority: 90 | | Priority: 80 |
+--------------------+ +--------------------+
Replication Characteristics¶
Asynchronous replication. Write operations are acknowledged to the application before being replicated to standbys. This design prevents a slow or unreachable standby from blocking writes on the primary. The trade-off is that in a catastrophic primary failure, the last few milliseconds of committed writes might not have reached the standbys.
Sub-millisecond lag. On the LAN (10.0.0.0/24), measured replication lag is consistently under 1 millisecond. In practice, the window of potential data loss during an unplanned failover is negligible -- typically less than one database transaction.
pgvector data included. The pgvector extension stores embedding vectors as regular PostgreSQL data. All vector embeddings are replicated along with every other table, ensuring that RAG functionality is fully available on any node that is promoted to primary.
VIP-Based Replication: The Key Innovation¶
The most important architectural decision in Astra's database layer is how standby nodes locate the primary for replication. The primary_conninfo setting in each standby's PostgreSQL configuration points to the Virtual IP (VIP) address 10.0.0.50, not to any specific node's fixed IP address.
This distinction is critical and warrants detailed explanation.
The Standard Approach (Fragile)¶
In a conventional PostgreSQL replication setup, each standby's primary_conninfo contains the primary node's fixed IP address:
# Standard approach -- standby points to a fixed IP
primary_conninfo = 'host=10.0.0.11 port=5432 user=replicator password=...'
This works as long as 10.0.0.11 (C1's node1) is the primary. But when C1 fails and C2 is promoted to primary:
- C2 is now the primary, serving on
10.0.0.21 - C3 is still a standby, but its
primary_conninfostill points to10.0.0.11(the dead C1 node) - C3 cannot replicate from C2 because it does not know C2 is the new primary
- Manual reconfiguration is required on C3 to update
primary_conninfoto10.0.0.21
This manual step defeats the purpose of automatic failover.
The Astra Approach (Self-Healing)¶
In Astra, every standby's primary_conninfo points to the VIP:
# Astra approach -- standby points to the VIP
primary_conninfo = 'host=10.0.0.50 port=5432 user=replicator password=...'
Now when C1 fails and C2 is promoted:
- C2 becomes the primary and receives the VIP (
10.0.0.50) from Keepalived - C3 is still a standby, and its
primary_conninfostill points to10.0.0.50 - C3 automatically reconnects to C2 via the VIP -- zero manual intervention
- Replication resumes from C2 to C3 without any configuration change
Why this matters for space
In a space station deployment, there is no system administrator to SSH into a node and update configuration files during a failure. The VIP-based approach ensures that the replication topology reconfigures itself automatically, regardless of which node fails or which node is promoted.
Timeline Following¶
Combined with recovery_target_timeline = 'latest' (the PostgreSQL 16 default setting), standbys automatically follow timeline switches that occur when a standby promotes to primary. When a standby is promoted, PostgreSQL creates a new WAL timeline. The latest setting tells remaining standbys to follow the newest timeline, enabling them to seamlessly reconnect to the newly promoted primary.
Cascade Failover Walkthrough¶
The VIP-based replication design enables cascade failover -- the ability to survive multiple sequential node failures while preserving all data. The following scenario demonstrates the full cascade.
Scenario: Kill C1, Then Kill C2¶
Initial state: C1 is PRIMARY (priority 100), C2 and C3 are STANDBYs.
Step 1: C1 fails (power pulled from PoE Switch 1)
+---------------------+
| k3s-c1-node1 |
| DOWN |
| (power lost) |
+---------------------+
+---------------------+ +---------------------+
| k3s-c2-node1 | | k3s-c3-node1 |
| STANDBY -> PRIMARY| | STANDBY (RO) |
| Gets VIP | | Points to VIP |
+---------+-----------+ +---------+-----------+
| |
+--------- Replication --------+
- Keepalived on C2 and C3 detects that C1's VRRP advertisements have stopped (within 1-2 seconds)
- C2 (priority 90, higher than C3's 80) transitions to MASTER and receives the VIP
- C2's Keepalived notify script promotes local PostgreSQL from standby to primary
- C3, still pointing to VIP
10.0.0.50, automatically reconnects to C2 - Replication resumes from C2 to C3
Result: System is operational with C2 as primary and C3 as standby. All data from C1 is present on both C2 and C3.
Step 2: C2 fails (power pulled from PoE Switch 2)
+---------------------+ +---------------------+
| k3s-c1-node1 | | k3s-c2-node1 |
| DOWN | | DOWN |
+---------------------+ +---------------------+
+---------------------+
| k3s-c3-node1 |
| STANDBY -> PRIMARY|
| Gets VIP |
| ALL DATA PRESENT |
+---------------------+
- Keepalived on C3 detects that C2's VRRP advertisements have stopped
- C3 (the only surviving node) transitions to MASTER and receives the VIP
- C3's notify script promotes local PostgreSQL from standby to primary
- C3 is now the sole survivor with all data from C1 and C2
Result: System is operational on a single node. All historical data, conversations, user accounts, and embeddings are intact. Users experience a brief page reload and can continue working.
Cascade order does not matter
This cascade works regardless of which nodes fail and in what order. Whether you kill C1 then C2, C2 then C1, C1 then C3, or any other combination, the last surviving node always has a complete copy of all data. This is because every standby is continuously replicated from the primary, and the VIP ensures that replication always points to the current primary.
Recovery After Cascade¶
When killed nodes are powered back on, the pg-autoheal system (see Auto-Heal System) automatically handles recovery:
- The rebooted node detects that it does not hold the VIP (another node is the current primary)
- It waits for the VIP to be reachable
- It applies a hostname-based stagger delay to prevent multiple nodes from recovering simultaneously
- It rebuilds itself as a standby via
pg_basebackupfrom the VIP - Streaming replication resumes
The system returns to full three-way redundancy without manual intervention.
Replication Lag Monitoring¶
Replication lag can be monitored from the current primary:
Expected output for a healthy system:
client_addr | state | replay_lag
----------------+-----------+------------
10.0.0.21 | streaming | 00:00:00
10.0.0.31 | streaming | 00:00:00
A replay_lag value of 00:00:00 indicates sub-second replication lag. Under normal LAN conditions, the actual lag is consistently under 1 millisecond.
Replication lag during recovery
During pg_basebackup (when a node is rebuilding), the primary sends both live WAL data to the existing standby and a full database copy to the recovering node. This can temporarily increase replication lag to the existing standby. The lag returns to sub-millisecond levels once the basebackup completes.
Manual Standby Rebuild¶
If the automatic pg-autoheal process fails, a standby can be manually rebuilt:
# Stop PostgreSQL on the node being rebuilt
sudo systemctl stop postgresql
# Remove the old data directory
sudo rm -rf /var/lib/postgresql/16/main
# Run pg_basebackup from the VIP
sudo -u postgres PGPASSWORD=replicator_hunch pg_basebackup \
-h 10.0.0.50 -U replicator -D /var/lib/postgresql/16/main \
-Fp -Xs -P -R
# Start PostgreSQL
sudo systemctl start postgresql
# Verify standby status
sudo -u postgres psql -c "SELECT pg_is_in_recovery();"
# Expected: t (true)
pg_basebackup from VIP sets primary_conninfo automatically
When pg_basebackup is run with the -R flag and the VIP as the host, it automatically configures the new standby's primary_conninfo to point to the VIP. This ensures the rebuilt standby participates correctly in future cascade failovers.