Hot & Cold Swap¶
Overview¶
Astra supports both hot swap and cold swap of computing boards. A failed node can be removed and replaced while the system continues serving users (hot swap), or the system can be fully powered down for a replacement (cold swap). In both cases, the replacement node automatically integrates itself into the cluster through the pg-autoheal self-repair system, requiring zero manual software configuration.
This capability is particularly important for a space deployment, where replacement hardware may need to be installed by crew members with basic technical training and no IT infrastructure expertise.
System-Level Hot Swap¶
Hot swap refers to the ability to remove and replace a computing board while the rest of the system remains operational and continues serving users without interruption.
How It Works¶
Astra's hot swap capability is a direct consequence of its triple-redundant architecture and Keepalived failover system:
-
A node is physically unplugged. The Ethernet cable is disconnected, which simultaneously removes network connectivity and PoE power. The node powers off immediately.
-
Keepalived detects the failure. The surviving nodes stop receiving VRRP advertisements from the removed node within 1--2 seconds. If the removed node was the MASTER (holding the VIP), the highest-priority surviving node transitions to MASTER and takes over the VIP.
-
Users continue without interruption. Traffic to the VIP (
10.0.0.50) is now handled by the new MASTER. Users may experience a brief page reload but their conversation history and all data are preserved through PostgreSQL streaming replication. -
The board is physically replaced. The old board is removed from the rack enclosure and a new board (or the same board after repair) is mounted in its place.
-
The replacement is plugged back in. The Ethernet cable is reconnected, providing both power (via PoE) and network connectivity through a single cable.
-
Automatic recovery begins. The replacement board boots, pg-autoheal detects that it is not the primary, and automatically rebuilds itself as a PostgreSQL standby via
pg_basebackupfrom the current primary (accessed through the VIP). Full redundancy is restored within approximately 60 seconds.
Single-cable design enables clean swap
The PoE (Power over Ethernet) design is a deliberate architectural choice that simplifies hot swap. Each K3s node receives both power and data through a single Ethernet cable. Removing one cable cleanly disconnects the node from the system. There is no separate power cable to manage, no risk of disconnecting data but not power (leaving a zombie node), and no need to coordinate the order of cable disconnections. One cable in, one cable out.
What Constitutes a Valid Replacement¶
A replacement board must have:
- The same operating system image (Ubuntu Server ARM64) with K3s, PostgreSQL 16, and Keepalived pre-installed
- A matching static IP address on
eth0for the LAN - The pg-autoheal script and systemd service installed
The replacement does not need to contain any application data. The database is rebuilt from the current primary via pg_basebackup, and container images are pulled (or pre-cached) as part of the standard K3s agent join process.
SD card cloning for faster replacement
A replacement board can be prepared in advance by cloning the microSD card of a working node. However, with the pg-autoheal system, even a board with a fresh OS installation will automatically rebuild its database from the current primary, making pre-cloning optional.
Cold Swap¶
Cold swap refers to replacing a board while the entire system is powered down. This is the simpler but slower option, appropriate when the system is not actively in use.
Procedure¶
- Power down the system. Unplug all PoE switches and the network router.
- Remove the failed board. Unscrew it from the rack enclosure and disconnect the Ethernet cable.
- Mount the replacement board. Secure it in the same rack position and connect the Ethernet cable.
- Power up the system. Plug in the network router and all PoE switches.
- Automatic recovery. Keepalived runs its priority-based election to determine the MASTER. The pg-autoheal system on non-primary nodes detects their status and rebuilds them as PostgreSQL standbys. The system reaches full operational status within approximately 2 minutes.
Astronaut Swap Procedure¶
The following six-step procedure is designed for crew members with basic technical training. It requires no command-line access, no SSH sessions, and no software configuration knowledge.
Step 1: Identify the Failed Board¶
Open the status dashboard at http://10.0.0.50:9090 from any device connected to the infra-net WiFi network. The failed device appears with an OFFLINE status indicator and a red highlight on its card in the 9-device grid.
Dashboard shows which physical position to check
Each card in the dashboard grid corresponds to a labeled physical position in the rack enclosure. The card's hostname (e.g., k3s-c1-node1) maps directly to a labeled slot in the DeskPi RackMate T1.
Step 2: Disconnect the Ethernet Cable¶
Unplug the single Ethernet cable from the failed board. On PoE-powered nodes, this simultaneously removes power and network connectivity. The board powers off immediately.
Step 3: Remove the Board¶
Unscrew the board from the rack enclosure mounting hardware. Remove it from its slot.
Step 4: Mount the Replacement Board¶
Insert the replacement board into the same physical slot and secure it with the mounting screws.
Step 5: Reconnect the Ethernet Cable¶
Plug the Ethernet cable back in. The board powers on via PoE and begins its boot sequence automatically.
Step 6: Verify Recovery on the Dashboard¶
Monitor the status dashboard. The replacement board's card transitions through the following states:
| Dashboard State | Meaning |
|---|---|
| OFFLINE | Board is not yet reachable on the network |
| BOOTING | Board is reachable but services are still starting |
| HEALING | pg-autoheal is running pg_basebackup to rebuild the standby database |
| READY | Board is fully operational as a standby, redundancy restored |
The entire recovery process -- from cable insertion to READY status -- typically completes within 60--90 seconds.
No CLI required
The entire swap procedure is physical: identify, unplug, remove, mount, plug in, verify. The software recovery is fully automatic. This is critical for a space deployment where crew time is limited and IT expertise cannot be assumed.
Comparison: Hot Swap vs. Cold Swap¶
| Characteristic | Hot Swap | Cold Swap |
|---|---|---|
| System remains operational during swap | Yes | No |
| User impact | Brief page reload (seconds) | Full outage during power cycle |
| Recovery time after replacement | ~60 seconds | ~2 minutes |
| Requires coordination | None -- unplug anytime | Must power down entire system |
| When to use | Active operation, single board failure | Scheduled maintenance, multiple replacements |
| Complexity for crew | Slightly higher (must check dashboard first) | Lower (power off, swap, power on) |