Skip to content

Hot & Cold Swap

Overview

Astra supports both hot swap and cold swap of computing boards. A failed node can be removed and replaced while the system continues serving users (hot swap), or the system can be fully powered down for a replacement (cold swap). In both cases, the replacement node automatically integrates itself into the cluster through the pg-autoheal self-repair system, requiring zero manual software configuration.

This capability is particularly important for a space deployment, where replacement hardware may need to be installed by crew members with basic technical training and no IT infrastructure expertise.

System-Level Hot Swap

Hot swap refers to the ability to remove and replace a computing board while the rest of the system remains operational and continues serving users without interruption.

How It Works

Astra's hot swap capability is a direct consequence of its triple-redundant architecture and Keepalived failover system:

  1. A node is physically unplugged. The Ethernet cable is disconnected, which simultaneously removes network connectivity and PoE power. The node powers off immediately.

  2. Keepalived detects the failure. The surviving nodes stop receiving VRRP advertisements from the removed node within 1--2 seconds. If the removed node was the MASTER (holding the VIP), the highest-priority surviving node transitions to MASTER and takes over the VIP.

  3. Users continue without interruption. Traffic to the VIP (10.0.0.50) is now handled by the new MASTER. Users may experience a brief page reload but their conversation history and all data are preserved through PostgreSQL streaming replication.

  4. The board is physically replaced. The old board is removed from the rack enclosure and a new board (or the same board after repair) is mounted in its place.

  5. The replacement is plugged back in. The Ethernet cable is reconnected, providing both power (via PoE) and network connectivity through a single cable.

  6. Automatic recovery begins. The replacement board boots, pg-autoheal detects that it is not the primary, and automatically rebuilds itself as a PostgreSQL standby via pg_basebackup from the current primary (accessed through the VIP). Full redundancy is restored within approximately 60 seconds.

Single-cable design enables clean swap

The PoE (Power over Ethernet) design is a deliberate architectural choice that simplifies hot swap. Each K3s node receives both power and data through a single Ethernet cable. Removing one cable cleanly disconnects the node from the system. There is no separate power cable to manage, no risk of disconnecting data but not power (leaving a zombie node), and no need to coordinate the order of cable disconnections. One cable in, one cable out.

What Constitutes a Valid Replacement

A replacement board must have:

  • The same operating system image (Ubuntu Server ARM64) with K3s, PostgreSQL 16, and Keepalived pre-installed
  • A matching static IP address on eth0 for the LAN
  • The pg-autoheal script and systemd service installed

The replacement does not need to contain any application data. The database is rebuilt from the current primary via pg_basebackup, and container images are pulled (or pre-cached) as part of the standard K3s agent join process.

SD card cloning for faster replacement

A replacement board can be prepared in advance by cloning the microSD card of a working node. However, with the pg-autoheal system, even a board with a fresh OS installation will automatically rebuild its database from the current primary, making pre-cloning optional.

Cold Swap

Cold swap refers to replacing a board while the entire system is powered down. This is the simpler but slower option, appropriate when the system is not actively in use.

Procedure

  1. Power down the system. Unplug all PoE switches and the network router.
  2. Remove the failed board. Unscrew it from the rack enclosure and disconnect the Ethernet cable.
  3. Mount the replacement board. Secure it in the same rack position and connect the Ethernet cable.
  4. Power up the system. Plug in the network router and all PoE switches.
  5. Automatic recovery. Keepalived runs its priority-based election to determine the MASTER. The pg-autoheal system on non-primary nodes detects their status and rebuilds them as PostgreSQL standbys. The system reaches full operational status within approximately 2 minutes.

Astronaut Swap Procedure

The following six-step procedure is designed for crew members with basic technical training. It requires no command-line access, no SSH sessions, and no software configuration knowledge.

Step 1: Identify the Failed Board

Open the status dashboard at http://10.0.0.50:9090 from any device connected to the infra-net WiFi network. The failed device appears with an OFFLINE status indicator and a red highlight on its card in the 9-device grid.

Dashboard shows which physical position to check

Each card in the dashboard grid corresponds to a labeled physical position in the rack enclosure. The card's hostname (e.g., k3s-c1-node1) maps directly to a labeled slot in the DeskPi RackMate T1.

Step 2: Disconnect the Ethernet Cable

Unplug the single Ethernet cable from the failed board. On PoE-powered nodes, this simultaneously removes power and network connectivity. The board powers off immediately.

Step 3: Remove the Board

Unscrew the board from the rack enclosure mounting hardware. Remove it from its slot.

Step 4: Mount the Replacement Board

Insert the replacement board into the same physical slot and secure it with the mounting screws.

Step 5: Reconnect the Ethernet Cable

Plug the Ethernet cable back in. The board powers on via PoE and begins its boot sequence automatically.

Step 6: Verify Recovery on the Dashboard

Monitor the status dashboard. The replacement board's card transitions through the following states:

Dashboard State Meaning
OFFLINE Board is not yet reachable on the network
BOOTING Board is reachable but services are still starting
HEALING pg-autoheal is running pg_basebackup to rebuild the standby database
READY Board is fully operational as a standby, redundancy restored

The entire recovery process -- from cable insertion to READY status -- typically completes within 60--90 seconds.

No CLI required

The entire swap procedure is physical: identify, unplug, remove, mount, plug in, verify. The software recovery is fully automatic. This is critical for a space deployment where crew time is limited and IT expertise cannot be assumed.

Comparison: Hot Swap vs. Cold Swap

Characteristic Hot Swap Cold Swap
System remains operational during swap Yes No
User impact Brief page reload (seconds) Full outage during power cycle
Recovery time after replacement ~60 seconds ~2 minutes
Requires coordination None -- unplug anytime Must power down entire system
When to use Active operation, single board failure Scheduled maintenance, multiple replacements
Complexity for crew Slightly higher (must check dashboard first) Lower (power off, swap, power on)