Skip to content

Network and Power Infrastructure

Astra is built on a self-contained network that requires zero external connectivity. The network and power infrastructure is designed around a single organizing principle: physical independence between power domains, ensuring that a failure in one domain cannot propagate to another.


Network Router

The central network component is a compact router that serves as both the network switch and WiFi access point for the entire system. The current prototype uses a GL.iNet OpenWrt-based router. In a Gateway deployment, this role would be filled by the station's internal network infrastructure.

Property Value
LAN Subnet 10.0.0.0/24
DHCP Range 10.0.0.100--200
WiFi SSID infra-net
WiFi Security WPA3-PSK
LAN Ports 3 (WAN port converted to LAN)

The network router provides three critical functions:

  1. Central switching. All three PoE switches connect to the router's LAN ports, creating a flat Layer 2 network where all 9 devices can communicate directly.
  2. WiFi access point. The WiFi network provides access for connected devices. No coordination with external IT infrastructure is required.
  3. Portability. The router is smaller than a smartphone. The entire Astra system can be set up in any location without depending on existing network infrastructure.

Self-contained operation

The router does not need an internet uplink to function. It creates a completely self-contained LAN. All AI models, container images, embedding caches, and application code are pre-loaded on every device. The system operates identically whether or not the WAN port is connected to the internet.


PoE Switches and Power Domains

Three Power over Ethernet (PoE) switches form the backbone of both the network and power delivery system. Each switch is connected to a separate power strip, creating three independent power and network domains.

Power & Network Domain Layout

Power & Network       Power & Network       Power & Network
Domain A              Domain B              Domain C
(PoE Switch 1)        (PoE Switch 2)        (PoE Switch 3)
====================  ====================  ====================
k3s-c1-node1 (.11)    k3s-c2-node1 (.21)    k3s-c3-node1 (.31)
k3s-c1-node2 (.12)    k3s-c2-node2 (.22)    k3s-c3-node2 (.32)
ai-server1   (.41)    ai-server2   (.42)    ai-server3   (.43)

Each domain contains exactly one K3s cluster (2 nodes) and one AI inference server. The K3s nodes receive both power and networking from the PoE switch. The AI servers are connected to the PoE switch for networking only -- they have their own independent power supplies (USB-C or DC barrel jack). This distribution ensures that any single-domain failure removes at most one cluster from service.

PoE+ Power Delivery (802.3at)

The K3s compute nodes (Raspberry Pi 4 units) are powered via PoE+ (IEEE 802.3at), which delivers both power and data over a single ethernet cable through a PoE HAT (Hardware Attached on Top) mounted on each Pi.

Standard Voltage Max Power per Port Cable
802.3at (PoE+) 48V DC 25.5W Cat5e or better

This single-cable design is critical for the hot-swap capability: disconnecting the ethernet cable from a Pi simultaneously removes its network connection and its power, cleanly shutting it down. Reconnecting the cable restores both power and network, triggering the automatic boot and self-heal sequence.

Single-cable simplicity for astronauts

The PoE design means that swapping a failed board requires only unplugging and re-plugging a single ethernet cable. No separate power cable management is needed. This simplicity is important for crew members performing hardware maintenance in a space station environment where cable management space is limited.


Failure Scenarios

The three-domain architecture provides predictable behavior under progressive failure:

Loss of One Domain

Impact: One K3s cluster (2 nodes) and one AI server go offline.

System behavior: Keepalived detects the lost node1 within approximately 1--2 seconds. The highest-priority surviving node1 receives the Virtual IP and promotes its PostgreSQL to primary (approximately 2--3 seconds for promotion). Total failover time is approximately 5 seconds. The remaining standby automatically reconnects to the new primary via the VIP. Users experience at most a brief page reload. AI inference continues on the two remaining servers.

Redundancy remaining: 2 clusters, 2 AI servers. Full data redundancy maintained.

Loss of Two Domains

Impact: Two K3s clusters (4 nodes) and two AI servers go offline.

System behavior: The sole surviving cluster takes ownership of the VIP and serves all traffic. Its PostgreSQL instance contains a complete copy of all data (replicated before the failures). One AI server continues handling inference requests.

Redundancy remaining: 1 cluster, 1 AI server. The system is fully functional but has no redundancy -- a third failure would cause a total outage.

Loss of All Three Domains

Impact: Total system outage. All devices are powered off.

System behavior: No service is available. However, all data is preserved on disk across all nodes. When power is restored, the system self-heals automatically:

  1. Keepalived elects the highest-priority surviving node1 as the primary.
  2. The pg-autoheal service on the other node1s detects they are not the primary, waits for the VIP to become reachable, and rebuilds themselves as streaming standbys via pg_basebackup.
  3. Full triple redundancy is restored within approximately 2 minutes.

Stagger delays during simultaneous recovery

When all three nodes boot simultaneously, the pg-autoheal system uses hostname-based stagger delays (C1: 0 seconds, C2: 30 seconds, C3: 60 seconds) to prevent all nodes from attempting pg_basebackup at the same time. Without this staggering, the primary's WAL sender capacity can be overwhelmed, causing all recovery attempts to fail.


Keepalived Virtual IP (VIP)

Keepalived manages a Virtual IP address (10.0.0.50) using the VRRP (Virtual Router Redundancy Protocol). This VIP is the single entry point for all user traffic and always resolves to the active primary cluster.

Property Value
Virtual IP 10.0.0.50/24
Interface eth0 (LAN)
Protocol VRRP (IP protocol 112)
Communication Unicast between the three node1s
Advertisement interval 1 second

Priority and Failover Order

Node Priority Role
k3s-c1-node1 100 (highest) Preferred primary
k3s-c2-node1 90 First failover target
k3s-c3-node1 80 (lowest) Last resort

When a higher-priority node fails, the next-highest-priority surviving node transitions to MASTER state, receives the VIP on its eth0 interface, and executes the notify script to promote its local PostgreSQL and restart the Open WebUI deployment.

Why physical ethernet for VRRP?

VRRP uses IP protocol 112, which is neither TCP nor UDP. VPN tunnels (such as WireGuard) only transport TCP, UDP, and ICMP traffic, so VRRP packets sent over a VPN interface are silently dropped. Keepalived must run on the physical eth0 interface with direct Layer 2 connectivity on the LAN.


Network Topology Summary

                    Internet (optional, not required)
                              |
                    +-------------------+
                    |  Network Router   |
                    |  10.0.0.1 (gw)    |
                    |  WiFi: infra-net  |
                    +---+-----+-----+--+
                        |     |     |
              +---------+  +--+--+  +----------+
              |            |     |             |
         PoE Switch 1  PoE Switch 2    PoE Switch 3
              |            |     |             |
    +---------+---+   +----+----+   +----+-----+
    | .11  .12  .41|  | .21 .22 .42|  | .31 .32 .43|
    | C1N1 C1N2 AI1|  | C2N1 C2N2 AI2|  | C3N1 C3N2 AI3|
    +------+-------+  +------+------+  +------+------+
           |                  |                |
           +------ VIP: 10.0.0.50 (floats) ---+

All user-facing access goes through the VIP at 10.0.0.50. The status dashboard is available at http://10.0.0.50:9090. Individual devices can also be accessed directly by their LAN IP addresses for diagnostic purposes.


Physical Enclosure

All computing hardware is mounted in a DeskPi RackMate T1 compact rack enclosure. This enclosure provides:

  • Physical containment and protection for all boards and cabling
  • Structured airflow with an open rear panel and top/bottom ventilation
  • Mounting points for securing boards against vibration
  • A compact form factor suitable for installation in a stowage locker on Gateway

In a space station deployment, the rack would be secured inside a standard storage locker, limiting physical access to authorized crew members and protecting the hardware from floating debris in the microgravity environment.