Skip to content

Data Redundancy

Overview

Astra's data redundancy strategy is built around a core principle: in a space environment, the most relevant failure mode is entire-node failure (caused by cosmic radiation, power events, thermal issues, or physical damage), not individual disk failure. The system uses distributed replication across physically separated nodes rather than local RAID arrays, providing superior protection against the failure modes most likely to occur aboard a spacecraft.

Data redundancy operates at two independent levels:

  1. PostgreSQL streaming replication for structured data (users, conversations, settings, vector embeddings)
  2. lsyncd file synchronization for unstructured data (uploaded documents, shared files)

Storage Hardware

All storage in Astra uses solid-state media. There are no spinning disks, which is important for a space application where vibration during launch and microgravity conditions make mechanical drives unreliable.

Device Storage Type Capacity Notes
K3s Cluster Nodes (6x Raspberry Pi) MicroSD (solid-state flash) 64--128 GB each Boot, OS, and application storage
AI Server 1 (Raspberry Pi 5) MicroSD (solid-state flash) 128 GB OS and Ollama model storage
AI Server 2 (Raspberry Pi 5) MicroSD (solid-state flash) 128 GB OS and Ollama model storage
AI Server 3 (Jetson Orin Nano) eMMC (embedded solid-state) 128 GB Soldered to the compute module; no external media

Distributed Replication vs. Local RAID

Why Not RAID

A traditional RAID array (such as RAID 1 mirroring) protects against the failure of an individual disk within a single machine. The disks are inside the same chassis, on the same power supply, exposed to the same thermal environment, and subject to the same radiation exposure. RAID protects against one specific failure mode: a single disk dying while the machine continues to operate.

In a space environment, the more dangerous failure modes render RAID irrelevant:

  • A cosmic ray strike that corrupts a node's memory or processor affects all disks on that node equally
  • A power supply failure takes out the entire node, including all its RAID disks
  • Thermal runaway affects the entire enclosure, not individual disks
  • Physical damage from a micrometeorrite or equipment impact is not limited to a single disk

Comparison

Characteristic Local RAID (single node) Astra's Distributed Replication
Protects against Individual disk failure Disk failure, node failure, power domain failure, any 2-of-3 node failures
Does NOT protect against Node failure, power loss, radiation damage to the node Simultaneous loss of all 3 nodes
Hardware required Multiple disks per node Standard storage on each node (no extra disks)
Recovery after failure Automatic RAID rebuild (if node survives) Automatic pg_basebackup rebuild from VIP
Power overhead Additional disks require additional power No additional power -- uses existing storage
Physical space Requires space for multiple disks per node No additional space
Failure domain independence No -- all disks share the same failure domain Yes -- each copy is on a separate power domain

Why Not Physical SSDs on Raspberry Pi

Raspberry Pi boards do not have native SATA or NVMe interfaces. Adding external SSDs requires USB adapters, which introduce several problems in Astra's design context:

  • Power draw: USB-attached SSDs draw additional power from the Pi's USB bus. For PoE-powered nodes with a limited power budget (the PoE standard provides a fixed wattage per port), this additional draw can push the node past its power allocation, causing instability or undervoltage throttling.
  • Physical bulk: USB adapters and cables add volume and cable management complexity to an already compact rack enclosure.
  • Additional failure points: Each USB adapter, cable, and SSD connector is a potential point of failure. In a vibration-prone environment (launch, docking), these additional connections are liabilities.
  • PoE single-cable design conflict: Astra's hot-swap capability depends on each node having a single cable (Ethernet/PoE) that provides both power and data. Adding a separate USB SSD breaks this clean single-cable abstraction and complicates the swap procedure.

The Jetson Orin Nano uses onboard eMMC (128 GB embedded solid-state storage soldered directly to the compute module), which provides SSD-like performance without any external adapters.

Engineering trade-off

The decision to use microSD storage with distributed replication instead of local SSDs is a deliberate engineering trade-off. Distributed replication provides better fault tolerance for the space use case (protecting against node-level failures, not just disk-level failures) while preserving the PoE single-cable hot-swap design. The trade-off is accepted lower I/O performance on individual nodes, which is acceptable for Astra's workload profile (primarily database reads/writes and web application serving, not I/O-intensive data processing).

PostgreSQL Streaming Replication

Architecture

PostgreSQL 16 runs bare metal on each of the three node1 servers. At any given time, one node is the PRIMARY (accepting read and write operations) and the other two are STANDBYs (read-only replicas that receive a continuous stream of write-ahead log records from the primary).

Node Default Role Priority
k3s-c1-node1 PRIMARY (writable) 100 (highest)
k3s-c2-node1 STANDBY (read-only) 90
k3s-c3-node1 STANDBY (read-only) 80 (lowest)

Replication Characteristics

Parameter Value
Replication mode Asynchronous streaming
Measured replication lag Sub-millisecond on the network router LAN
WAL level replica
Maximum WAL senders 4
Replication transport TCP over network router LAN (VIP-based)
Vector embeddings Replicated (pgvector extension installed on all nodes)

Asynchronous replication means that writes are acknowledged to the application before being confirmed as replicated to standbys. This prevents a slow or temporarily unreachable standby from blocking writes on the primary. The trade-off is that in a catastrophic primary failure, the last few milliseconds of writes might not have reached the standbys. In practice, with sub-millisecond replication lag on the LAN, this window is negligible.

VIP-Based primary_conninfo

Each standby's primary_conninfo points to the VIP (10.0.0.50), not to a specific node's fixed IP address. This design decision is critical for enabling automatic cascade failover.

When the current primary fails and a standby is promoted (receiving the VIP via Keepalived), all remaining standbys automatically reconnect to the new primary through the same VIP address. No configuration change is required on any standby. Combined with recovery_target_timeline = 'latest' (the PostgreSQL 16 default), standbys follow timeline switches that occur during promotion, enabling seamless cascade failover.

For a detailed description of the cascade failover mechanism, see the Keepalived & VIP and Cascade Failover documentation.

What Is Replicated

All data in the openwebui database is replicated, including:

  • User accounts and credentials (bcrypt-hashed passwords)
  • Conversation history (all messages and responses)
  • Application settings and configuration (including Ollama endpoint URLs)
  • Vector embeddings for RAG (stored via the pgvector extension)
  • Session data

Auto-Heal Recovery

When a failed node is powered back on, the pg-autoheal system automatically detects that the node is not the primary and rebuilds it as a standby via pg_basebackup from the current primary (accessed through the VIP). This process is fully automatic and typically completes within 60 seconds.

For a detailed description of the auto-heal system, see the Auto-Heal System documentation.

lsyncd File Synchronization

Architecture

lsyncd (Live Syncing Daemon) runs on all three node2 (agent) servers. It provides bidirectional file synchronization of the /data/shared/ directory across all three node2s.

Parameter Value
Service lsyncd
Configuration file /etc/lsyncd/lsyncd.conf.lua
Synchronized directory /data/shared/
Transport rsync over SSH
Propagation time Approximately 5 seconds
Authentication SSH keys distributed between all node2s

How It Works

lsyncd monitors the /data/shared/ directory for filesystem events (file creation, modification, deletion) using Linux's inotify subsystem. When a change is detected, lsyncd triggers an rsync operation to replicate the change to the other two node2s via SSH.

The synchronization is bidirectional: a file created on any node2 appears on the other two within approximately 5 seconds. Deletions are also propagated -- removing a file on one node2 removes it from all three.

Purpose

lsyncd provides RAID-like redundancy for files that are not stored in the PostgreSQL database. This includes:

  • Documents uploaded to Open WebUI for RAG processing
  • Shared configuration files
  • Any files placed in /data/shared/ by operators or the application

This complements the PostgreSQL replication layer: the database handles structured data, and lsyncd handles unstructured file data. Together, they ensure that all user data is replicated across at least three physically separated nodes.

Why Node2s Instead of Node1s

File synchronization runs on the node2 (agent) servers rather than the node1 (server) servers to separate concerns:

  • Node1 servers are already running K3s control plane, PostgreSQL, Keepalived, the dashboard, and pg-autoheal. Adding file synchronization would increase the workload on these already-busy nodes.
  • Node2 servers have a lighter workload (K3s agent only) and can absorb the I/O overhead of file synchronization without impacting critical services.
  • Placing file storage on the agent nodes provides an independent failure domain -- a node1 failure does not affect file storage, and a node2 failure does not affect the database or failover system.