Skip to content

Tested Failover Scenarios

Overview

Astra's fault-tolerance claims are backed by systematic testing of every plausible failure mode. Each scenario described below has been executed on the physical hardware, with results verified through the status dashboard, PostgreSQL replication queries, and direct user interaction with Open WebUI.

This level of testing is essential for a system intended for a space application. In space, there is no on-call engineer to diagnose unexpected behavior, no remote SSH session to patch a misconfiguration, and no way to physically intervene until the next crew rotation. Every failure mode the system might encounter must be anticipated, tested, and shown to resolve automatically. The scenarios below demonstrate that Astra handles single-node failures, cascade failures, full power cycles, and partial infrastructure loss without manual intervention.

Scenario Matrix

The following table summarizes all tested failover scenarios. Each scenario has been executed on the physical hardware and verified to produce the expected behavior.

# Scenario Expected Behavior Verified
1 Kill C1 (primary, priority 100) C2 promotes to primary within approximately 2 seconds. C3 automatically re-replicates from C2 via the VIP. Users experience a brief page reload. Yes
2 Kill C2 (after C2 was promoted) The highest-priority survivor (C1 if available, otherwise C3) promotes to primary. Remaining standby reconnects via VIP. Yes
3 Kill C1, then kill C2 (cascade) C2 promotes first after C1 dies. When C2 is subsequently killed, C3 promotes with all data from both C1 and C2 intact. Yes
4 Kill C2, then kill C1 (reverse cascade) C1 remains primary after C2 dies (C1 has higher priority). After C1 is killed, C3 promotes with all data. Yes
5 Kill any 2 nodes, plug both back in Both killed nodes auto-heal as standbys via pg-autoheal within approximately 60 seconds. Full triple redundancy is restored. Yes
6 Pod crash (single cluster) K3s detects the failed pod and restarts it on the same node or reschedules it to the other node in the cluster within approximately 3 minutes. Yes
7 Full power cycle (all 3 off, then all 3 on) Keepalived elects a primary based on priority. Other nodes stagger-heal as standbys via pg-autoheal. Full recovery within approximately 2 minutes. Yes
8 AI server failure Open WebUI continues to function with the remaining AI servers. Model availability is reduced, but the chat interface, database, and all non-AI features remain fully operational. Yes

Detailed Scenario Descriptions

Scenario 1: Single Primary Failure

Test procedure: Physically unplug the PoE switch powering Cluster 1 (the default primary with priority 100).

What this proves: The most basic failover case. The system must detect the failure, elect a new primary, promote its PostgreSQL instance, and resume serving users -- all without human intervention.

Observed behavior: Within approximately 2 seconds of power loss to C1, Keepalived on C2 and C3 detected the missing VRRP advertisements. C2 (priority 90) transitioned to MASTER, received the VIP, promoted its local PostgreSQL to primary, and restarted Open WebUI. C3 remained a standby and automatically reconnected to C2 via the VIP. The conversation in progress on the browser survived with a brief page reload.

Scenario 2: Secondary Failure After Promotion

Test procedure: After C2 has been promoted to primary (following the kill of C1), unplug the PoE switch powering Cluster 2.

What this proves: The system can handle a failure of a promoted node, not just the original primary. The failover logic must work regardless of which node is currently serving as primary.

Observed behavior: C3 detected the loss of C2's VRRP advertisements and promoted to MASTER. PostgreSQL on C3 was promoted from standby to primary. All data written to C2 during its time as primary was present on C3, confirming that streaming replication kept the standby current.

Scenario 3: Cascade Failure (Kill C1, Then C2)

Test procedure: Unplug C1. Wait for C2 to promote. Then unplug C2. Only C3 remains.

What this proves: The system can survive the loss of two out of three clusters in sequence, preserving all data across both failures. This is the most demanding failover scenario because it tests the VIP-based primary_conninfo design -- C3 must automatically follow the primary from C1 to C2, and then promote itself when C2 dies.

Observed behavior: After C1 was killed, C2 promoted and began accepting writes. C3 reconnected to C2 via the VIP without any configuration change. After C2 was killed, C3 promoted to MASTER. All data from both C1 and C2 was present and accessible on C3. This confirms that the VIP-based replication design enables true cascade failover.

Scenario 4: Reverse Cascade (Kill C2, Then C1)

Test procedure: Unplug C2 first (a standby). Then unplug C1 (the primary). Only C3 remains.

What this proves: Cascade failover works regardless of the order in which nodes fail. Whether the primary fails first or a standby fails first, the end result is the same: the last surviving node has a complete copy of all data.

Observed behavior: When C2 was killed, C1 remained primary (it was unaffected by the loss of a standby). When C1 was subsequently killed, C3 promoted to MASTER and had all data. The replication lag at the time of C1's failure was sub-millisecond, so no writes were lost.

Scenario 5: Recovery After Dual Failure

Test procedure: Kill any two nodes. Wait for the survivor to stabilize as primary. Then plug both killed nodes back in simultaneously.

What this proves: The auto-heal system (pg-autoheal) correctly rebuilds returning nodes as standbys without human intervention, even when two nodes return at the same time. The stagger delays prevent thundering herd issues during simultaneous recovery.

Observed behavior: Both nodes booted, ran pg-autoheal, and rebuilt themselves as PostgreSQL standbys via pg_basebackup from the surviving primary (accessed through the VIP). The hostname-based stagger delays (0s, 30s, 60s) ensured that only one node ran pg_basebackup at a time. Full triple redundancy was restored within approximately 60 seconds.

Scenario 6: Pod Crash Within a Single Cluster

Test procedure: Delete the Open WebUI pod on one cluster using kubectl delete pod.

What this proves: K3s provides intra-cluster fault tolerance. A crashed application pod is automatically restarted by the Kubernetes scheduler without affecting the database, other clusters, or the VIP.

Observed behavior: K3s detected the missing pod and created a replacement within seconds. The new pod was scheduled on the same node or the other node in the cluster (depending on resource availability) and was serving HTTP requests within approximately 3 minutes. No data was lost because the database runs independently of the application pod.

Scenario 7: Full Power Cycle

Test procedure: Simultaneously unplug all three PoE switches (total system power loss). Wait 10 seconds. Plug all three back in simultaneously.

What this proves: The system can recover from a complete power outage without any manual intervention. This is the most realistic simulation of a power event on a space station, where an entire rack might lose power due to a circuit breaker trip or solar panel issue.

Observed behavior: All three nodes booted simultaneously. Keepalived ran its election and determined the MASTER based on priority (C1, priority 100). The pg-autoheal system on C2 and C3 detected that they were not the primary, waited for the VIP to become reachable, applied their stagger delays, and rebuilt as standbys. The entire system was fully operational with triple redundancy within approximately 2 minutes.

Scenario 8: AI Server Failure

Test procedure: Power off one AI server (any of the three).

What this proves: The AI inference layer is independent of the application layer. Losing an AI server reduces model availability but does not cause any downtime in the chat interface, database, or failover system.

Observed behavior: Open WebUI displayed the remaining AI servers' models in the model selector. Requests to models hosted only on the downed server returned an error, but models available on the surviving servers continued to work. When the AI server was powered back on, its models reappeared in the selector within approximately 30 seconds (the Ollama health check interval).

Why This Testing Matters for Space

In terrestrial computing, system failures are inconvenient but rarely life-threatening. An engineer can SSH into a machine, restart a service, or drive to a data center. In space, none of these options exist:

  • Communication delay makes remote troubleshooting impractical for urgent issues
  • Crew training focuses on mission objectives, not IT infrastructure management
  • Physical access to computing hardware may be restricted by module layout and crew schedules
  • Replacement parts are not available until the next resupply mission

Every scenario tested above represents a failure that the system handles without any human intervention. The system detects the failure, reconfigures itself, and resumes operation -- exactly the behavior required for a medical support system in a deep-space environment where help from Earth may be hours away.