Requirements Traceability¶
This page provides a comprehensive mapping of every requirement from the NASA HUNCH Mini Rack Scope of Work to Astra's implementation. Each requirement is listed with its current status and a detailed description of how the Astra system addresses it.
The requirements traceability matrix demonstrates that Astra meets or exceeds all specified requirements while operating within the constraints of commercial off-the-shelf hardware.
Traceability Matrix¶
| # | Requirement | Status | How Astra Addresses It |
|---|---|---|---|
| 1 | Design and develop a Raspberry Pi / NVIDIA Jetson server farm suitable for space conditions | Complete | Astra deploys 9 devices organized into 3 independent K3s clusters (6 Raspberry Pi 4 units) plus 3 AI inference servers (2 Raspberry Pi 5 units and 1 NVIDIA Jetson Orin Nano). The system provides triple redundancy with verified cascade failover, surviving the loss of any 2 of 3 clusters with zero data loss. |
| 2 | Utilize a Single or Double Storage Locker | Complete | All hardware is mounted in a DeskPi RackMate T1 compact rack enclosure designed to fit within a standard stowage locker form factor. The enclosure provides physical containment, structured airflow, and vibration-resistant mounting points. |
| 3 | Access to Multiple Sensors for Application Farm Health Test | Complete | Every device exposes hardware sensors monitored in real time. Raspberry Pi units provide CPU temperature, core voltage, and throttle/undervoltage detection via vcgencmd. The Jetson Orin Nano provides 9 thermal zones and INA3221 power rail monitoring (voltage, current, power). All sensor data is displayed in the live status dashboard with color-coded indicators and click-to-expand detail views. |
| 4 | Test performance and reliability in simulated space conditions | Complete | Eight distinct failover scenarios have been tested and verified: single-cluster failure, cascade failure (kill 2 of 3 clusters in sequence), reverse cascade, full power cycle recovery, pod crash recovery, AI server failure, and simultaneous multi-node recovery. Cascade failover has been verified to preserve all data through sequential loss of any 2 clusters in any order. |
| 5 | Evaluate potential applications and benefits in space | Complete | Astra implements a medical AI assistant using domain-specific models. Meditron (EPFL/Yale) provides medical text Q&A trained on medical literature. MedGemma 1.5 (Google) provides medical imaging and text analysis. These models demonstrate the value of on-board AI for astronaut health support during deep-space missions where communication with ground-based physicians is delayed or impossible. |
| 6 | Utilize Linux for all applications | Complete | All 9 devices run Linux. The K3s compute nodes run Ubuntu Server (ARM64). The Raspberry Pi 5 AI servers run Debian GNU/Linux 13 (trixie). The Jetson Orin Nano runs Ubuntu 22.04.5 LTS with Jetson Linux R36.4.7 (JetPack 6.2.1). |
| 7 | Detect Hardware failure | Complete | The status dashboard monitors per-device hardware telemetry including CPU temperature, core voltage, throttle status, memory utilization, and disk utilization. Service health is checked via systemd status queries for all critical services (K3s, PostgreSQL, Keepalived, Ollama). Device-level failure is detected via network timeout. The pg-autoheal system detects and automatically recovers from power loss events without manual intervention. |
| 8 | Provide cold and hot Swappable Boards | Complete | Hot swap: A board can be physically unplugged from a running cluster while the system continues serving users. Keepalived fails over the VIP within approximately 2 seconds. When a replacement board is plugged back in, pg-autoheal automatically rebuilds it as a PostgreSQL standby, restoring full redundancy within approximately 60 seconds. Cold swap: Power down the system, replace the board, power up. Keepalived elects a primary and pg-autoheal rebuilds standbys automatically. No command-line access or manual configuration is required for either procedure. |
| 9 | Provide SSD and RAID Storage for all Applications | Addressed | The Jetson Orin Nano uses 128 GB eMMC (embedded solid-state storage soldered to the module), providing SSD-equivalent performance. Raspberry Pi units use microSD (solid-state flash). Rather than local RAID, Astra implements distributed replication that provides superior fault tolerance for the space use case: PostgreSQL 3-way streaming replication protects database data against entire-node failure (not just disk failure), and lsyncd 3-way file synchronization provides RAID-like redundancy for shared files. Physical SSDs on Raspberry Pi require USB adapters that increase power draw (problematic for PoE-constrained nodes), add cable complexity, and introduce additional failure points. The distributed approach protects against the failure mode most relevant in space: entire-node loss from radiation, thermal events, or power failures. |
| 10a | Power | Complete | PoE+ (IEEE 802.3at) power delivery provides single-cable power and data to K3s compute nodes. Three independent PoE switches create three separate power domains, providing independent failure isolation. The AI servers have persistent static IPs and dedicated power connections within their respective domains. |
| 10b | Communication | Complete | The network router creates a self-contained LAN (10.0.0.0/24) with WiFi access (SSID: infra-net, WPA3-PSK). All inter-device communication uses this LAN with static IP addresses. Keepalived VRRP manages the Virtual IP for user-facing traffic. No external internet is required for any operational function. In a Gateway deployment, the system would connect to the station's internal network infrastructure. |
| 10c | Understand Thermals and manage | Complete | Hardware thermal sensors are monitored on all 9 devices. Raspberry Pi units expose CPU temperature via vcgencmd and the Linux thermal subsystem. The Jetson exposes 9 thermal zones (CPU, GPU, CV accelerators, SoC, junction temperature) and INA3221 power rail monitoring. All readings are displayed in real time on the status dashboard with color-coded thresholds (green below 60 degrees C, yellow 60--75 degrees C, red above 75 degrees C). |
| 10d | Quiet cooling | Complete | Cooling uses small-form-factor fans integrated into PoE HATs (Pi 4 units) and the Jetson's built-in PWM fan. No high-speed server fans are used. The acoustic profile is comparable to a desktop computer at idle, suitable for continuous operation in a crew habitat environment where persistent loud noise contributes to crew fatigue. |
| 10e | Secure (Physical and Software, APPs and OS) | Complete | Physical security: All hardware is contained within the DeskPi RackMate T1 enclosure. Network security: Air-gapped LAN with WPA3 WiFi encryption and no public-facing services. Application security: OpenWebUI enforces user authentication with bcrypt-hashed passwords. Database security: PostgreSQL requires password authentication for all connections. Management security: Encrypted VPN provides authenticated remote access for ground-based development, used only for management operations. |
| 11 | Monitor individual Raspberry PIs and JETSON for failure and swap | Complete | The status dashboard displays a 9-device grid where each device card shows real-time health metrics: temperature (with color coding), memory utilization, disk utilization, and service statuses. Cards are clickable to expand full detail views including load averages, per-service systemd status, throttle registers (Pis), individual thermal zone readings (Jetson), and power rail data (Jetson). A device that becomes unreachable displays an "OFFLINE" indicator. The Failover Readiness panel at the top of the dashboard shows overall system health as READY, HEALING, NO REDUNDANCY, or NOT READY. |
| 12 | Easily swappable by astronauts | Complete | The PoE single-cable design enables tool-free hot swap. Procedure: (1) Identify the failed board on the dashboard (shows OFFLINE with red indicator). (2) Disconnect the ethernet cable (PoE-powered boards power off immediately). (3) Remove and replace the board in the rack. (4) Reconnect the ethernet cable (board powers on and auto-heals). (5) Verify recovery on the dashboard (transitions from OFFLINE to HEALING to READY). No command-line access, SSH, or software configuration is required at any step. |
| 13 | Containerization Docker + Kubernetes or Ansible | Complete | Astra uses K3s, a lightweight CNCF-certified Kubernetes distribution designed for edge computing and ARM devices. K3s provides the same pod scheduling, self-healing, and service abstraction as full Kubernetes with significantly lower resource overhead. Applications are deployed via Helm charts. Three independent K3s clusters (not federated) provide cross-cluster redundancy. |
| 14 | Overlay a LLM Open Source and Demo Simple Application | Complete | Open WebUI v0.8.12 provides a ChatGPT-like web interface accessible via any web browser. Seven open-source AI models are deployed across three Ollama inference servers: Qwen3 (4B and 8B variants), Phi-4 Mini, Gemma3 (1B), Gemma4 E2B (multimodal), Meditron (medical), and MedGemma 1.5 (medical imaging). Users interact through a familiar chat interface with model selection, conversation history, and multi-turn dialogue. |
| 15 | Load a Narrow AI/ML Model simple Application | Complete | Two domain-specific medical AI models are deployed: Meditron (3.8 GB), a medical text Q&A model trained on medical literature and developed by EPFL and Yale University, and MedGemma 1.5 (3.3 GB, quantized to Q4_K_M), a medical imaging and text model by Google capable of analyzing medical images and answering clinical questions. Both models demonstrate narrow AI applied to astronaut health support. |
| 16 | Research where to install SSDs on a Raspberry Pi | Complete | Research was completed and documented. Raspberry Pi boards lack native SATA or NVMe interfaces; adding SSDs requires USB adapters that increase power draw (problematic for PoE-powered nodes with limited 25.5W budgets), add physical bulk and cable complexity, and introduce additional failure points. The engineering decision was to use distributed replication (PostgreSQL streaming + lsyncd file sync) over local SSDs, as distributed replication protects against entire-node failure -- the dominant failure mode in a space environment -- rather than only individual disk failure. The Jetson Orin Nano uses onboard 128 GB eMMC, providing solid-state performance without external adapters. |
Requirements Summary¶
| Category | Total | Complete | Addressed | Remaining |
|---|---|---|---|---|
| Hardware design | 4 | 4 | 0 | 0 |
| Sensors and monitoring | 3 | 3 | 0 | 0 |
| Testing and evaluation | 2 | 2 | 0 | 0 |
| Infrastructure (power, comms, thermal, security) | 5 | 5 | 0 | 0 |
| Storage | 2 | 1 | 1 | 0 |
| Software and applications | 3 | 3 | 0 | 0 |
| Total | 19 | 18 | 1 | 0 |
Addressed vs. Complete
Requirement 9 (SSD and RAID Storage) is marked as "Addressed" rather than "Complete" because the implementation uses distributed replication instead of traditional local RAID. This is a deliberate engineering trade-off: distributed replication provides superior data protection for the space use case (protects against entire-node failure, not just disk failure) while maintaining compatibility with the PoE single-cable power design. The requirement's intent -- reliable, redundant data storage -- is fully satisfied.
Tested Failover Scenarios¶
The following scenarios have been tested and verified as part of requirement 4 (performance and reliability testing):
| Scenario | Expected Behavior | Verified |
|---|---|---|
| Single cluster failure (kill C1) | C2 promotes to primary within approximately 2 seconds. C3 re-replicates from C2 via VIP. | Yes |
| Kill promoted cluster (kill C2 after C2 promoted) | Remaining highest-priority node promotes. | Yes |
| Cascade failure (kill C1, then kill C2) | C2 promotes first. After C2 is killed, C3 promotes with all data from both C1 and C2. | Yes |
| Reverse cascade (kill C2, then kill C1) | C1 remains primary after C2 dies. After C1 is killed, C3 promotes with all data. | Yes |
| Recovery (kill 2, plug both back in) | Both nodes auto-heal as standbys via pg-autoheal within approximately 60 seconds. | Yes |
| Pod crash (single cluster) | K3s restarts the pod on the same or other node within approximately 3 minutes. | Yes |
| Full power cycle (all 3 off, then all 3 on) | Keepalived elects primary. Others stagger-heal. Full recovery within approximately 2 minutes. | Yes |
| AI server failure | OpenWebUI continues functioning with remaining AI servers. Reduced model availability but no downtime. | Yes |