Astra: A Fault-Tolerant AI Medical Assistant for Space¶
NASA HUNCH Mini Rack Project -- Mischa Nelson & Urvi Dhenge
Astra is a fully redundant, offline-capable AI medical assistant system designed for deployment aboard the NASA Gateway space station and future deep-space missions. Built entirely by high school students under the NASA HUNCH (High School Students United with NASA to Create Hardware) program using commercial off-the-shelf hardware, Astra demonstrates that enterprise-grade fault tolerance is achievable with accessible components.
The Problem¶
Astronauts operating beyond low Earth orbit face a convergence of risks that no existing system adequately addresses:
- Communication delay. While Gateway's lunar orbit has near-real-time communication (~1.3 seconds), future missions to Mars and beyond will face delays up to 24 minutes one way. Astra is designed for both scenarios -- providing immediate assistance regardless of communication latency. In a medical emergency on a deep-space mission, waiting nearly an hour for a single exchange with a ground-based physician could be fatal.
- Limited crew medical training. Crew members receive basic medical training, but they are not physicians. Complex diagnoses and treatment decisions require expert guidance that may not be immediately available.
- No reliable internet in space. Deep-space missions will have severely limited or nonexistent internet connectivity. Any system that depends on cloud services is unusable.
- Hardware fails. Cosmic radiation, thermal cycling, vibration during launch, and component aging all contribute to hardware failures. A single server running a medical AI is a single point of failure.
The Solution¶
Astra addresses every one of these challenges:
- On-board AI inference. The medical assistant runs locally with zero network latency. Astronauts get immediate responses without waiting for Earth.
- Medical-specialized models. The system runs AI models specifically trained on medical literature alongside general-purpose models, providing domain-specific guidance.
- Fully offline operation. Every component -- the web interface, the AI models, the database, the embeddings -- runs locally with zero internet dependency.
- Triple redundancy with automatic failover. Three independent clusters, each on its own power domain, with automatic failover and self-healing. The system survives losing any two of three clusters and recovers automatically when power is restored.
Key Statistics¶
| Metric | Value |
|---|---|
| Total physical devices | 9 |
| Independent K3s clusters | 3 (6 Raspberry Pi nodes) |
| AI inference servers | 3 (2x Pi 5 + 1x Jetson Orin Nano) |
| AI models deployed | 7 (including 2 medical-specialized) |
| Database replication | 3-way PostgreSQL streaming, sub-millisecond lag |
| Failover time | Approximately 5 seconds |
| Cascade failover | Tested -- survives loss of any 2 of 3 clusters |
| Internet dependency | Zero -- fully offline capable |
| Recovery after total power loss | Fully automatic within 2 minutes |
Team¶
- Mischa Nelson -- Lead engineer and system architect.
- Urvi Dhenge -- Project contributor.
Documentation¶
This wiki provides comprehensive technical documentation of every aspect of the Astra system.
System Design¶
- Architecture Overview -- System topology, functional layers, and network design
- Requirements Traceability -- Complete mapping of NASA HUNCH requirements to Astra's implementation
Hardware¶
- Compute Nodes -- The 6 Raspberry Pi units running K3s clusters
- AI Servers -- The 3 dedicated AI inference servers
- Network and Power -- Network router, PoE switches, and power domain design
- Sensors and Thermal Management -- Hardware monitoring and cooling systems
Kubernetes & Deployment¶
- K3s Overview -- Why K3s, how it works, ARM64 support
- Cluster Topology -- 3 independent clusters, server vs agent roles
- Helm & Deployment -- Helm charts, environment variables, offline caching
AI & Medical Models¶
- Ollama Engine -- Inference server configuration and API
- Model Selection -- All 7 models, selection rationale, performance benchmarks
- Medical AI -- Meditron, MedGemma, and the medical use case
Database & Replication¶
- PostgreSQL Architecture -- Database design, pgvector, bare-metal deployment
- Cascade Failover -- VIP-based replication and cascade failover
- Auto-Heal System -- Self-repairing after power loss
High Availability¶
- Keepalived & VIP -- VRRP configuration, notify scripts, failover timeline
- Tested Scenarios -- 8 verified failover scenarios
- Hot & Cold Swap -- Board replacement procedures
Application & Operations¶
- Open WebUI -- Application layer, offline operation, voice mode
- Monitoring & Dashboard -- 9-device hardware monitoring dashboard
- Networking & Security -- Air-gapped LAN, authentication, security model
- Data Redundancy -- Distributed replication vs RAID, storage architecture
Engineering¶
- Design Decisions -- Rationale behind every major architectural choice
- Challenges & Solutions -- Technical problems solved during development
Reference¶
- Demo Guide -- Step-by-step live demonstration procedure
- Requirements Traceability -- NASA HUNCH requirements mapped to implementation
- Replication Guide -- Bill of materials and build instructions for future teams