Demo Guide¶
Overview¶
This guide covers the complete procedure for demonstrating Astra at NASA HUNCH events and presentations. The demo is designed to showcase the system's core capabilities: AI-powered medical assistance, real-time failover, cascade failure survival, and automatic self-healing recovery.
Pre-Demo Setup¶
Complete these steps before the audience arrives. The entire setup takes approximately 2--3 minutes.
Step 1: Power On¶
Plug in all three PoE switches and the network router. Ensure all Ethernet cables are connected between the network router, the PoE switches, and the computing devices.
Cable check
Verify that each PoE switch has its power indicator lit and that the link LEDs are active on each port connected to a Raspberry Pi or AI server. A missing link LED may indicate a loose or faulty cable.
Step 2: Wait for Boot¶
Allow approximately 2 minutes for all 9 devices to boot. The Raspberry Pis boot in approximately 30--45 seconds, but the full stack (K3s, PostgreSQL, Keepalived, Ollama) requires additional initialization time.
Step 3: Connect to WiFi¶
On a laptop, tablet, or phone, connect to the WiFi network:
| Parameter | Value |
|---|---|
| SSID | infra-net |
| Security | WPA3-PSK |
The WiFi password is provided separately to demo attendees.
Step 4: Open the Dashboard¶
Navigate to http://10.0.0.50:9090 in a web browser. This is the status dashboard, served through the Keepalived VIP.
Step 5: Verify Readiness¶
Wait for the Failover Readiness Panel at the top of the dashboard to show a green READY indicator. This confirms:
- All three clusters are operational
- PostgreSQL replication is active on all standbys
- The VIP is assigned and routing traffic correctly
If READY does not appear within 3 minutes
Check the dashboard for specific cluster status. Common causes of delayed readiness:
- A standby is still running pg-autoheal (status shows HEALING -- wait for it to complete)
- A PoE cable is loose (device shows OFFLINE -- reseat the cable)
- A node is still booting (K3s or PostgreSQL service not yet started)
Live Demo Sequence¶
The following five-step sequence is designed for maximum impact during live presentations.
Step 1: Show the Dashboard¶
Navigate to: http://10.0.0.50:9090
Talking points:
- Point out the Failover Readiness Panel: "This green indicator means all three independent clusters are healthy and fully synchronized. The system can survive losing any two of the three right now."
- Show the per-cluster cards: "Each of these cards represents an independent cluster. C1 is the current primary -- it holds the Virtual IP and is accepting all user traffic. C2 and C3 are hot standbys, receiving a continuous stream of replicated data."
- Show the AI server status: "These three cards represent the AI inference servers. The Jetson Orin Nano has a CUDA GPU and is our fastest server. The two Pi 5 units provide CPU-based inference."
- Point out temperature readings: "Every device reports its CPU temperature, voltage, and throttle status in real time. In space, thermal monitoring is critical because there is no convective cooling in microgravity."
Step 2: Use Open WebUI¶
Navigate to: http://10.0.0.50:30080
Login credentials:
| Field | Value |
|---|---|
hunch@example.com |
|
| Password | hunch |
Talking points:
- "This is Open WebUI, our AI chat interface. It looks and works like ChatGPT, but everything runs locally on the hardware you see in front of you. No internet, no cloud, no external servers."
- Select the model
ai3.qwen3:4bfor the fastest response times (runs on the Jetson GPU at approximately 25--40 tokens per second). - Start a conversation. For a medical demonstration, try a question such as: "What are the symptoms of acute radiation syndrome, and what immediate first aid steps should be taken?"
- "The AI model running on our Jetson device is responding in real time. Every word you see is generated locally on the hardware in this rack."
Model selection for different demo purposes
- Fastest responses: Select
ai3.qwen3:4b(Jetson GPU, ~25--40 tok/s) - Medical questions: Select
meditronormedgemma-1.5-4b-it - Speed demo ("wow factor"): Select
gemma3:1bon any server (near-instant responses due to the tiny model size)
Step 3: Demonstrate Failover¶
While the conversation is still active and visible on the screen:
- Announce what you are about to do: "I am going to physically unplug one-third of the hardware. Watch what happens to the conversation."
- Physically unplug one PoE switch cable. This kills one entire power domain: 2 K3s nodes and 1 AI server.
- Show the audience: The page may briefly reload (approximately 5 seconds), but the conversation continues with all history intact.
- Show the dashboard: Navigate back to
http://10.0.0.50:9090. The lost cluster appears as offline, and the Failover Readiness Panel transitions from READY to HEALING or NO REDUNDANCY. - Continue the conversation to prove the system is fully functional with only two-thirds of its hardware.
Talking points:
- "I just removed one-third of the server hardware -- two Raspberry Pis and one AI server. The system detected the failure in under 2 seconds, promoted a standby to primary, and continued serving. The user never lost their conversation."
- "In space, this is what happens when a solar flare damages one of the three server racks. The system keeps working."
Step 4: Demonstrate Cascade Failover¶
With one switch already unplugged:
- Announce: "Now I am going to unplug a second rack. Two-thirds of the hardware will be gone."
- Unplug a second PoE switch cable.
- Show the dashboard: Only one cluster remains green. The system is in NO REDUNDANCY mode.
- Continue the conversation on Open WebUI. It still works.
Talking points:
- "The system just survived losing two-thirds of its hardware. Only one cluster and one AI server remain, and the conversation is still running with all data intact."
- "This is cascade failover. The data was replicated from the first cluster to the second, and from the second to the third. The last survivor has a complete copy of everything."
Step 5: Demonstrate Recovery¶
- Plug both PoE switches back in.
- Show the dashboard: Watch the clusters appear as devices come online. The Failover Readiness Panel shows HEALING as pg-autoheal rebuilds the standbys.
- Wait approximately 60--90 seconds. The dashboard transitions back to READY.
Talking points:
- "Watch the dashboard. The system is now automatically rebuilding. Each returning node detects that it is not the primary, connects to the surviving primary, and copies a fresh replica of the entire database. No human intervention required."
- "Within about a minute, we are back to full triple redundancy. The system healed itself."
- "This is the key capability for a space application: the system survived losing two-thirds of its hardware, kept working the entire time, and automatically recovered to full redundancy when the hardware came back. Zero manual intervention, zero data loss, zero downtime."
Voice Mode Setup¶
Open WebUI supports voice input for hands-free interaction. This requires Chrome with a specific flag enabled.
Configuration¶
- Open Chrome and navigate to
chrome://flags/#unsafely-treat-insecure-origin-as-secure - In the text field, enter:
http://10.0.0.50:30080 - Set the dropdown to Enabled
- Click Relaunch to restart Chrome
- Navigate to
http://10.0.0.50:30080and grant microphone permissions when prompted
Safari limitation
Safari does not support microphone access over HTTP under any configuration. Voice mode requires Chrome with the insecure-origin flag described above. This is a browser security policy limitation, not an Astra limitation.
Using Voice Mode¶
After enabling the flag, click the microphone icon in the Open WebUI chat input area. Speak your question, and the browser's speech-to-text engine will transcribe it into the chat input. The AI response will appear as text (text-to-speech output depends on browser support).
Troubleshooting¶
VIP Not Responding¶
If http://10.0.0.50:30080 does not load:
- Check WiFi connection. Ensure your device is connected to
infra-net, not a different network. - Try individual node IPs. Access Open WebUI directly at:
http://10.0.0.11:30080(Cluster 1)http://10.0.0.21:30080(Cluster 2)http://10.0.0.31:30080(Cluster 3)
- Check the dashboard. Navigate to
http://10.0.0.11:9090(or .21 or .31) to see which nodes are online. - Verify Keepalived. If no individual node IPs work either, the issue is likely at the network layer (check cable connections and PoE switch power).
Dashboard Shows HEALING for Extended Period¶
If a cluster remains in HEALING state for more than 2 minutes:
- The pg-autoheal script may be running a
pg_basebackupoperation, which can take 30--60 seconds depending on database size. - Check the hostname-based stagger: C2 waits 30 seconds and C3 waits 60 seconds before starting their recovery, so a full system restart may take up to 2 minutes before all nodes are healthy.
- If HEALING persists beyond 3 minutes, there may be a connectivity issue preventing
pg_basebackupfrom reaching the primary via the VIP.
AI Models Not Responding¶
If the AI chat returns errors or timeouts:
- Check the AI server status on the dashboard. If all three AI servers show OFFLINE, the AI inference layer has no available servers.
- Verify that at least one AI server is powered on and connected to the network router LAN.
- If an AI server shows ONLINE but models are not responding, the Ollama service may need a restart (this requires SSH access, which is outside the scope of the demo procedure).
Open WebUI Shows Database Error on Standby Clusters¶
Open WebUI instances on standby clusters (C2 and C3 under normal conditions) may display database errors because the local PostgreSQL is a read-only standby. This is expected behavior. Only the primary cluster (the one holding the VIP) will have a fully functional Open WebUI instance. Access the system through the VIP (10.0.0.50) to ensure you reach the active primary.