Monitoring & Dashboard¶
Overview¶
Astra includes a custom-built status dashboard that provides real-time visibility into the health of all 9 devices, the state of PostgreSQL replication, AI server availability, and overall failover readiness. The dashboard is a single Python file with zero external dependencies -- it uses only the Python standard library -- and runs on all three node1 servers on port 9090.
The dashboard is accessible through the Keepalived VIP at http://10.0.0.50:9090, which means it automatically follows the active primary. It can also be reached directly via any node1's LAN IP on the same port.
Design Philosophy¶
Why a Custom Dashboard Instead of Grafana and Prometheus¶
Industry-standard monitoring solutions like Grafana (visualization) and Prometheus (metrics collection) are excellent tools for large-scale infrastructure. However, they introduce significant overhead that is poorly suited to Astra's resource-constrained environment:
| Concern | Grafana/Prometheus | Astra Dashboard |
|---|---|---|
| Dependencies | Go runtime, Node.js, multiple databases, plugin ecosystem | Zero -- Python standard library only |
| Memory footprint | Hundreds of MB for Prometheus TSDB + Grafana server | Minimal -- single-threaded Python process |
| Storage requirements | Time-series database for historical metrics | None -- stateless, real-time only |
| Additional services to monitor | Prometheus, Grafana, alertmanager, node-exporter | None -- one process per node |
| Setup complexity | Multi-service deployment with configuration files | Copy one Python file and one systemd unit |
| Student accessibility | Requires knowledge of PromQL, Grafana dashboarding | Standard Python -- readable and modifiable by students |
The dashboard provides exactly the information needed for operational monitoring and live demonstrations, with nothing extra. Every line of code can be understood by a student learning Python, and the entire application fits in a single file.
Accessing the Dashboard¶
| Access Method | URL |
|---|---|
| Via VIP (recommended for demos) | http://10.0.0.50:9090 |
| Direct to C1 node1 (LAN) | http://10.0.0.11:9090 |
| Direct to C2 node1 (LAN) | http://10.0.0.21:9090 |
| Direct to C3 node1 (LAN) | http://10.0.0.31:9090 |
| Direct via LAN | http://10.0.0.11:9090, http://10.0.0.21:9090, http://10.0.0.31:9090 |
The dashboard runs identically on all three node1 servers. When accessed via the VIP, it is served by whichever node is currently the Keepalived MASTER.
Dashboard Panels¶
Failover Readiness Panel¶
The most prominent element on the dashboard is the Failover Readiness Panel at the top of the page. It provides an at-a-glance assessment of the system's ability to survive a hardware failure.
| Status | Meaning | Visual Indicator |
|---|---|---|
| READY | All three clusters are healthy and replicating. The system can survive the loss of any two clusters. | Solid green |
| HEALING | One or more clusters are in the process of rebuilding after a failure. The system is operational but not yet fully redundant. | Pulsing amber animation |
| NO REDUNDANCY | Only one cluster is operational. The system is serving users but cannot survive another failure. | Solid amber/orange |
| NOT READY | No clusters are reporting healthy status. | Solid red |
Below the overall status indicator, three per-cluster cards show the individual state of each cluster:
- Role: PRIMARY or STANDBY (indicates whether this cluster holds the VIP and is accepting writes)
- PostgreSQL status: Running/stopped, primary/standby mode, replication lag
- K3s pod health: Whether the Open WebUI pod is running
- Readiness: Per-cluster assessment of operational status
9-Device Health Grid¶
The main body of the dashboard displays a grid of cards representing all 9 devices in the system (6 K3s nodes + 3 AI servers). Each card shows:
| Metric | Source | Threshold Indicators |
|---|---|---|
| CPU Temperature | vcgencmd measure_temp (Pis) or thermal zone sysfs (Jetson) |
Green (<60 C), Yellow (60--75 C), Red (>75 C) |
| Core Voltage | vcgencmd measure_volts (Pis only) |
Nominal range check |
| Throttle/Undervoltage | vcgencmd get_throttled (Pis only) |
0x0 = healthy; any other value = warning |
| Memory Utilization | /proc/meminfo |
Color-coded by percentage |
| Disk Utilization | df output |
Color-coded by percentage |
| Device Status | Network reachability + service checks | ONLINE / OFFLINE |
Each card is clickable to expand a detail panel showing:
- Load averages (1m, 5m, 15m)
- Full memory breakdown (total, used, available, cached)
- Disk usage per mount point
- Individual service statuses (K3s, PostgreSQL, Keepalived, Ollama, etc.)
- For the Jetson: individual thermal zone temperatures (CPU, GPU, CV, SoC, junction) and power rail readings (VDD_IN, VDD_CPU_GPU_CV, VDD_SOC)
AI Server Status¶
A dedicated section shows the status of each AI inference server:
- Ollama health: Whether the Ollama HTTP API responds on port 11434
- Currently loaded model: Which model is currently in memory (via the
/api/psendpoint). Shows "No model loaded" when idle. - Server role identifier: The prefix ID (
ai1,ai2,ai3) that maps to the server in the Open WebUI model selector
PostgreSQL Replication Status¶
Displays the current replication topology and health:
- Which node is the current primary (holds the VIP)
- Replication lag for each standby (sourced from
pg_stat_replicationon the primary) - Connection state for each standby (streaming, catchup, or disconnected)
Service Status Summary¶
A consolidated view of critical systemd services across all node1 servers:
k3s-- Kubernetes runtimepostgresql-- Databasekeepalived-- Failover managementhunch-status-dashboard-- This dashboardpg-autoheal-- Auto-recovery after reboot
Data Collection¶
Local Sensor Reads¶
The dashboard reads hardware sensor data from the local node using system calls:
vcgencmdcommands for Raspberry Pi temperature, voltage, and throttle status/sys/class/thermal/thermal_zone*/tempfor Linux thermal framework readings/proc/meminfoanddffor memory and disk metricssystemctl is-activefor service status checks
Remote Device Polling¶
For devices other than the local node, the dashboard collects sensor data via SSH. Every 15 seconds, it connects to each remote device, runs the same sensor commands, and parses the results.
The 15-second polling interval is a deliberate balance: frequent enough to detect failures promptly during a live demo, but infrequent enough to minimize the SSH overhead on resource-constrained Raspberry Pi hardware.
Auto-Refresh¶
The dashboard page automatically refreshes every 5 seconds in the browser. This provides near-real-time monitoring during live demonstrations without requiring the user to manually reload the page. The 5-second browser refresh interval is independent of the 15-second backend polling interval -- the browser always displays the most recent data available from the last polling cycle.
Systemd Service¶
The dashboard runs as a systemd service on each node1:
| Parameter | Value |
|---|---|
| Service name | hunch-status-dashboard.service |
| Script location | /opt/hunch-status-dashboard/dashboard.py |
| Port | 9090 |
| Runs on | All 3 node1 servers |
| Restart policy | Automatic restart on failure |
The service starts automatically at boot and runs continuously. Since the dashboard runs on all three node1 servers, it is available regardless of which node is the current Keepalived MASTER. Accessing it via the VIP ensures that the user always reaches the dashboard instance running on the active primary.
Hardware Failure Detection¶
The dashboard's sensor monitoring enables detection of several failure conditions that are particularly relevant in a space environment:
| Condition | Detection Method | Dashboard Indication |
|---|---|---|
| Overheating | CPU temperature exceeds threshold values | Temperature badge turns yellow or red |
| Power supply issue | vcgencmd get_throttled returns non-zero value |
Throttle status shows warning flag |
| Undervoltage | Core voltage reading below expected range | Voltage reading flagged in detail panel |
| Device offline | SSH or network timeout during polling | Card shows OFFLINE status |
| Service crash | systemctl is-active returns non-active |
Service listed as stopped in detail panel |
| Memory exhaustion | Memory utilization percentage exceeds threshold | Memory badge turns yellow or red |
| Disk full | Disk utilization percentage exceeds threshold | Disk badge turns yellow or red |
Thermal monitoring in space
In a space station environment, thermal management is more complex than on Earth. Without convective cooling (no air currents in microgravity), heat dissipation relies entirely on conduction and radiation. The dashboard's temperature monitoring allows crew members to identify devices that are approaching thermal limits before they throttle or fail, enabling proactive intervention such as repositioning equipment or adjusting workloads.