Skip to content

Monitoring & Dashboard

Overview

Astra includes a custom-built status dashboard that provides real-time visibility into the health of all 9 devices, the state of PostgreSQL replication, AI server availability, and overall failover readiness. The dashboard is a single Python file with zero external dependencies -- it uses only the Python standard library -- and runs on all three node1 servers on port 9090.

The dashboard is accessible through the Keepalived VIP at http://10.0.0.50:9090, which means it automatically follows the active primary. It can also be reached directly via any node1's LAN IP on the same port.

Design Philosophy

Why a Custom Dashboard Instead of Grafana and Prometheus

Industry-standard monitoring solutions like Grafana (visualization) and Prometheus (metrics collection) are excellent tools for large-scale infrastructure. However, they introduce significant overhead that is poorly suited to Astra's resource-constrained environment:

Concern Grafana/Prometheus Astra Dashboard
Dependencies Go runtime, Node.js, multiple databases, plugin ecosystem Zero -- Python standard library only
Memory footprint Hundreds of MB for Prometheus TSDB + Grafana server Minimal -- single-threaded Python process
Storage requirements Time-series database for historical metrics None -- stateless, real-time only
Additional services to monitor Prometheus, Grafana, alertmanager, node-exporter None -- one process per node
Setup complexity Multi-service deployment with configuration files Copy one Python file and one systemd unit
Student accessibility Requires knowledge of PromQL, Grafana dashboarding Standard Python -- readable and modifiable by students

The dashboard provides exactly the information needed for operational monitoring and live demonstrations, with nothing extra. Every line of code can be understood by a student learning Python, and the entire application fits in a single file.

Accessing the Dashboard

Access Method URL
Via VIP (recommended for demos) http://10.0.0.50:9090
Direct to C1 node1 (LAN) http://10.0.0.11:9090
Direct to C2 node1 (LAN) http://10.0.0.21:9090
Direct to C3 node1 (LAN) http://10.0.0.31:9090
Direct via LAN http://10.0.0.11:9090, http://10.0.0.21:9090, http://10.0.0.31:9090

The dashboard runs identically on all three node1 servers. When accessed via the VIP, it is served by whichever node is currently the Keepalived MASTER.

Dashboard Panels

Failover Readiness Panel

The most prominent element on the dashboard is the Failover Readiness Panel at the top of the page. It provides an at-a-glance assessment of the system's ability to survive a hardware failure.

Status Meaning Visual Indicator
READY All three clusters are healthy and replicating. The system can survive the loss of any two clusters. Solid green
HEALING One or more clusters are in the process of rebuilding after a failure. The system is operational but not yet fully redundant. Pulsing amber animation
NO REDUNDANCY Only one cluster is operational. The system is serving users but cannot survive another failure. Solid amber/orange
NOT READY No clusters are reporting healthy status. Solid red

Below the overall status indicator, three per-cluster cards show the individual state of each cluster:

  • Role: PRIMARY or STANDBY (indicates whether this cluster holds the VIP and is accepting writes)
  • PostgreSQL status: Running/stopped, primary/standby mode, replication lag
  • K3s pod health: Whether the Open WebUI pod is running
  • Readiness: Per-cluster assessment of operational status

9-Device Health Grid

The main body of the dashboard displays a grid of cards representing all 9 devices in the system (6 K3s nodes + 3 AI servers). Each card shows:

Metric Source Threshold Indicators
CPU Temperature vcgencmd measure_temp (Pis) or thermal zone sysfs (Jetson) Green (<60 C), Yellow (60--75 C), Red (>75 C)
Core Voltage vcgencmd measure_volts (Pis only) Nominal range check
Throttle/Undervoltage vcgencmd get_throttled (Pis only) 0x0 = healthy; any other value = warning
Memory Utilization /proc/meminfo Color-coded by percentage
Disk Utilization df output Color-coded by percentage
Device Status Network reachability + service checks ONLINE / OFFLINE

Each card is clickable to expand a detail panel showing:

  • Load averages (1m, 5m, 15m)
  • Full memory breakdown (total, used, available, cached)
  • Disk usage per mount point
  • Individual service statuses (K3s, PostgreSQL, Keepalived, Ollama, etc.)
  • For the Jetson: individual thermal zone temperatures (CPU, GPU, CV, SoC, junction) and power rail readings (VDD_IN, VDD_CPU_GPU_CV, VDD_SOC)

AI Server Status

A dedicated section shows the status of each AI inference server:

  • Ollama health: Whether the Ollama HTTP API responds on port 11434
  • Currently loaded model: Which model is currently in memory (via the /api/ps endpoint). Shows "No model loaded" when idle.
  • Server role identifier: The prefix ID (ai1, ai2, ai3) that maps to the server in the Open WebUI model selector

PostgreSQL Replication Status

Displays the current replication topology and health:

  • Which node is the current primary (holds the VIP)
  • Replication lag for each standby (sourced from pg_stat_replication on the primary)
  • Connection state for each standby (streaming, catchup, or disconnected)

Service Status Summary

A consolidated view of critical systemd services across all node1 servers:

  • k3s -- Kubernetes runtime
  • postgresql -- Database
  • keepalived -- Failover management
  • hunch-status-dashboard -- This dashboard
  • pg-autoheal -- Auto-recovery after reboot

Data Collection

Local Sensor Reads

The dashboard reads hardware sensor data from the local node using system calls:

  • vcgencmd commands for Raspberry Pi temperature, voltage, and throttle status
  • /sys/class/thermal/thermal_zone*/temp for Linux thermal framework readings
  • /proc/meminfo and df for memory and disk metrics
  • systemctl is-active for service status checks

Remote Device Polling

For devices other than the local node, the dashboard collects sensor data via SSH. Every 15 seconds, it connects to each remote device, runs the same sensor commands, and parses the results.

The 15-second polling interval is a deliberate balance: frequent enough to detect failures promptly during a live demo, but infrequent enough to minimize the SSH overhead on resource-constrained Raspberry Pi hardware.

Auto-Refresh

The dashboard page automatically refreshes every 5 seconds in the browser. This provides near-real-time monitoring during live demonstrations without requiring the user to manually reload the page. The 5-second browser refresh interval is independent of the 15-second backend polling interval -- the browser always displays the most recent data available from the last polling cycle.

Systemd Service

The dashboard runs as a systemd service on each node1:

Parameter Value
Service name hunch-status-dashboard.service
Script location /opt/hunch-status-dashboard/dashboard.py
Port 9090
Runs on All 3 node1 servers
Restart policy Automatic restart on failure

The service starts automatically at boot and runs continuously. Since the dashboard runs on all three node1 servers, it is available regardless of which node is the current Keepalived MASTER. Accessing it via the VIP ensures that the user always reaches the dashboard instance running on the active primary.

Hardware Failure Detection

The dashboard's sensor monitoring enables detection of several failure conditions that are particularly relevant in a space environment:

Condition Detection Method Dashboard Indication
Overheating CPU temperature exceeds threshold values Temperature badge turns yellow or red
Power supply issue vcgencmd get_throttled returns non-zero value Throttle status shows warning flag
Undervoltage Core voltage reading below expected range Voltage reading flagged in detail panel
Device offline SSH or network timeout during polling Card shows OFFLINE status
Service crash systemctl is-active returns non-active Service listed as stopped in detail panel
Memory exhaustion Memory utilization percentage exceeds threshold Memory badge turns yellow or red
Disk full Disk utilization percentage exceeds threshold Disk badge turns yellow or red

Thermal monitoring in space

In a space station environment, thermal management is more complex than on Earth. Without convective cooling (no air currents in microgravity), heat dissipation relies entirely on conduction and radiation. The dashboard's temperature monitoring allows crew members to identify devices that are approaching thermal limits before they throttle or fail, enabling proactive intervention such as repositioning equipment or adjusting workloads.