Sensors and Thermal Management¶
Astra monitors hardware health telemetry on every device in the system. All sensor data is displayed in real time on the status dashboard (port 9090) and refreshes automatically. This continuous monitoring enables early detection of thermal, power, and resource issues before they affect system availability.
Raspberry Pi Sensors (All 8 Pi Units)¶
Every Raspberry Pi in the system -- both the six Pi 4 compute nodes and the two Pi 5 AI servers -- exposes hardware sensors through the vcgencmd interface and the Linux thermal subsystem.
Available Sensor Readings¶
| Sensor | Interface | Example Output | Significance |
|---|---|---|---|
| CPU Temperature | vcgencmd measure_temp |
temp=54.5'C |
Primary indicator of thermal health. Throttling begins at 80 degrees C on Pi 4. |
| Throttle / Undervoltage | vcgencmd get_throttled |
throttled=0x0 |
Bitfield register indicating current and historical power/thermal events. A value of 0x0 indicates no issues. |
| Core Voltage | vcgencmd measure_volts |
volt=0.9260V |
Monitors power delivery stability. Abnormal readings indicate power supply degradation. |
| CPU Temperature (sysfs) | /sys/class/thermal/thermal_zone0/temp |
54530 (millidegrees) |
Alternative temperature reading via the Linux kernel thermal framework. |
Throttle Status Register¶
The throttle status register (vcgencmd get_throttled) is a critical diagnostic tool for a space application. It encodes both current conditions and historical events since the last boot in a single hexadecimal value:
| Bit | Hex Mask | Meaning |
|---|---|---|
| 0 | 0x1 |
Undervoltage detected (currently active) |
| 1 | 0x2 |
ARM frequency capped (currently active) |
| 2 | 0x4 |
Currently throttled |
| 3 | 0x8 |
Soft temperature limit active |
| 16 | 0x10000 |
Undervoltage has occurred since boot |
| 17 | 0x20000 |
ARM frequency capping has occurred since boot |
| 18 | 0x40000 |
Throttling has occurred since boot |
| 19 | 0x80000 |
Soft temperature limit has occurred since boot |
A reading of 0x0 indicates healthy operation. Any non-zero value signals a power or thermal issue that should be investigated. In a space environment, an undervoltage condition could indicate degrading power distribution hardware or insufficient current delivery under load.
NVIDIA Jetson Orin Nano Sensors (ai-server3)¶
The Jetson Orin Nano provides significantly more detailed hardware telemetry than the Raspberry Pi, with multiple thermal zones and power rail monitoring via dedicated sensor hardware.
Thermal Zones¶
The Jetson exposes nine distinct thermal zones, each corresponding to a different functional block of the system-on-chip:
| Zone | Description | Monitoring Significance |
|---|---|---|
cpu-thermal |
CPU core temperature | Indicates processor thermal load |
gpu-thermal |
CUDA GPU temperature | Critical during AI inference workloads |
cv0-thermal |
Computer vision accelerator 0 | Specialized hardware block temperature |
cv1-thermal |
Computer vision accelerator 1 | Specialized hardware block temperature |
cv2-thermal |
Computer vision accelerator 2 | Specialized hardware block temperature |
soc0-thermal |
System-on-chip zone 0 | Overall SoC thermal envelope |
soc1-thermal |
System-on-chip zone 1 | Overall SoC thermal envelope |
soc2-thermal |
System-on-chip zone 2 | Overall SoC thermal envelope |
tj-thermal |
Junction temperature | Die-level maximum temperature; the single most important thermal reading |
The junction temperature (tj-thermal) represents the highest temperature at the silicon die and is the metric most closely correlated with thermal throttling and long-term reliability. NVIDIA specifies maximum junction temperature limits for each Jetson module.
Power Rails (INA3221 Current/Voltage Monitor)¶
The Jetson includes an INA3221 three-channel current and voltage sensor that provides real-time power consumption data across the major power rails:
| Rail | Description | Nominal Voltage | Monitoring Significance |
|---|---|---|---|
VDD_IN |
Main input voltage | 5.0V | Overall power supply health |
VDD_CPU_GPU_CV |
CPU, GPU, and CV power domain | Variable | Active compute power consumption |
VDD_SOC |
System-on-chip power domain | Variable | SoC infrastructure power consumption |
The INA3221 reports voltage (mV), current (mA), and power (mW) for each rail. This data is valuable for:
- Thermal budget validation: Confirming that total power consumption stays within the thermal design power (TDP) of the cooling solution.
- Power supply adequacy: Detecting voltage droop under heavy inference loads that could indicate an undersized power supply.
- Workload characterization: Understanding the power profile of different AI models to predict battery life in a portable or space deployment.
Dashboard Integration¶
The status dashboard reads all sensor data from every device and presents it in a unified 9-device monitoring grid.
Temperature Color Coding¶
| Color | Temperature Range | Interpretation |
|---|---|---|
| Green | Below 60 degrees C | Normal operating temperature |
| Yellow | 60--75 degrees C | Elevated. Monitor for sustained high readings. |
| Red | Above 75 degrees C | Critical. Investigate cooling immediately. Approaching throttle threshold. |
Device Card Layout¶
Each device is represented as a card in the dashboard grid. The card displays:
- Device name and status (ONLINE / OFFLINE)
- CPU temperature with color-coded badge
- Memory utilization percentage
- Disk utilization percentage
- Service statuses for critical services (K3s, PostgreSQL, Keepalived, Ollama, etc.)
Cards are clickable to expand full sensor details, including:
- Load averages (1, 5, 15 minute)
- Per-service systemd status
- Throttle/undervoltage register (Pis)
- Individual thermal zone readings (Jetson)
- Power rail voltage, current, and power readings (Jetson)
Hardware Failure Detection¶
The sensor monitoring system detects the following failure conditions in real time:
| Condition | Detection Method | Dashboard Indication | Severity |
|---|---|---|---|
| Overheating | CPU temperature exceeds threshold | Temperature badge turns yellow or red | Warning / Critical |
| Power supply degradation | vcgencmd get_throttled returns non-zero value |
Throttle status shows warning flag | Warning |
| Undervoltage | Core voltage below expected range | Voltage reading flagged | Critical |
| Device offline | SSH/network timeout during health check | Card shows "OFFLINE" status | Critical |
| Service crash | systemctl is-active returns non-active state |
Service status in detail panel shows failure | Warning |
| Memory exhaustion | Memory utilization exceeds threshold | Memory badge turns yellow or red | Warning |
| Disk full | Disk utilization exceeds threshold | Disk badge turns yellow or red | Warning |
Monitoring refresh rate
The dashboard refreshes sensor data from all 9 devices every 15 seconds. The dashboard web page itself auto-refreshes every 5 seconds during demos, ensuring that any status change is visible to observers within seconds of occurring.
Thermal Management and Cooling¶
Cooling Design¶
The Astra system uses a combination of passive and active cooling appropriate for the compact DeskPi RackMate T1 enclosure:
K3s Compute Nodes (6x Raspberry Pi 4):
- Each Pi 4 is powered via a PoE+ HAT (802.3at) which includes an integrated cooling fan mounted directly above the board.
- The PoE HAT fan provides forced-air cooling across the processor and surrounding components.
- Under typical operating conditions (serving web traffic, running database replication), CPU temperatures remain well within the green range (below 60 degrees C).
AI Server 1 and 2 (2x Raspberry Pi 5):
- One Pi 5 (ai-server1) includes both passive heatsinks and an active cooling solution for sustained AI inference workloads.
- The second Pi 5 (ai-server2) currently has no heatsink or fan installed. This is a known limitation -- adding cooling hardware to this unit is a planned improvement.
- CPU-intensive inference generates more sustained heat than the intermittent loads on the K3s nodes, making adequate cooling important for maintaining inference throughput without thermal throttling.
AI Server 3 (NVIDIA Jetson Orin Nano):
- The Jetson's developer kit carrier board includes a built-in fan with PWM (Pulse Width Modulation) speed control.
- Fan speed adjusts automatically based on thermal zone temperatures -- running at low speed during idle and ramping up under sustained GPU inference load.
- GPU inference workloads generate more concentrated heat than CPU workloads, making active cooling essential for preventing thermal throttling during extended inference sessions.
Rack Airflow¶
The DeskPi RackMate T1 enclosure supports natural convective airflow through its ventilation design:
- Front panel: Closed, preventing debris ingress.
- Rear panel: Primarily open, serving as the primary heat exhaust path.
- Top and bottom panels: Ventilated, allowing vertical convective airflow. Heat rises naturally through the enclosure from bottom to top.
- Internal fans: The PoE HAT fans on each Pi 4 and the Jetson's built-in fan create localized forced-air flow across their respective boards, supplementing the passive convective flow.
Acoustic Profile¶
The cooling solution produces significantly less noise than traditional server hardware:
- PoE HAT fans on Pi 4 units are small-form-factor fans. They are audible at close range but produce a fraction of the noise of standard 1U/2U rack server fans.
- The Jetson's PWM-controlled fan runs at low speed during idle periods and ramps only under sustained GPU load, minimizing noise during typical operation.
- No high-speed server fans are present in the system. The overall acoustic output is comparable to a desktop computer at idle -- suitable for a crew habitat environment where persistent loud noise would be unacceptable.
Noise considerations for space
Acoustic noise is a documented concern aboard crewed space stations, where crew members live and work in close proximity to equipment for extended periods. The small-form-factor fans are audible at close range but produce significantly less noise than traditional rack-mounted server fans. For a crewed habitat, replacing the PoE HAT fans with lower-noise alternatives would be a recommended improvement.