Challenges & Solutions¶
Overview¶
Building Astra involved solving numerous technical challenges across networking, database management, memory constraints, authentication systems, and distributed systems coordination. This section documents the eight most significant challenges encountered during development, including the investigation process, root cause analysis, and the solutions applied.
These challenges are documented for future teams who may encounter similar issues when building distributed systems on resource-constrained hardware.
Challenge 1: Network Router WiFi Client Isolation¶
Problem¶
After initial assembly of the hardware rack, WiFi clients connected to the GL.iNet network router could not reach any wired LAN devices (the Raspberry Pi nodes). Devices connected via WiFi could see the router itself but could not ping or access any of the PoE-connected computing nodes. The system appeared completely non-functional to anyone connecting via WiFi.
Investigation¶
Hours of debugging were spent on software configuration: changing subnets, flushing iptables rules on the router, modifying bridge interface settings, converting the WAN port to a LAN port, and performing multiple router reboots. All software configurations appeared correct -- the routing tables showed the right entries, the bridge interface included both WiFi and LAN interfaces, and the firewall rules were permissive.
Root Cause¶
A faulty Ethernet cable between the router and the first PoE switch. The cable had an intermittent connection that passed basic link detection (the link LED illuminated) but dropped packets under any meaningful load. The physical layer was broken while the software layer appeared healthy.
Solution¶
Replaced the faulty Ethernet cable with a known-good cable. WiFi-to-LAN connectivity worked immediately.
Lesson Learned¶
Always verify the physical layer before debugging software. Hours of configuration changes are wasted when the underlying hardware connection is broken. This principle is especially relevant for a space-rated system, where physical connections must be verified, tested, and secured before any software commissioning begins.
Challenge 2: Simultaneous pg_basebackup Failure (Thundering Herd)¶
Problem¶
When all three nodes booted simultaneously after a full power cycle, each node's pg-autoheal script attempted to run pg_basebackup from the primary at the same time. The primary's WAL sender process could not handle three simultaneous full-database copies. All three basebackup operations failed, leaving the standby nodes with wiped data directories (deleted in preparation for the basebackup) and no automatic recovery path.
Investigation¶
Log analysis revealed that all three pg-autoheal instances detected that they were not the primary and immediately attempted recovery in parallel. The primary's max_wal_senders setting (4) was sufficient for concurrent connections, but the I/O contention from three simultaneous full-database copies overwhelmed the microSD storage on the primary node.
Root Cause¶
A thundering herd problem: multiple recovery processes triggering identical resource-intensive operations at the same instant, overwhelming the shared resource (the primary's storage subsystem).
Solution¶
Implemented hostname-based stagger delays in the pg-autoheal script:
| Node | Stagger Delay |
|---|---|
k3s-c1-node1 |
0 seconds |
k3s-c2-node1 |
30 seconds |
k3s-c3-node1 |
60 seconds |
Additionally, retry logic was added: up to 3 attempts with increasing wait times (30s, 60s, 90s between retries). After each stagger or retry delay, the script re-checks whether conditions have changed (for example, whether the node was promoted during the delay) before proceeding. This ensures that only one node at a time runs pg_basebackup.
Lesson Learned¶
Distributed systems must account for thundering herd problems. When multiple nodes recover simultaneously after a shared failure event, they must coordinate or stagger their recovery actions to avoid overwhelming shared resources. This is a fundamental principle of distributed systems design.
Challenge 3: Jetson Out-of-Memory with Large Models¶
Problem¶
Loading gemma4:e2b (7.2 GB) on the NVIDIA Jetson Orin Nano (8 GB unified RAM) consumed all available memory. The Linux OOM (Out Of Memory) killer could not recover the system because the CUDA driver held the allocated memory in a way that prevented clean reclamation. The device froze completely, requiring a physical power cycle to recover.
Investigation¶
The Jetson Orin Nano uses a unified memory architecture: the 8 GB of RAM is shared between the CPU and GPU. When Ollama loaded the 7.2 GB model into GPU memory via CUDA, only approximately 800 MB remained for the operating system, kernel, systemd services, and the Ollama process itself. The CUDA driver's memory allocations are not subject to normal Linux memory management, so the OOM killer could not free the GPU-held memory.
Root Cause¶
The model size (7.2 GB) exceeded the safe operating envelope for a device with 8 GB of unified (shared CPU/GPU) memory. Unlike discrete GPU systems where GPU VRAM is separate from system RAM, the Jetson's unified architecture means that GPU memory allocation directly reduces available system memory.
Solution¶
Added Ollama configuration to enforce conservative memory usage. OLLAMA_MAX_LOADED_MODELS=1 (only one model in memory at any time) is set on all three servers to prevent out-of-memory conditions on any hardware. Additionally, Jetson-specific settings were applied:
OLLAMA_CONTEXT_LENGTH=3072-- Reduced context window to lower per-request memory usageOLLAMA_GPU_OVERHEAD=1073741824-- Reserve 1 GB for the OS and CUDA runtime
Operational rule established: only models 5 GB or smaller are loaded on the Jetson. The Pi 5 units (16 GB RAM) are used for larger models.
Lesson Learned¶
On devices with unified memory (shared CPU/GPU RAM), memory management must be significantly more conservative than on devices with dedicated GPU VRAM. The operating system, drivers, and runtime overhead must always have a guaranteed memory reservation that cannot be consumed by application workloads. Always leave headroom.
Challenge 4: OpenWebUI Authentication Bypass Breaking Login¶
Problem¶
To simplify demonstrations, WEBUI_AUTH was set to False to bypass the authentication screen. However, with existing user accounts already in the database, this caused a 403 Forbidden error on the signup endpoint. The application displayed a dead onboarding screen that could not be dismissed or bypassed, rendering it completely unusable.
Investigation¶
The OpenWebUI codebase checks for existing users when WEBUI_AUTH=False. If users exist, it enters an inconsistent state: it skips the login form (because auth is disabled) but also blocks the signup form (because users already exist and it treats the system as already configured). The result is an unrecoverable dead screen.
Root Cause¶
Disabling authentication in an application designed with authentication as a core assumption creates unexpected state conflicts. The auth-disabled code path was not designed for environments where users already exist in the database.
Solution¶
Kept WEBUI_AUTH=True (authentication enabled) and instead reset the demo user's password via direct database update. The bcrypt password hash was generated inside the OpenWebUI container (to ensure the correct library version and hash format), then applied via a psql UPDATE statement on the primary database.
Lesson Learned¶
Disabling authentication in an application designed for authentication often creates unexpected state conflicts. It is typically safer and more reliable to manage credentials properly (reset passwords, create demo accounts) than to bypass the authentication system entirely.
Challenge 5: VRRP Protocol Compatibility¶
Problem¶
The initial Keepalived configuration attempted to use a VPN tunnel interface for VRRP communication between nodes. However, VRRP never functioned: no node ever transitioned to MASTER, and the VIP was never assigned.
Investigation¶
Network packet captures showed VRRP packets being sent but never received by peer nodes. The packets appeared to leave the local node but never arrived at their destinations. Standard TCP/UDP traffic between nodes worked correctly, ruling out general connectivity issues.
Root Cause¶
VRRP uses IP protocol number 112 -- it is not a TCP or UDP application. It operates directly at the IP layer, similar to ICMP (protocol 1) or OSPF (protocol 89). VPN tunnels typically only transport TCP, UDP, and ICMP traffic. All VRRP packets sent over the VPN tunnel were silently dropped.
Solution¶
Switched Keepalived to the eth0 interface on the network router LAN, using unicast peer addresses (10.0.0.11, 10.0.0.21, 10.0.0.31). VRRP functions correctly over standard Ethernet because Ethernet does not filter by IP protocol number.
Lesson Learned¶
Not all network protocols work over VPN tunnels. VRRP, OSPF, and other routing protocols that use raw IP protocol numbers (not TCP/UDP port numbers) require direct Layer 2 or Layer 3 connectivity. VPN tunnels that only encapsulate TCP, UDP, and ICMP will silently drop these protocols. Always verify protocol compatibility before assuming any tunnel can carry arbitrary traffic.
Challenge 6: Shell Escaping Bcrypt Hashes via SSH¶
Problem¶
When updating user passwords in PostgreSQL via remote SSH, the bcrypt hash string (which contains dollar signs, e.g., $2b$12$...) was corrupted before reaching the database. The stored hash did not match the expected value, and users could not log in with the new password.
Investigation¶
The bcrypt hash format uses dollar signs ($) as field separators. When passed through an SSH command line, the shell on the remote side interpreted these dollar signs as variable references. For example, $2b was interpreted as the value of a shell variable named 2b (which is empty), causing portions of the hash to be silently removed.
Root Cause¶
Multi-layer shell escaping. The bcrypt hash passed through at least two shell interpretation layers: the local shell (on the operator's laptop), the SSH transport, and the remote shell (on the target node). Each layer performs its own variable expansion and special character interpretation, and the dollar signs in the bcrypt hash were mangled during this process.
Solution¶
Generated the bcrypt hash and executed the database UPDATE statement in the same remote command, avoiding the need to pass the hash through multiple shell layers. By using a single-quoted heredoc or generating the hash within the remote shell context, variable expansion was prevented entirely.
Lesson Learned¶
When passing strings containing $, backticks, or other shell metacharacters through SSH, always generate and consume the string in the same shell context. Cross-shell string passing through SSH is a common source of escaping bugs, especially with strings that contain characters that are meaningful to the shell (dollar signs, backticks, exclamation marks, double quotes).
Challenge 7: Obsolete pg-bridge Service Causing False Alerts¶
Problem¶
The pg-bridge service (a socat relay bridging the Kubernetes pod network at 10.42.0.1:5433 to localhost:5432) was originally needed when OpenWebUI pods connected to PostgreSQL via the pod network. After switching DATABASE_URL to use stable overlay network IPs directly, pg-bridge became redundant. However, it continued running (and failing at boot, before the K3s CNI network was ready), generating false alerts in the dashboard and polluting logs.
Investigation¶
Dashboard monitoring showed pg-bridge repeatedly failing and restarting. Log analysis confirmed that the failures were expected -- the service was trying to bind to a pod network address that did not yet exist during early boot. The alerts were false positives because the service was no longer needed for any functional purpose.
Root Cause¶
An obsolete service that was not removed when its purpose was eliminated. The service continued to run because systemd was configured to start it, and continued to be monitored because the dashboard was configured to check it.
Solution¶
Stopped and disabled pg-bridge on all nodes. Removed pg-bridge from the dashboard's monitoring configuration.
Lesson Learned¶
Remove obsolete services promptly when their purpose is eliminated. Dead services that fail at boot create noise in monitoring systems and can mask real problems. In an operational system, every monitored service should serve an active purpose -- otherwise, false alerts erode confidence in the monitoring system itself.
Challenge 8: Prefix ID Configuration Not Persisting After Failover¶
Problem¶
The OLLAMA_API_CONFIGS setting (which maps human-readable prefix identifiers like ai1/, ai2/, ai3/ to specific AI server URLs) was loaded in-memory per OpenWebUI pod instance. After a Keepalived failover event, the newly activated pod on the promoted cluster did not have this configuration. Instead of displaying friendly model names like ai3.qwen3:4b, the interface showed raw Ollama URLs.
Investigation¶
The prefix ID configuration was initially set via environment variables or the in-app admin panel, both of which result in the configuration being stored in the pod's process memory. When a pod restarts (either from a crash or a failover-triggered restart), the in-memory configuration is lost.
Root Cause¶
Ephemeral pod memory is not a suitable storage location for configuration in a system with failover. Configuration that exists only in a running pod's memory is lost whenever that pod restarts, and the replacement pod starts with default settings.
Solution¶
The configuration is now stored in the PostgreSQL config table, which is replicated to all nodes via streaming replication. When a new pod starts (on any cluster), it reads the configuration from the database on startup. The prefix ID mappings, Ollama base URLs, and API configurations all persist across pod restarts and failovers.
If the configuration ever needs to be re-applied manually, it can be set via a POST request to the /ollama/config/update API endpoint.
Lesson Learned¶
In a system with failover, all configuration must be stored in replicated persistent storage (the database) or in durable configuration files that are present on every node. Ephemeral pod memory is inherently unreliable in any system where pods can restart or be rescheduled. This principle applies broadly to any Kubernetes-based application that may fail over between nodes.