Compute Nodes¶
The application layer of Astra is built on six Raspberry Pi 4 single-board computers organized into three independent Kubernetes (K3s) clusters. Each cluster contains two nodes with distinct roles, providing both intra-cluster and cross-cluster fault tolerance.
Cluster and Node Inventory¶
| Node | Role | LAN IP | Cluster | OS |
|---|---|---|---|---|
| k3s-c1-node1 | Server (primary) | 10.0.0.11 | Cluster 1 | Ubuntu Server 24.04.3 LTS (ARM64) |
| k3s-c1-node2 | Agent (worker) | 10.0.0.12 | Cluster 1 | Ubuntu Server 24.04.4 LTS (ARM64) |
| k3s-c2-node1 | Server (primary) | 10.0.0.21 | Cluster 2 | Ubuntu Server 24.04.3 LTS (ARM64) |
| k3s-c2-node2 | Agent (worker) | 10.0.0.22 | Cluster 2 | Ubuntu Server 24.04.4 LTS (ARM64) |
| k3s-c3-node1 | Server (primary) | 10.0.0.31 | Cluster 3 | Ubuntu Server 24.04.3 LTS (ARM64) |
| k3s-c3-node2 | Agent (worker) | 10.0.0.32 | Cluster 3 | Ubuntu Server 24.04.4 LTS (ARM64) |
All six nodes run Ubuntu Server 24.04 LTS and have static IP addresses on the LAN (10.0.0.0/24).
Hardware Specifications¶
All compute nodes use the Raspberry Pi 4 Model B:
| Specification | Value |
|---|---|
| Processor | Broadcom BCM2711, quad-core Cortex-A72 (ARM v8), 1.8 GHz |
| RAM | 4 GB+ LPDDR4-3200 SDRAM |
| Storage | MicroSD (solid-state flash), 64--128 GB per node |
| Networking | Gigabit Ethernet (primary), 2.4 GHz / 5 GHz WiFi (unused) |
| Power | PoE+ (802.3at) via PoE HAT |
| GPIO | 40-pin header (used for PoE HAT connection) |
Node Roles and Responsibilities¶
Node1 (Server Nodes)¶
Each node1 operates as the K3s control plane server for its cluster and runs several critical services:
| Service | Description |
|---|---|
| K3s Server | Kubernetes control plane (API server, scheduler, controller manager) and kubelet |
| PostgreSQL 16 | Local database instance storing all Open WebUI data (users, conversations, embeddings) |
| Keepalived | VRRP daemon managing the Virtual IP (10.0.0.50) for automatic failover |
| Status Dashboard | Custom Python monitoring service on port 9090 |
| pg-autoheal | Systemd oneshot service that automatically rebuilds the node as a PostgreSQL standby after power loss |
Node1s are the critical infrastructure components. One node1 serves as the PostgreSQL primary (writable), while the other two operate as streaming standbys (read-only replicas). The Keepalived priority determines which node1 is preferred as the primary:
| Node | Keepalived Priority | Preferred Role |
|---|---|---|
| k3s-c1-node1 | 100 (highest) | Primary |
| k3s-c2-node1 | 90 | First failover target |
| k3s-c3-node1 | 80 (lowest) | Last resort |
Node2 (Agent Nodes)¶
Each node2 operates as a K3s worker agent within its cluster:
| Service | Description |
|---|---|
| K3s Agent | Kubernetes worker (kubelet only, no control plane components) |
| lsyncd | Live Syncing Daemon providing bidirectional file replication of /data/shared/ across all three node2s |
The architecture supports scheduling pods to either node1 or node2 within a cluster. In the current prototype, pods are pinned to node1 via a Kubernetes nodeSelector as a workaround for WiFi client isolation on the home network used during development. This constraint can be removed to allow pods to run on node2, enabling intra-cluster fault tolerance with a typical recovery time of approximately 3 minutes.
Node2 also provides:
- Additional compute capacity for future workloads beyond the primary Open WebUI application.
- File synchronization via lsyncd, replicating
/data/shared/bidirectionally across all three node2s.
Pod scheduling
Application persistence is disabled (persistence.enabled: false in the Helm chart) so that pods can freely move between node1 and node2 within a cluster. The database runs bare metal on node1 (not in a container), so database availability is unaffected by pod scheduling decisions.
K3s Configuration¶
Node1 servers run K3s with the following configuration at /etc/rancher/k3s/config.yaml:
| Setting | Purpose |
|---|---|
disable-network-policy: true |
Prevents boot loops caused by the network policy controller initialization on ARM hardware. Without this setting, K3s intermittently fails to start with a network interface error. |
tls-san |
Adds an additional IP to the K3s API server's TLS certificate, enabling remote kubectl access for management purposes. |
Agent nodes use default K3s agent configuration with no custom config.yaml required.
ARM-specific stability fix
The disable-network-policy: true setting is required on all K3s server nodes running on ARM hardware. Without it, the K3s server process enters an intermittent boot loop caused by the network policy controller failing to find the correct network interface during initialization. This was identified and resolved during development.
Storage¶
All compute nodes use microSD cards (solid-state flash memory) for boot and data storage. While microSD does not offer the performance of NVMe or SATA SSDs, the distributed replication architecture (PostgreSQL streaming replication across node1s, lsyncd file sync across node2s) provides data redundancy that protects against both individual media failure and entire-node loss.
See Data Redundancy for a detailed comparison of Astra's distributed replication approach versus traditional RAID. See the Network and Power section for details on the physical redundancy domains that protect these nodes, and the AI Servers section for the dedicated inference hardware.