Skip to content

Compute Nodes

The application layer of Astra is built on six Raspberry Pi 4 single-board computers organized into three independent Kubernetes (K3s) clusters. Each cluster contains two nodes with distinct roles, providing both intra-cluster and cross-cluster fault tolerance.


Cluster and Node Inventory

Node Role LAN IP Cluster OS
k3s-c1-node1 Server (primary) 10.0.0.11 Cluster 1 Ubuntu Server 24.04.3 LTS (ARM64)
k3s-c1-node2 Agent (worker) 10.0.0.12 Cluster 1 Ubuntu Server 24.04.4 LTS (ARM64)
k3s-c2-node1 Server (primary) 10.0.0.21 Cluster 2 Ubuntu Server 24.04.3 LTS (ARM64)
k3s-c2-node2 Agent (worker) 10.0.0.22 Cluster 2 Ubuntu Server 24.04.4 LTS (ARM64)
k3s-c3-node1 Server (primary) 10.0.0.31 Cluster 3 Ubuntu Server 24.04.3 LTS (ARM64)
k3s-c3-node2 Agent (worker) 10.0.0.32 Cluster 3 Ubuntu Server 24.04.4 LTS (ARM64)

All six nodes run Ubuntu Server 24.04 LTS and have static IP addresses on the LAN (10.0.0.0/24).


Hardware Specifications

All compute nodes use the Raspberry Pi 4 Model B:

Specification Value
Processor Broadcom BCM2711, quad-core Cortex-A72 (ARM v8), 1.8 GHz
RAM 4 GB+ LPDDR4-3200 SDRAM
Storage MicroSD (solid-state flash), 64--128 GB per node
Networking Gigabit Ethernet (primary), 2.4 GHz / 5 GHz WiFi (unused)
Power PoE+ (802.3at) via PoE HAT
GPIO 40-pin header (used for PoE HAT connection)

Node Roles and Responsibilities

Node1 (Server Nodes)

Each node1 operates as the K3s control plane server for its cluster and runs several critical services:

Service Description
K3s Server Kubernetes control plane (API server, scheduler, controller manager) and kubelet
PostgreSQL 16 Local database instance storing all Open WebUI data (users, conversations, embeddings)
Keepalived VRRP daemon managing the Virtual IP (10.0.0.50) for automatic failover
Status Dashboard Custom Python monitoring service on port 9090
pg-autoheal Systemd oneshot service that automatically rebuilds the node as a PostgreSQL standby after power loss

Node1s are the critical infrastructure components. One node1 serves as the PostgreSQL primary (writable), while the other two operate as streaming standbys (read-only replicas). The Keepalived priority determines which node1 is preferred as the primary:

Node Keepalived Priority Preferred Role
k3s-c1-node1 100 (highest) Primary
k3s-c2-node1 90 First failover target
k3s-c3-node1 80 (lowest) Last resort

Node2 (Agent Nodes)

Each node2 operates as a K3s worker agent within its cluster:

Service Description
K3s Agent Kubernetes worker (kubelet only, no control plane components)
lsyncd Live Syncing Daemon providing bidirectional file replication of /data/shared/ across all three node2s

The architecture supports scheduling pods to either node1 or node2 within a cluster. In the current prototype, pods are pinned to node1 via a Kubernetes nodeSelector as a workaround for WiFi client isolation on the home network used during development. This constraint can be removed to allow pods to run on node2, enabling intra-cluster fault tolerance with a typical recovery time of approximately 3 minutes.

Node2 also provides:

  • Additional compute capacity for future workloads beyond the primary Open WebUI application.
  • File synchronization via lsyncd, replicating /data/shared/ bidirectionally across all three node2s.

Pod scheduling

Application persistence is disabled (persistence.enabled: false in the Helm chart) so that pods can freely move between node1 and node2 within a cluster. The database runs bare metal on node1 (not in a container), so database availability is unaffected by pod scheduling decisions.


K3s Configuration

Node1 servers run K3s with the following configuration at /etc/rancher/k3s/config.yaml:

disable-network-policy: true
tls-san:
  - <additional-management-ip>
Setting Purpose
disable-network-policy: true Prevents boot loops caused by the network policy controller initialization on ARM hardware. Without this setting, K3s intermittently fails to start with a network interface error.
tls-san Adds an additional IP to the K3s API server's TLS certificate, enabling remote kubectl access for management purposes.

Agent nodes use default K3s agent configuration with no custom config.yaml required.

ARM-specific stability fix

The disable-network-policy: true setting is required on all K3s server nodes running on ARM hardware. Without it, the K3s server process enters an intermittent boot loop caused by the network policy controller failing to find the correct network interface during initialization. This was identified and resolved during development.


Storage

All compute nodes use microSD cards (solid-state flash memory) for boot and data storage. While microSD does not offer the performance of NVMe or SATA SSDs, the distributed replication architecture (PostgreSQL streaming replication across node1s, lsyncd file sync across node2s) provides data redundancy that protects against both individual media failure and entire-node loss.

See Data Redundancy for a detailed comparison of Astra's distributed replication approach versus traditional RAID. See the Network and Power section for details on the physical redundancy domains that protect these nodes, and the AI Servers section for the dedicated inference hardware.