Skip to content

Replication Guide

Overview

This guide provides sufficient detail for a future team to replicate the entire Astra system from scratch. It covers the bill of materials, the high-level build order, and the critical gotchas that can derail a build if not anticipated.

This is a replication guide, not a step-by-step tutorial

This document provides the architectural roadmap and critical warnings needed to build an equivalent system. Each step references specific technologies and configurations that have their own detailed documentation (K3s installation guides, PostgreSQL replication tutorials, Keepalived configuration references). A team replicating Astra should be prepared to consult those upstream resources for the detailed procedures within each step.

Bill of Materials

The following hardware is required to build a complete Astra system:

Qty Item Purpose Notes
6 Raspberry Pi 4 (4 GB+ RAM) K3s cluster nodes (3 servers + 3 agents) 4 GB minimum, 8 GB recommended
2 Raspberry Pi 5 (16 GB RAM) AI inference servers (CPU-based) 16 GB is required for running 5--8 GB models
1 NVIDIA Jetson Orin Nano (8 GB) AI inference server (GPU-accelerated) Developer kit recommended for carrier board with fan
3 PoE switch (4+ ports each) Independent power domains Must support 802.3af/at PoE for Pi power delivery
6 PoE HAT or PoE+ HAT Power delivery to K3s Pis via Ethernet Must match Pi model (Pi 4 HAT vs. Pi 5 HAT)
1 Compact network router LAN switch + WiFi access point OpenWrt-based recommended; must have 3+ LAN ports (or WAN-to-LAN conversion)
10+ Ethernet cables (Cat5e or better) Network wiring Include spares; a faulty cable can waste hours
9+ MicroSD cards (32 GB+ each) Boot and OS storage for each Pi 64 GB+ recommended; Class A2 preferred for better random I/O
1 DeskPi RackMate T1 (or similar) Physical enclosure Any compact rack enclosure that fits the hardware
3 Power strips Independent power feeds for each domain Each power strip powers one PoE switch

Optional but recommended:

Qty Item Purpose
1 USB keyboard + micro-HDMI adapter Initial Pi configuration (headless setup is possible but a display helps)
1 Ethernet cable tester Verify cables before installation (prevents Challenge 1)
1 Label maker Physical labeling of each device in the rack

High-Level Build Order

The following 13-step sequence represents the recommended build order. Each step builds on the previous steps, so they should be completed in order.

Step 1: Flash OS Images

Install the operating system on all computing devices:

  • K3s cluster Pis (6 units): Ubuntu Server (ARM64). Use the Raspberry Pi Imager or dd to flash the image to microSD cards.
  • AI server Pis (2 units): Ubuntu Server (ARM64) or Raspberry Pi OS (64-bit).
  • Jetson Orin Nano: NVIDIA JetPack (the official Jetson Linux distribution). Flash using the NVIDIA SDK Manager or the L4T flash tools.

Configure SSH access and a default user account on each device during the flash process.

Step 2: Configure Static IPs

Assign static IP addresses on eth0 for all devices on the network router LAN (10.0.0.0/24):

Device Static IP
Cluster 1, Node 1 10.0.0.11
Cluster 1, Node 2 10.0.0.12
Cluster 2, Node 1 10.0.0.21
Cluster 2, Node 2 10.0.0.22
Cluster 3, Node 1 10.0.0.31
Cluster 3, Node 2 10.0.0.32
AI Server 1 10.0.0.41
AI Server 2 10.0.0.42
AI Server 3 10.0.0.43

Use NetworkManager (nmcli) to set static IPs persistently. Set the wired route-metric high (e.g., 9999) so that if a device also has WiFi, the WiFi interface is preferred for default route.

Step 3: Set Up Remote Management (Optional)

For ground-based development, install a secure mesh VPN on all 9 devices. This provides stable IP addresses for remote management and cross-network database connectivity during development. This step is not required for a flight deployment where all devices are on the same physical network.

Step 4: Install K3s

Install K3s on the cluster nodes:

  • Node1s (server mode): Install K3s as a server node on each of the three node1s. Add disable-network-policy: true to /etc/rancher/k3s/config.yaml to prevent boot loops on ARM hardware. Add any additional IPs (such as management network IPs) to the tls-san list for remote kubectl access.
  • Node2s (agent mode): Install K3s as an agent on each node2, joining it to its cluster's node1 server.

Verify each cluster with kubectl get nodes -- each cluster should show two nodes (one server, one agent).

Step 5: Install PostgreSQL 16

Install PostgreSQL 16 on all three node1 servers (bare metal, not in a container):

  • Create the openwebui database and openwebui user
  • Create the replicator replication user
  • Install the pgvector extension
  • Configure one node1 as the primary (writable) and the other two as streaming standbys

Use VIP in primary_conninfo

Configure each standby's primary_conninfo to point to the VIP (10.0.0.50), not to a specific node's IP. This is critical for automatic cascade failover. See the Design Decisions documentation for the detailed rationale.

Set wal_level=replica and max_wal_senders=4 on the primary. Configure pg_hba.conf to allow replication connections from all nodes.

Step 6: Install Keepalived

Install Keepalived on all three node1 servers:

  • Configure VRRP on eth0 (physical Ethernet, not a VPN interface) with unicast peer addresses
  • Set priorities: node1-cluster1 = 100, node1-cluster2 = 90, node1-cluster3 = 80
  • Set VIP to 10.0.0.50/24
  • Create the notify script at /etc/keepalived/notify.sh that promotes PostgreSQL and restarts Open WebUI on MASTER transition
  • Add systemd overrides to ensure Keepalived starts after all required network services

Step 7: Deploy Open WebUI

Deploy Open WebUI via Helm chart on each K3s cluster:

  • Use the open-webui/open-webui Helm chart
  • Set DATABASE_URL to point to the local node1's IP address on port 5432
  • Set WEBUI_SECRET_KEY to the same value on all three clusters
  • Set ENABLE_DB_MIGRATIONS=False
  • Set HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1
  • Set ENABLE_OLLAMA_API=True
  • Configure NodePort service on port 30080
  • Disable persistence (persistence.enabled: false)

Step 8: Install Ollama

Install Ollama on all three AI servers:

  • Configure systemd overrides at /etc/systemd/system/ollama.service.d/override.conf
  • Set OLLAMA_HOST=0.0.0.0:11434 to listen on all interfaces
  • Set OLLAMA_KEEP_ALIVE=30m and OLLAMA_MAX_LOADED_MODELS=1
  • For the Jetson, add CUDA-specific environment variables
  • Pull the desired models on each server using ollama pull

Configure the Ollama base URLs in the OpenWebUI database config table using the LAN IPs of the AI servers.

Step 9: Deploy pg-autoheal

Copy the pg-autoheal.sh script and pg-autoheal.service systemd unit to each node1:

  • Script location: /opt/hunch-status-dashboard/pg-autoheal.sh
  • Ensure the script includes hostname-based stagger delays (C1=0s, C2=30s, C3=60s)
  • Ensure the script includes retry logic (3 attempts with increasing backoff)
  • Enable the systemd service to run at boot

Step 10: Deploy the Dashboard

Copy the dashboard.py script and hunch-status-dashboard.service systemd unit to each node1:

  • Script location: /opt/hunch-status-dashboard/dashboard.py
  • The dashboard listens on port 9090
  • Enable the systemd service to run at boot

Step 11: Set Up lsyncd

Install and configure lsyncd on all three node2 (agent) servers:

  • Create the /data/shared/ directory on each node2
  • Configure /etc/lsyncd/lsyncd.conf.lua for bidirectional sync between all three node2s
  • Distribute SSH keys between all node2s for passwordless rsync
  • Enable and start the lsyncd service

Step 12: Pre-Cache Offline Assets

Ensure all required assets are cached locally for offline operation:

  • Embedding models: Download and cache at /opt/owui-cache/ on all three node1s (approximately 924 MB)
  • Container images: Pull all required container images on every node (both node1s and node2s) so that K3s does not need to pull from a registry
  • Ollama models: Pull all desired models on each AI server

Do not skip this step

Without pre-cached assets, the system will attempt to download files from the internet on first use. In an offline environment, this will cause application startup failures, missing embedding functionality, and unavailable AI models. Pre-caching must be completed while internet access is available.

Step 13: Test Failover

Execute every scenario from the Tested Failover Scenarios documentation:

  • Single node failure
  • Cascade failure (kill two nodes sequentially)
  • Full power cycle (all nodes off, then all on)
  • Pod crash recovery
  • AI server failure

Do not consider the build complete until all scenarios pass.

Critical Gotchas

The following list represents hard-won lessons from the Astra build. Each item has caused significant debugging time or data loss during development. Future teams should address every item proactively.

Read all of these before starting the build

  1. VRRP does not work over VPN tunnels. VRRP uses IP protocol 112, which is silently dropped by VPN tunnels. Keepalived must use a physical Ethernet interface (eth0), not a VPN interface.

  2. Set ENABLE_DB_MIGRATIONS=False in Open WebUI Helm values. Version 0.8.12 has a peewee ORM bug that crashes the migration runner on startup. Without this flag, the application will not start.

  3. Set HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1 for offline operation. Without these flags, the application hangs on startup while attempting to download models from the internet.

  4. Share the same WEBUI_SECRET_KEY across all three clusters. Different keys cause session cookies to be invalidated during failover, forcing users to re-authenticate.

  5. Use VIP (not fixed IPs) in primary_conninfo. This is the single most important architectural decision for enabling automatic cascade failover. Standbys must point to 10.0.0.50, not to a specific node's IP address.

  6. Add stagger delays to pg-autoheal. Without stagger delays, simultaneous boot causes all nodes to run pg_basebackup at the same time, overwhelming the primary and causing all recovery attempts to fail.

  7. Do not load models larger than 5 GB on 8 GB Jetson devices. The unified memory architecture means GPU memory allocation directly reduces system RAM. A model that is too large will freeze the device, requiring a physical power cycle.

  8. Disable network policy in K3s (disable-network-policy: true in /etc/rancher/k3s/config.yaml) on ARM hardware. The network policy controller's initialization can cause boot loops on Raspberry Pi nodes.

  9. Do not enable VPN exit nodes on K3s nodes. VPN exit node mode modifies iptables rules in a way that conflicts with K3s ClusterIP routing, breaking pod-to-service communication. If an exit node is needed temporarily (e.g., for downloading packages), disable it immediately after use.

  10. Pre-cache all container images on every node before demo day. Without pre-cached images, K3s will attempt to pull images from a container registry. In an offline environment, this pull will fail, and the application pods will be stuck in ImagePullBackOff state indefinitely.