Skip to content

Cluster Topology

Design Philosophy: Three Independent Clusters

Astra uses three independent K3s clusters rather than a single large Kubernetes cluster. This is a deliberate architectural decision driven by the project's primary design goal: fault tolerance through physical isolation.

Why Not One Large Cluster

A single Kubernetes cluster with six nodes would be simpler to manage -- one Helm deployment, one control plane, one set of configurations. However, it introduces a critical vulnerability: the control plane is a single point of failure. If the server node hosting the control plane fails, the entire cluster loses its ability to schedule pods, manage services, and respond to failures.

Kubernetes does support multi-master (highly available) control planes, but they require a minimum of three server nodes running etcd with a quorum-based consensus protocol. This consumes significant resources on each server node and adds complexity that is difficult to justify on Raspberry Pi hardware.

The Three-Cluster Approach

By splitting the hardware into three independent clusters, each mapped to a separate physical power domain (PoE switch), Astra achieves a property that is impossible with a single cluster: the failure of any power domain affects only one cluster. The remaining two clusters continue operating with full functionality.

Power Domain A          Power Domain B          Power Domain C
(PoE Switch 1)          (PoE Switch 2)          (PoE Switch 3)
================        ================        ================
  Cluster 1               Cluster 2               Cluster 3
  k3s-c1-node1            k3s-c2-node1            k3s-c3-node1
  k3s-c1-node2            k3s-c2-node2            k3s-c3-node2
  ai-server1              ai-server2              ai-server3

The trade-off is management complexity: three separate Helm deployments, three sets of configurations, and cross-cluster coordination via PostgreSQL replication and Keepalived. The Astra team accepts this trade-off because fault tolerance is the overriding design requirement.

Cluster Composition

Each cluster consists of exactly two nodes:

Server Node (node1)

The server node runs the K3s control plane and is the cluster's primary compute resource. Each server node hosts:

Component Purpose
K3s control plane API server, scheduler, controller manager
K3s kubelet Container runtime and pod management
PostgreSQL 16 Database for OpenWebUI (user data, conversations, embeddings)
Keepalived VIP management and VRRP failover
Status dashboard Real-time health monitoring (port 9090)
pg-autoheal Automatic standby recovery after power loss

Agent Node (node2)

The agent node acts as a worker that can run pods scheduled by the server's control plane. Each agent node hosts:

Component Purpose
K3s kubelet Container runtime and pod management
lsyncd Bidirectional file synchronization across all node2s

Agent nodes provide several important capabilities:

  • Pod scheduling flexibility: If the server node is under heavy load or fails, K3s can schedule the OpenWebUI pod on the agent node instead. The architecture fully supports scheduling to either node.
  • Additional compute capacity: Agent nodes provide a second pool of CPU and memory resources within each cluster.
  • File synchronization: Each agent node runs lsyncd for bidirectional file synchronization across all three clusters, providing RAID-like redundancy for shared files (see Data Redundancy).
  • Future workload expansion: Additional services or workloads can be deployed to agent nodes without competing for resources with the database and control plane on the server node.

Node Inventory

Node Cluster Role LAN IP
k3s-c1-node1 Cluster 1 Server 10.0.0.11
k3s-c1-node2 Cluster 1 Agent 10.0.0.12
k3s-c2-node1 Cluster 2 Server 10.0.0.21
k3s-c2-node2 Cluster 2 Agent 10.0.0.22
k3s-c3-node1 Cluster 3 Server 10.0.0.31
k3s-c3-node2 Cluster 3 Agent 10.0.0.32

Application Namespace

All OpenWebUI application workloads run in the open-webui Kubernetes namespace. This namespace is created by the Helm chart during initial deployment and keeps application pods logically separated from K3s system components in kube-system.

# List all pods in the application namespace
sudo KUBECONFIG=/etc/rancher/k3s/k3s.yaml kubectl get pods -n open-webui

Service Type: NodePort

OpenWebUI is exposed as a NodePort service on port 30080. A NodePort service makes the application accessible on a static port on every node in the cluster. This means:

  • Users can reach OpenWebUI at http://<any-node-ip>:30080
  • The Keepalived VIP (10.0.0.50) routes to port 30080 on whichever node1 is the current MASTER
  • No external load balancer or ingress controller is required

Why NodePort instead of LoadBalancer or Ingress

LoadBalancer services require a cloud provider or MetalLB. Ingress controllers add complexity and resource consumption. NodePort is the simplest approach that works on bare-metal ARM hardware with no external dependencies, and it integrates cleanly with Keepalived VIP failover.

Persistence Configuration

Persistent volume storage is disabled in the OpenWebUI Helm deployment (persistence.enabled: false). This is a critical configuration choice that enables pod mobility.

Why Persistence Is Disabled

With persistent volumes enabled, a pod is bound to the specific node where its volume resides. If that node fails, the pod cannot be rescheduled to the other node because the volume is not available there.

By disabling persistence:

  • The OpenWebUI pod can run on either node1 or node2 within the cluster
  • If one node fails, K3s can reschedule the pod to the surviving node
  • All persistent data (conversations, settings, user accounts) is stored in PostgreSQL, not in the pod's local filesystem
  • Embedding model caches are mounted as read-only hostPath volumes from /opt/owui-cache/ (pre-cached on all node1s)

Pod data is ephemeral

With persistence disabled, any files written inside the pod's container filesystem are lost when the pod restarts. This is by design -- all important data is stored in the PostgreSQL database, which is replicated across all three clusters.

Pod Scheduling: nodeSelector

OpenWebUI pods are currently pinned to node1 (the server node) using a Kubernetes nodeSelector. This is a prototype workaround for a specific networking constraint in the current demo configuration, not an architectural limitation.

Current Constraint

In the current prototype, demo users connect via the network router's WiFi. On certain consumer-grade routers, WiFi clients cannot reach devices on wired LAN ports that are behind a different internal bridge. The node1 servers have static IPs on the wired LAN and are always reachable from WiFi clients. Agent nodes (node2s) may have connectivity issues depending on the router's bridge configuration.

By constraining pods to node1 via nodeSelector, the system guarantees that the OpenWebUI pod is always reachable from WiFi clients connecting through the VIP.

Architecture Supports Full Mobility

The underlying architecture fully supports scheduling pods to either node1 or node2 within a cluster. Persistent volume storage is disabled (all data is in PostgreSQL on node1), and embedding model caches are pre-cached on all nodes. Removing the nodeSelector constraint would enable full pod mobility between nodes, allowing K3s to reschedule pods to the surviving node if either node fails.

In a production deployment with proper network routing (such as a space station's internal network infrastructure), the nodeSelector constraint would not be needed.

Failure Modes and Recovery

Intra-Cluster Failure (Pod Level)

If the OpenWebUI pod crashes on one node, K3s restarts it automatically. If the pod's node becomes unavailable, K3s reschedules to the other node (subject to the nodeSelector constraint). Recovery time is approximately 3 minutes.

Cross-Cluster Failure (Power Domain Level)

If an entire power domain fails (PoE switch unplugged), the affected cluster goes offline completely. Keepalived detects the failure within 1-2 seconds and promotes a surviving cluster. Users are redirected to the new active cluster via the VIP with minimal disruption.

See the Cascade Failover and Keepalived documentation for the full cross-cluster failover mechanism.