Skip to content

Architecture Overview

Astra consists of 9 physical computing devices organized into two functional layers -- an Application Layer that serves the user-facing interface and manages data, and an AI Inference Layer that runs language model inference. Both layers are connected through a self-contained local area network that requires zero external connectivity.


System Topology

The following diagram shows the physical layout of the Astra system. A network router serves as the central network hub and WiFi access point. Three PoE switches create three independent power and network domains, each containing one complete set of application nodes and one AI inference server.

                        +-------------------+
                        |  Network Router   |
                        |  10.0.0.0/24 LAN  |
                        +--------+----------+
                                 |
                 +---------------+---------------+
                 |               |               |
          +------+------+ +-----+-------+ +-----+-------+
          | PoE Switch 1| | PoE Switch 2| | PoE Switch 3|
          | (Power &    | | (Power &    | | (Power &    |
          |  Network    | |  Network    | |  Network    |
          |  Domain A)  | |  Domain B)  | |  Domain C)  |
          +------+------+ +-----+-------+ +-----+-------+
                 |               |               |
          +------+------+ +-----+-------+ +-----+-------+
          |  Cluster 1  | |  Cluster 2  | |  Cluster 3  |
          |  node1+node2| |  node1+node2| |  node1+node2|
          |  + ai-srv1  | |  + ai-srv2  | |  + ai-srv3  |
          +-------------+ +-------------+ +-------------+

The router provides a WiFi access point (SSID: infra-net) and a wired LAN. In the current prototype, a GL.iNet OpenWrt-based router is used. In a Gateway deployment, this role would be filled by the station's internal network infrastructure.

Connected devices access the system through a single Virtual IP address (10.0.0.50), which is automatically managed by Keepalived to always point to the active primary cluster.


Functional Layers

Application Layer (6 Raspberry Pi Units)

The Application Layer consists of six Raspberry Pi 4 units organized into three independent Kubernetes (K3s) clusters. Each cluster contains two nodes:

  • Node1 (Server): Runs the K3s control plane, PostgreSQL 16 database, Keepalived failover daemon, the status dashboard, and the pg-autoheal self-repair service.
  • Node2 (Agent): Runs the K3s worker agent and lsyncd file synchronization daemon.

Each cluster operates as a fully independent Kubernetes deployment. Open WebUI (the user-facing chat application) runs as a pod within each cluster and connects to the local node1's PostgreSQL instance for data storage.

At any given time, one cluster's PostgreSQL instance is the primary (writable) and the other two are streaming standbys (read-only replicas). All user data -- accounts, conversations, settings, and AI embeddings -- is replicated across all three nodes in real time with sub-millisecond lag.

Why three independent clusters instead of one large cluster?

A single large Kubernetes cluster would be simpler to manage but creates a single point of failure at the control plane level. Three independent clusters on separate power domains ensure that losing one power domain only affects one cluster. The remaining two clusters continue operating with full data integrity.

AI Inference Layer (3 Servers)

The AI Inference Layer consists of three dedicated servers running Ollama, a lightweight AI inference engine:

Server Hardware Inference Type Performance
ai-server1 Raspberry Pi 5 (16 GB) CPU-only ~3--8 tokens/second
ai-server2 Raspberry Pi 5 (16 GB) CPU-only ~3--8 tokens/second
ai-server3 NVIDIA Jetson Orin Nano (8 GB) CUDA GPU ~25--40 tokens/second

All three AI servers are configured as separate named Ollama endpoints in every OpenWebUI cluster, each with a distinct prefix ID (ai1/, ai2/, ai3/). Users explicitly select which server to run their model on -- this is not random round-robin. This deliberate design enables running different models on different servers simultaneously without unload/load delays, explicit control over which hardware runs inference, and the ability to keep one model loaded on each server for instant switching between them. If one AI server goes offline, the remaining servers continue handling requests with no interruption to the user interface.

Separation of concerns

AI inference is deliberately separated from the application layer. The K3s nodes handle web serving, database operations, orchestration, and replication -- all of which require consistent, predictable performance. AI inference is CPU/GPU-intensive and would compete for the same resources, potentially causing database replication lag or interface slowdowns. Separate servers also provide independent failure modes: losing a K3s node does not affect AI inference, and losing an AI server does not affect the web application or database.


Network Layer

The entire system runs on a self-contained local area network with no external dependencies.

Component Role Details
Network Router Central switch + WiFi AP Subnet 10.0.0.0/24, WiFi SSID infra-net
PoE Switch 1 Power & Network Domain A Powers Cluster 1 nodes, networking for ai-server1
PoE Switch 2 Power & Network Domain B Powers Cluster 2 nodes, networking for ai-server2
PoE Switch 3 Power & Network Domain C Powers Cluster 3 nodes, networking for ai-server3
Virtual IP User entry point 10.0.0.50 -- floats to the active primary via Keepalived VRRP

The network router provides a completely self-contained LAN and WiFi access point. Connected devices access Astra at http://10.0.0.50:30080. No coordination with external IT infrastructure is required.

No internet required

The system is designed to function with zero external connectivity. All AI models, embedding caches, container images, and application code are pre-loaded on every device. This mirrors the operational reality of a deep-space mission where internet connectivity is unavailable.


Software Stack Summary

Layer Component Technology Purpose
Container Orchestration K3s Lightweight Kubernetes (v1.34.x) Pod scheduling, self-healing, service abstraction
Application Open WebUI v0.8.12 (Helm chart v13.3.1) ChatGPT-like web interface for AI interaction
AI Inference Ollama v0.20.7 Local model serving via REST API
Database PostgreSQL v16 with pgvector User data, conversations, vector embeddings
High Availability Keepalived VRRP on eth0 Virtual IP management and automatic failover
Self-Repair pg-autoheal Custom shell script + systemd Automatic standby rebuild after power loss
File Sync lsyncd Live Syncing Daemon Bidirectional file replication across node2s
Remote Management Tailscale WireGuard mesh VPN Secure remote access for ground-based development and testing only (not needed in a flight deployment)
Monitoring Status Dashboard Custom Python (zero dependencies) Real-time 9-device health monitoring

Data Flow

The following describes the path of a user request through the system:

  1. A user connects to WiFi network infra-net and navigates to http://10.0.0.50:30080.
  2. The Virtual IP (managed by Keepalived) routes the request to the active primary cluster's node1.
  3. The K3s NodePort service forwards the request to the Open WebUI pod.
  4. Open WebUI authenticates the user against the local PostgreSQL database.
  5. The user selects an AI model and sends a message.
  6. Open WebUI forwards the inference request to the appropriate Ollama server on the local network.
  7. Ollama loads the model (if not already in memory), generates a response, and streams it back.
  8. Open WebUI stores the conversation in PostgreSQL and displays the response to the user.
  9. PostgreSQL asynchronously replicates the new conversation data to the standby nodes.

If the active primary fails at any point during this flow, Keepalived promotes a standby within approximately 5 seconds and the user can resume with their full conversation history intact.