Skip to content

Design Decisions

Overview

Every major architectural decision in Astra was made deliberately, with alternatives considered and trade-offs explicitly accepted. This section documents the reasoning behind each choice. Understanding these decisions is critical for evaluating the system's design, for future teams considering modifications, and for NASA engineers assessing the project's engineering rigor.

Each decision follows the same structure: the choice that was made, the alternatives that were considered, and the reasoning for why the chosen approach best serves Astra's mission requirements.


Three Independent Clusters vs. One Large Cluster

Decision: Deploy three separate, independent K3s clusters (each with 2 nodes), rather than one large 6-node Kubernetes cluster.

Alternatives considered:

  • A single 6-node Kubernetes cluster with pod anti-affinity rules to spread workloads
  • A single 3-node cluster with external database replication
  • Two clusters in an active/passive configuration

Why this choice: A single large Kubernetes cluster has a single control plane, which is itself a single point of failure. If the control plane node fails, the entire cluster loses its scheduling, self-healing, and service discovery capabilities. While Kubernetes supports multi-master (HA control plane) configurations, this requires a minimum of 3 control plane nodes with etcd consensus, consuming significant resources on Raspberry Pi hardware.

By splitting into three independent clusters, each on its own PoE switch (and therefore its own power domain), Astra achieves a stronger isolation guarantee: losing one power domain only affects one cluster. The remaining two clusters continue to operate with fully independent control planes, databases, and application instances. There is no shared state between clusters that could create a correlated failure.

The trade-off is management complexity. Three clusters means three Helm deployments, three sets of environment variables, and three Keepalived configurations that must be kept in sync. This additional operational burden is accepted because fault tolerance is the primary design goal, and the three-cluster architecture provides the strongest failure isolation achievable with the available hardware.


K3s Instead of Full Kubernetes

Decision: Use K3s (a lightweight, CNCF-certified Kubernetes distribution) instead of upstream Kubernetes (kubeadm, kops, or managed K8s).

Alternatives considered:

  • Full upstream Kubernetes via kubeadm
  • Docker Swarm (lighter-weight orchestration)
  • Direct systemd service management (no orchestrator)

Why this choice: Full upstream Kubernetes requires significantly more RAM and CPU than K3s. The control plane components (etcd, kube-apiserver, kube-controller-manager, kube-scheduler) consume several hundred megabytes of RAM each. On Raspberry Pi hardware with 4--8 GB of RAM, this overhead leaves insufficient resources for the application workload.

K3s replaces etcd with an embedded SQLite database (or external database), combines multiple control plane components into a single binary, and strips out cloud-provider-specific code. The result is the same Kubernetes API surface -- all standard tooling (kubectl, Helm, standard manifests) works identically -- but with a fraction of the resource footprint.

Docker Swarm was considered but lacks the ecosystem maturity, Helm chart support, and CNCF certification that Kubernetes provides. Direct systemd management was considered but would sacrifice automatic pod rescheduling, health checks, and rolling updates -- all of which K3s provides as standard features.


PostgreSQL Streaming Replication Instead of a Distributed Database

Decision: Use PostgreSQL 16 with native streaming replication across three nodes, rather than a distributed database system.

Alternatives considered:

  • CockroachDB (distributed SQL with built-in consensus)
  • YugabyteDB (distributed SQL, PostgreSQL-compatible)
  • SQLite with custom replication (application-level sync)

Why this choice: Distributed databases like CockroachDB and YugabyteDB provide built-in multi-node consistency through Raft or Paxos consensus protocols. However, they require a minimum of three nodes with significant RAM (4+ GB per node recommended) and CPU resources for consensus operations. On Raspberry Pi hardware, these resource requirements leave insufficient headroom for the application workload.

PostgreSQL streaming replication is lightweight, battle-tested at every scale (from single-board computers to enterprise data centers), and provides sub-millisecond replication lag on a LAN. The trade-off is that standbys are read-only -- only the primary can accept writes. Astra's Keepalived and auto-promotion system makes this transparent to users: when the primary fails, a standby is promoted within seconds, and the write restriction is lifted automatically.

SQLite with custom replication was considered but would require building a bespoke replication protocol, introducing significant development effort and a large surface area for bugs. PostgreSQL's streaming replication is a mature, well-documented protocol that has been in production use for over a decade.


VIP-Based primary_conninfo Instead of Fixed IPs

Decision: Configure PostgreSQL standbys to replicate from the VIP (10.0.0.50) rather than from a specific node's fixed IP address.

Alternatives considered:

  • Standbys point to a fixed primary IP (e.g., 10.0.0.11)
  • Application-level failover logic that reconfigures standbys after promotion
  • DNS-based failover with short TTL

Why this choice: If standbys pointed to a fixed primary IP, any failover event would require reconfiguring every standby's primary_conninfo to point to the new primary's IP address. This reconfiguration must happen quickly and reliably -- exactly the kind of manual intervention that Astra is designed to eliminate.

By pointing primary_conninfo to the VIP, standbys automatically follow whichever node currently holds the VIP. When Keepalived promotes a new MASTER and assigns it the VIP, all remaining standbys reconnect to the new primary without any configuration change. This is the key enabler of zero-intervention cascade failover: the system can lose two nodes sequentially, and the last survivor has all data, with no human touching any configuration file.

DNS-based failover was rejected because DNS TTL caching introduces propagation delays, and the system operates on an air-gapped LAN without a DNS server (all connections use IP addresses directly).


Ollama Instead of vLLM or Other Inference Engines

Decision: Use Ollama as the AI inference engine on all three AI servers.

Alternatives considered:

  • vLLM (high-throughput inference engine)
  • llama.cpp directly (raw inference runtime)
  • TensorRT-LLM (NVIDIA-optimized inference)

Why this choice: Ollama provides the simplest deployment path across Astra's heterogeneous hardware: ARM CPU (Raspberry Pi 5) and NVIDIA CUDA GPU (Jetson Orin Nano). It handles model downloading, quantization format management, memory allocation, and serving via a clean HTTP API -- all through a single binary with minimal configuration.

vLLM offers higher throughput for concurrent users but requires more complex setup, does not support ARM CPU inference, and is designed for GPU-heavy deployments with dedicated VRAM. TensorRT-LLM is NVIDIA-specific and does not run on the Raspberry Pi 5 units at all. Running llama.cpp directly would work on both platforms but would require building a custom HTTP API layer, model management system, and memory management logic -- all of which Ollama provides out of the box.

For a system with heterogeneous hardware and a priority on reliability over throughput, Ollama's approach of simplicity and broad hardware support is the correct engineering choice.


Separate AI Servers Instead of Co-Located Inference

Decision: Run AI inference on three dedicated servers rather than on the K3s cluster nodes.

Alternatives considered:

  • Running Ollama as a pod within each K3s cluster
  • Running Ollama as a sidecar container alongside Open WebUI
  • Running inference directly on the node1 servers (bare metal, alongside K3s)

Why this choice: Separation of concerns. The K3s nodes handle the web application, database, container orchestration, Keepalived failover, and replication -- all of which require consistent, predictable performance. AI inference is CPU/GPU-intensive and highly variable in resource consumption (a large model can consume several gigabytes of RAM for minutes, then release it).

Co-locating inference with the application stack creates resource contention: a large model load could starve PostgreSQL of RAM, cause replication lag, or make the web interface unresponsive. Separating the AI layer onto dedicated hardware provides independent failure modes. Losing a K3s node does not affect AI inference capability. Losing an AI server does not affect the web application, database, or failover system. This independence simplifies troubleshooting and improves overall system resilience.


Compact Network Router as Network Foundation

Decision: Use a compact network router as the central network switch and WiFi access point.

Alternatives considered:

  • A full-size managed switch with a separate WiFi access point
  • A consumer home router
  • Direct node-to-node networking with no central switch

Why this choice: The system must be portable and self-contained, capable of being demonstrated in any location without depending on existing network infrastructure. A compact network router combines a multi-port switch and WiFi AP in a device small enough to fit inside the rack enclosure. In a Gateway deployment, the router would be replaced by the station's internal network switch.

A full-size managed switch would provide more ports and advanced features (VLANs, QoS) but is physically too large for the Mini Rack form factor and adds weight. A consumer home router would work functionally but is also too large. Direct node-to-node networking without a central switch would require a mesh topology with complex routing, and would not provide WiFi access for demo attendees.

The router's OpenWrt firmware provides the configuration flexibility needed (static DHCP, port conversion, firewall rules) while maintaining a minimal physical footprint.


Disabled Database Migrations

Decision: Set ENABLE_DB_MIGRATIONS=False in all Open WebUI deployments.

Alternatives considered:

  • Running migrations on every startup (the default behavior)
  • Running migrations manually during deployment
  • Patching the upstream code to fix the bug

Why this choice: Open WebUI v0.8.12 contains a bug in its db.py module where a # db = None line is commented out, causing the peewee ORM migration system to crash on startup. The crash prevents the application from starting at all.

Disabling migrations is safe because the database schema is already correct -- all necessary migrations ran successfully during the initial setup. The schema is stable at v0.8.12 and does not change between pod restarts.

Patching the upstream code was considered but introduces a maintenance burden: every future upgrade would require re-applying the patch or verifying that the fix is included upstream. Disabling migrations is the simplest and most reliable workaround.


Offline Operation Flags

Decision: Set HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1 in all Open WebUI deployments.

Alternatives considered:

  • Running without these flags and accepting occasional download attempts
  • Blocking outbound network access at the firewall level
  • Using a local HuggingFace model mirror

Why this choice: The HuggingFace and Transformers Python libraries have a built-in behavior of checking for newer model versions and downloading updates on every startup. In an offline environment, these network requests either hang indefinitely (blocking application startup) or crash with network errors.

Setting these environment variables is the cleanest solution: it tells the libraries to use only locally cached files (pre-loaded at /opt/owui-cache/), with no network activity whatsoever. Firewall-level blocking would work but is harder to debug when issues arise, and a local mirror adds complexity and resource requirements.


Custom Dashboard Instead of Grafana/Prometheus

Decision: Build a custom status dashboard in pure Python with zero external dependencies.

Alternatives considered:

  • Grafana with Prometheus for metrics collection and visualization
  • Netdata (lightweight agent-based monitoring)
  • Simple shell scripts with text output

Why this choice: Grafana and Prometheus are industry-standard monitoring tools, but they require significant resources: a Go runtime for Prometheus, a Node.js/Go runtime for Grafana, a time-series database for metrics storage, and configuration for scrape targets, dashboards, and alert rules. On Raspberry Pi hardware, these resource requirements compete directly with the application workload.

Astra's dashboard is a single Python file that uses only the Python standard library. It has zero dependencies to install, zero databases to manage, zero services to configure beyond the dashboard itself. It can be understood and modified by students learning Python, and it provides exactly the monitoring information needed for operations and demonstrations -- nothing more, nothing less.

Netdata was considered as a middle ground but still requires agent installation on every node and a central aggregation point. Simple shell scripts would work for operators but do not provide a browser-accessible visual dashboard suitable for live demonstrations.