Skip to content

Open WebUI Application Layer

Overview

Open WebUI (v0.8.12) is the user-facing application that provides Astra's AI medical assistant interface. It is an open-source, self-hosted web application that presents a familiar chat interface -- similar in appearance and interaction model to commercial AI assistants -- while running entirely on local hardware with no cloud dependency.

Open WebUI serves as the bridge between the astronaut and the AI inference backends. It handles user authentication, conversation management, model selection, and presentation of AI responses, while delegating the actual inference work to the Ollama servers running on the dedicated AI backend hardware.

Core Features

Chat Interface

The primary interface is a multi-turn conversational chat window. Users can:

  • Send text queries and receive AI-generated responses
  • Maintain conversation history across sessions (stored in PostgreSQL)
  • Select from multiple AI models with different capabilities
  • Switch between models mid-conversation
  • Use voice input (with browser support -- see Voice Mode below)

Multi-Backend AI Support

Open WebUI connects to multiple AI inference backends simultaneously:

  • Ollama backends (local): Three Ollama servers on the LAN provide offline AI inference with medical and general-purpose models
  • OpenAI-compatible APIs (cloud): An OpenRouter endpoint provides access to cloud-hosted models when internet connectivity is available

The Ollama backend URLs are stored in the PostgreSQL config table and are loaded by the application on startup. This database-backed configuration ensures that the Ollama endpoint settings survive pod restarts and failovers, because the database is replicated across all three clusters.

Authentication and Session Management

User authentication is enabled (WEBUI_AUTH=True). All users must log in with credentials stored in the PostgreSQL database. Passwords are hashed using bcrypt.

All three clusters share the same WEBUI_SECRET_KEY, which ensures that session cookies remain valid during failover. When the active cluster changes (due to a Keepalived failover), the user's browser still holds a valid session cookie, and the new cluster accepts it without requiring re-authentication.

Retrieval-Augmented Generation (RAG)

Open WebUI supports RAG through the pgvector PostgreSQL extension. The pgvector extension is installed on all database nodes, and the platform's RAG pipeline is supported. Documents could be uploaded, chunked, and embedded into vector representations stored in the database. When a user asks a question, the system could retrieve relevant document chunks and include them in the prompt context, allowing the AI model to ground its responses in specific reference material.

RAG status in current deployment

The pgvector extension is installed and RAG is supported by the Open WebUI platform, but document upload and RAG workflows have not been configured in the current deployment. This is a capability of the platform that could be enabled to allow the AI to reference uploaded medical protocols, drug interaction databases, or crew-specific medical histories.

Deployment Architecture

Helm Chart

Open WebUI is deployed on each K3s cluster using the official Helm chart:

Parameter Value
Chart open-webui/open-webui
Chart version v13.3.1
Application version v0.8.12
Namespace open-webui
Service type NodePort
NodePort 30080

The Helm chart is deployed independently on each of the three K3s clusters. Each deployment is a self-contained instance that connects to its own cluster's local PostgreSQL database.

Database Connection

Each cluster's Open WebUI deployment connects to its own node1's PostgreSQL instance:

Cluster DATABASE_URL Target
Cluster 1 k3s-c1-node1, port 5432
Cluster 2 k3s-c2-node1, port 5432
Cluster 3 k3s-c3-node1, port 5432

Ground-development configuration

In the current ground-based prototype, the DATABASE_URL uses stable overlay network IPs that provide cross-network reachability during development and testing. In a flight deployment, standard LAN IPs would be used since all devices would be on the same physical network.

Persistence Configuration

Persistent volume storage is disabled (persistence.enabled: false) in the Helm values. This is a deliberate choice: it allows Open WebUI pods to be freely rescheduled between node1 and node2 within a cluster without being tied to a specific node's local storage.

All persistent data (users, conversations, settings, embeddings) is stored in PostgreSQL, which runs independently of the application pod. File uploads that need to be shared across nodes are handled by the lsyncd file synchronization system running on the node2 agents.

Key Configuration Decisions

Disabled Database Migrations

ENABLE_DB_MIGRATIONS=False

Open WebUI v0.8.12 contains a known bug in its database migration system. The peewee ORM's db.py module has a commented-out line (# db = None) that causes the migration runner to crash on startup. Disabling migrations is a safe workaround because the database schema is already correct (all migrations ran successfully during initial setup). This flag must remain False until the upstream bug is fixed in a future release.

Offline Operation Flags

HF_HUB_OFFLINE=1
TRANSFORMERS_OFFLINE=1

Without these flags, the HuggingFace and Transformers Python libraries attempt to download model weights, tokenizers, and configuration files from the internet on every application startup. In an offline environment, these download attempts either hang indefinitely (blocking the application from starting) or crash with network errors.

Setting these environment variables instructs the libraries to use only locally cached files. The required embedding model cache is pre-loaded at /opt/owui-cache/ (924 MB) on every node1 and mounted into the pod as a hostPath volume.

Shared Secret Key

WEBUI_SECRET_KEY=<same value on all 3 clusters>

All three clusters use the same WEBUI_SECRET_KEY. This value is used to sign and verify session cookies. By sharing the key across clusters, a session cookie issued by one cluster is valid on any other cluster. This ensures that users are not forced to re-authenticate after a Keepalived failover transitions traffic to a different cluster.

Pre-Cached Embedding Models

The sentence-transformer models used for RAG embeddings are pre-downloaded and cached at /opt/owui-cache/ on all three node1 servers. This directory is mounted into the Open WebUI pod as a hostPath volume. The total cache size is approximately 924 MB.

Without this cache, the application would attempt to download the embedding models from HuggingFace on first use -- which would fail in an offline environment.

Pod Scheduling

Open WebUI pods are configured with a nodeSelector that pins them to node1 within each cluster. This is a prototype workaround for a WiFi client isolation issue where clients connected to the router WiFi cannot reach pods scheduled on node2. The architecture supports scheduling to either node, and this constraint would be removed in a production deployment with proper network routing.

Voice Mode

Open WebUI supports voice input (speech-to-text) for hands-free interaction with the AI. This is particularly relevant for a medical scenario where an astronaut may need to interact with the system while their hands are occupied with a patient.

Voice interaction requires a secure HTTPS connection. In the current prototype, this is achieved using a Chrome browser flag (chrome://flags/#unsafely-treat-insecure-origin-as-secure) to treat the HTTP endpoint as a secure origin for microphone access. In a production deployment, a proper TLS certificate would be installed on the router to enable voice mode natively across all browsers.

Browser compatibility

Safari does not support microphone access over HTTP under any configuration. In the current prototype, voice mode requires Chrome with the insecure-origin flag. A production deployment with HTTPS would enable voice mode in all modern browsers.

Cloud Fallback (OpenRouter)

When internet connectivity is available, Open WebUI can route requests to cloud-hosted AI models through OpenRouter, an OpenAI-compatible API aggregator. This provides access to large frontier models that cannot run on Astra's local hardware.

The OpenRouter endpoint is configured as an OpenAI-compatible backend:

Parameter Value
Base URL https://openrouter.ai/api/v1
Protocol OpenAI-compatible REST API

Cloud models require internet

The OpenRouter integration is a convenience feature for environments where internet is available. It is not part of Astra's core offline capability. All medical and general-purpose AI functionality works without any cloud connectivity using the local Ollama backends.

Environment Variable Reference

The following environment variables are set in the Helm deployment for each cluster:

Variable Value Purpose
DATABASE_URL postgresql://openwebui:...@<node1-ip>:5432/openwebui Per-cluster database connection string
WEBUI_SECRET_KEY (shared across all clusters) Session cookie signing key for failover compatibility
ENABLE_DB_MIGRATIONS False Workaround for peewee ORM migration bug in v0.8.12
HF_HUB_OFFLINE 1 Prevent HuggingFace library from attempting network downloads
TRANSFORMERS_OFFLINE 1 Prevent Transformers library from attempting network downloads
ENABLE_OLLAMA_API True Enable the local Ollama AI backend integration
ENABLE_WEBSOCKET_SUPPORT False Disabled (would require Redis, adding complexity)
VECTOR_DB pgvector Use PostgreSQL with pgvector extension for embedding storage

Standby Cluster Behavior

Expected behavior on standby clusters

Open WebUI instances running on standby clusters (C2 and C3 under normal conditions) will encounter read-only database errors when attempting write operations. This is expected behavior -- the PostgreSQL instances on standby nodes are streaming replicas that do not accept writes. The Open WebUI pod on a standby cluster may crash-loop or display error messages until that cluster's PostgreSQL is promoted to primary (either through a Keepalived failover or manual promotion). This is a known and accepted limitation of the architecture.