Skip to content

Keepalived & Virtual IP

Overview

Keepalived is the cornerstone of Astra's automatic failover system. It manages a single Virtual IP (VIP) address -- 10.0.0.50 -- that always points to whichever cluster node is currently serving as the active primary. Users and applications connect to this one stable address and are transparently routed to the correct node, even as the underlying primary changes due to hardware failure or maintenance.

Keepalived achieves this through the VRRP (Virtual Router Redundancy Protocol), an industry-standard protocol originally designed for router redundancy. In Astra's architecture, VRRP is repurposed to provide application-level high availability: the VIP floats between three physically separated Raspberry Pi nodes, each running on its own independent power domain.

VRRP Protocol Fundamentals

VRRP operates by having multiple nodes participate in a shared virtual router group. Each node is assigned a priority value, and the node with the highest priority among all currently reachable participants becomes the MASTER. The MASTER node:

  • Binds the VIP to its network interface, making it the target for all traffic destined for 10.0.0.50
  • Sends periodic VRRP advertisements (heartbeats) to all other participants at a configurable interval
  • Continues serving as MASTER until it either fails or is manually taken offline

All other nodes in the group operate in BACKUP state. They monitor the MASTER's advertisements and, if the MASTER stops sending them (indicating a failure), the highest-priority BACKUP transitions to MASTER and takes over the VIP.

Node Priorities and Election

All three of Astra's node1 servers participate in the VRRP group with the following priority assignments:

Node Priority LAN IP Role
k3s-c1-node1 100 (highest) 10.0.0.11 Preferred primary
k3s-c2-node1 90 10.0.0.21 First failover target
k3s-c3-node1 80 (lowest) 10.0.0.31 Last resort

All nodes start as BACKUP

Despite having different priorities, all three nodes are configured with state BACKUP in their Keepalived configuration. Keepalived handles the election process automatically based on priority values. This prevents split-brain scenarios during simultaneous boot, where two nodes might both believe they are the pre-configured MASTER.

When the system boots for the first time (or after a full power cycle), Keepalived runs an election. The node with the highest priority that is reachable and healthy becomes MASTER and receives the VIP. In normal operation, this is k3s-c1-node1 (priority 100).

Why Physical Ethernet (eth0)

VRRP uses IP protocol number 112 -- it is not a TCP or UDP application. It operates directly at the IP layer, similar to how ICMP (ping) uses protocol number 1. VPN tunnels typically only transport TCP, UDP, and ICMP traffic, so VRRP packets sent over VPN interfaces are silently dropped. Keepalived must run on a physical Ethernet interface with direct Layer 2 or Layer 3 connectivity between peers.

Astra runs Keepalived on eth0, the physical Ethernet interface connected to the LAN (10.0.0.0/24). On this network, VRRP packets are transmitted as standard Ethernet frames and are delivered without issue.

Unicast Communication

By default, VRRP uses multicast to communicate between nodes. However, Astra uses unicast instead. Each node is explicitly configured with the IP addresses of its peers.

The reason for this choice is practical: compact network routers may not reliably forward multicast traffic between their ports. Consumer-grade routers often have incomplete or buggy multicast forwarding implementations. Unicast eliminates this dependency entirely by sending VRRP advertisements directly to each peer's known IP address.

Configuration Reference

The following is a representative Keepalived configuration for k3s-c1-node1 (priority 100). The other two nodes use identical structure with adjusted priority, unicast_src_ip, and unicast_peer values.

vrrp_instance VI_1 {
    state BACKUP
    interface eth0
    virtual_router_id 51
    priority 100
    advert_int 1
    unicast_src_ip 10.0.0.11
    unicast_peer {
        10.0.0.21
        10.0.0.31
    }
    virtual_ipaddress {
        10.0.0.50/24
    }
    notify /etc/keepalived/notify.sh
}

Key configuration parameters:

Parameter Value Purpose
state BACKUP All nodes start as BACKUP; election determines MASTER
interface eth0 Physical Ethernet on the LAN
virtual_router_id 51 Shared group identifier (must match on all nodes)
priority 100 / 90 / 80 Determines election winner; highest active priority wins
advert_int 1 VRRP advertisement interval in seconds
unicast_src_ip Node's own LAN IP Source address for unicast VRRP packets
unicast_peer Other two nodes' LAN IPs Destinations for unicast VRRP packets
virtual_ipaddress 10.0.0.50/24 The VIP assigned to the MASTER
notify /etc/keepalived/notify.sh Script executed on state transitions

Notify Script: MASTER Promotion Sequence

When a node transitions to MASTER state, Keepalived executes the notify script (/etc/keepalived/notify.sh). This script orchestrates the critical promotion sequence that transforms a standby node into the active primary:

Step 1: VIP Assignment

Keepalived itself assigns 10.0.0.50/24 to the local eth0 interface. This happens automatically as part of the VRRP MASTER transition, before the notify script runs. From this point forward, all traffic destined for 10.0.0.50 is received by this node.

Step 2: PostgreSQL Promotion

The notify script promotes the local PostgreSQL instance from a read-only standby to a writable primary. This involves:

  • Removing the standby.signal file from the PostgreSQL data directory
  • Executing pg_ctl promote to end recovery mode and enable write operations
  • The promoted PostgreSQL instance begins accepting both read and write queries

Step 3: Open WebUI Restart

The script triggers a rolling restart of the Open WebUI Kubernetes deployment. This causes the application pod to restart and establish a new database connection to the now-writable local PostgreSQL instance. The restart ensures that the application does not continue attempting writes against a connection that was previously read-only.

BACKUP transition

When a node transitions to BACKUP state (for example, when a higher-priority node comes back online), no special action is taken. The node remains a standby and continues replicating from whichever node holds the VIP.

Systemd Integration

Keepalived runs as a systemd service with dependency overrides that ensure it starts after all required network services are initialized. This guarantees that the physical network interface is fully configured before Keepalived begins sending VRRP advertisements.

Failover Timeline

The following table describes the sequence of events from the moment a primary node loses power to the moment users can resume using the system:

Time Event
0 s Primary node loses power (or is physically unplugged)
~1--2 s Surviving nodes detect missing VRRP advertisements from the former MASTER
~2 s Highest-priority surviving node transitions to MASTER and receives the VIP
~2--3 s Notify script promotes local PostgreSQL from standby to primary
~3--5 s Open WebUI pod restarts and connects to the promoted database
~5 s System is fully operational on the new primary

In practice, a user who is mid-conversation at the moment of failure experiences only a brief page reload. Their conversation history is preserved because all data was replicated to the standby nodes via PostgreSQL streaming replication before the failure occurred.

Interaction with PostgreSQL Replication

The VIP managed by Keepalived is central to Astra's self-healing PostgreSQL replication architecture. Each standby node's primary_conninfo points to 10.0.0.50 (the VIP), not to any specific node's fixed IP address. This means that when a new node is promoted to MASTER and receives the VIP, all remaining standbys automatically reconnect to the new primary without any configuration change.

This design enables cascade failover -- the ability to survive multiple sequential node failures. For a detailed description of the replication architecture and cascade failover behavior, see the Database & Replication documentation.