Skip to content

AI Servers

Astra uses three dedicated AI inference servers running Ollama to serve language models. These servers are physically and logically separate from the K3s application clusters, providing independent failure modes and preventing AI workloads from competing with the application and database layers for resources.


Server Inventory

Server Hardware RAM LAN IP Storage Inference Type Approximate Speed
ai-server1 Raspberry Pi 5 16 GB 10.0.0.41 128 GB microSD CPU-only ~3--8 tokens/second
ai-server2 Raspberry Pi 5 16 GB 10.0.0.42 128 GB microSD CPU-only ~3--8 tokens/second
ai-server3 NVIDIA Jetson Orin Nano 8 GB (unified) 10.0.0.43 128 GB eMMC CUDA GPU ~25--40 tokens/second

All three servers run Ollama v0.20.7 and are configured as named Ollama endpoints in every OpenWebUI cluster. Each server listens on port 11434 for inference requests from the application layer.

Power delivery

Unlike the K3s compute nodes, the AI servers do not use PoE for power. They are powered via their own standard power supplies (USB-C for the Pi 5 units, DC barrel jack for the Jetson) and are connected to the PoE switches for networking only.


Raspberry Pi 5 Servers (ai-server1, ai-server2)

The two Raspberry Pi 5 units provide CPU-based AI inference with 16 GB of RAM each. Key characteristics:

Specification Value
Processor Broadcom BCM2712, quad-core Cortex-A76, 2.4 GHz
RAM 16 GB LPDDR4X
Storage 128 GB microSD (solid-state flash)
Inference CPU-only (no GPU acceleration)
OS Debian GNU/Linux 13 (trixie) / Linux ARM64
Power USB-C 5V

With 16 GB of RAM, the Pi 5 units can load and run models up to approximately 8 GB in size, including larger models like qwen3:8b (5.2 GB) and gemma4:e2b (7.2 GB) that cannot fit in the Jetson's 8 GB memory. The trade-off is speed: CPU-only inference on the Pi 5 produces approximately 3--8 tokens per second depending on model size, which is sufficient for demonstration but noticeably slower than the Jetson's GPU-accelerated inference.

Like the Jetson, the Pi 5 servers are configured with OLLAMA_MAX_LOADED_MODELS=1 to ensure only one model is loaded at a time, preventing memory exhaustion. OLLAMA_KEEP_ALIVE=30m keeps the loaded model in memory for 30 minutes of idle time before unloading.


NVIDIA Jetson Orin Nano (ai-server3)

The Jetson Orin Nano is the fastest inference device in the Astra system, providing approximately 5--10 times the throughput of the Pi 5 units for the same model.

Specification Value
Processor 6-core Arm Cortex-A78AE v8.2
GPU 1024-core NVIDIA Ampere architecture with 32 Tensor Cores
RAM 8 GB LPDDR5 (unified -- shared between CPU and GPU)
Storage 128 GB eMMC (embedded solid-state, soldered to module)
OS Ubuntu 22.04.5 LTS (Jetson Linux R36.4.7, JetPack 6.2.1)
CUDA JetPack 6.2.1 CUDA toolkit
Power DC barrel jack

Unified Memory Architecture

The Jetson Orin Nano uses a unified memory architecture where the 8 GB of LPDDR5 RAM is shared between the CPU and GPU. Unlike desktop GPUs that have dedicated VRAM separate from system RAM, every byte used by the CUDA GPU reduces the memory available to the operating system and applications.

This has direct implications for model loading:

  • The operating system, CUDA runtime, and system services consume approximately 1--1.5 GB at idle.
  • Ollama itself requires additional memory for its process.
  • The remaining memory is available for loading AI models.
  • In practice, models up to approximately 5 GB can be loaded safely. Models larger than this risk exhausting all available memory and freezing the device.

Memory constraints on the Jetson

Loading a model larger than approximately 5 GB on the Jetson can consume all available unified memory. The Linux OOM killer may not recover cleanly because the CUDA driver holds memory in a way that prevents normal reclamation. If this occurs, the device requires a physical power cycle. The Ollama configuration on ai-server3 includes protective settings (OLLAMA_GPU_OVERHEAD=1073741824, OLLAMA_CONTEXT_LENGTH=3072, OLLAMA_MAX_LOADED_MODELS=1) to mitigate this risk.

Jetson-Specific Ollama Configuration

The Jetson requires additional environment variables beyond the standard Ollama configuration:

[Service]
Environment="OLLAMA_MAX_LOADED_MODELS=1"
Environment="OLLAMA_CONTEXT_LENGTH=3072"
Environment="OLLAMA_GPU_OVERHEAD=1073741824"
Environment="OLLAMA_LLM_LIBRARY=cuda_jetpack6"
Environment="JETSON_JETPACK=6"
Environment="LD_LIBRARY_PATH=/usr/local/lib/ollama:/usr/local/lib/ollama/cuda_jetpack6:/usr/local/cuda/targets/aarch64-linux/lib"
Setting Purpose
OLLAMA_MAX_LOADED_MODELS=1 Ensures only one model is loaded at a time, preventing memory exhaustion
OLLAMA_CONTEXT_LENGTH=3072 Reduces the context window size to save memory (default is 4096)
OLLAMA_GPU_OVERHEAD=1073741824 Reserves 1 GB of memory for the OS and CUDA runtime
OLLAMA_LLM_LIBRARY=cuda_jetpack6 Selects the CUDA-accelerated inference library compiled for JetPack 6

Performance Comparison

The performance difference between CPU-only inference on the Pi 5 and GPU-accelerated inference on the Jetson is significant:

Metric Pi 5 (CPU) Jetson Orin Nano (GPU)
Inference speed (4B model) ~3--8 tokens/second ~25--40 tokens/second
Time to first token 2--5 seconds Under 1 second
RAM available for models ~14 GB (of 16 GB) ~6 GB (of 8 GB shared)
Largest model supported ~8 GB ~5 GB
Power consumption (inference) ~10W ~15W

The Jetson is the preferred server for real-time demonstrations and medical AI queries where response latency matters. The Pi 5 units serve as backup inference capacity and can run larger models that do not fit in the Jetson's constrained memory.


Installed Models

All three AI servers have the following models installed and available:

Model Size Purpose Recommended Server
qwen3:4b 2.5 GB Fastest general-purpose reasoning model with hybrid thinking mode Jetson (best quality/speed)
qwen3:8b 5.2 GB Higher-quality general-purpose model Pi 5 only (exceeds Jetson safe limit)
phi4-mini 2.5 GB Math, reasoning, and coding tasks All servers
gemma3:1b 815 MB Ultra-fast lightweight model for speed demonstrations All servers
gemma4:e2b 7.2 GB Multimodal reasoning (vision and audio input) Pi 5 only (too large for Jetson)
meditron 3.8 GB Medical text Q&A (developed by EPFL and Yale) All servers
medgemma-1.5-4b-it 3.3 GB Medical imaging and text (Google, quantized to Q4_K_M) Jetson (fast GPU inference)

Model selection for demonstrations

For the absolute fastest response times, use gemma3:1b (815 MB) -- it loads almost instantly and produces near-instant responses on any server. For the fastest general-purpose reasoning model with good quality, use qwen3:4b on ai-server3 (the Jetson). For medical demonstrations, use meditron or medgemma-1.5-4b-it.


Why Separate AI Servers

Separating AI inference from the K3s application clusters is a deliberate architectural decision based on two principles:

Resource isolation. AI inference is computationally intensive and consumes significant RAM. Running inference on the same nodes that handle the web application, database, and replication would create resource contention. A large model loading into memory could cause PostgreSQL replication lag, web interface timeouts, or K3s scheduling failures.

Independent failure modes. With separate servers, losing a K3s node does not reduce AI inference capacity, and losing an AI server does not affect the web application, database, or failover system. Each layer can fail and recover independently.

This mirrors the design philosophy used in production data centers, where compute, storage, and inference workloads run on dedicated hardware to prevent cross-layer interference.