AI Servers¶
Astra uses three dedicated AI inference servers running Ollama to serve language models. These servers are physically and logically separate from the K3s application clusters, providing independent failure modes and preventing AI workloads from competing with the application and database layers for resources.
Server Inventory¶
| Server | Hardware | RAM | LAN IP | Storage | Inference Type | Approximate Speed |
|---|---|---|---|---|---|---|
| ai-server1 | Raspberry Pi 5 | 16 GB | 10.0.0.41 | 128 GB microSD | CPU-only | ~3--8 tokens/second |
| ai-server2 | Raspberry Pi 5 | 16 GB | 10.0.0.42 | 128 GB microSD | CPU-only | ~3--8 tokens/second |
| ai-server3 | NVIDIA Jetson Orin Nano | 8 GB (unified) | 10.0.0.43 | 128 GB eMMC | CUDA GPU | ~25--40 tokens/second |
All three servers run Ollama v0.20.7 and are configured as named Ollama endpoints in every OpenWebUI cluster. Each server listens on port 11434 for inference requests from the application layer.
Power delivery
Unlike the K3s compute nodes, the AI servers do not use PoE for power. They are powered via their own standard power supplies (USB-C for the Pi 5 units, DC barrel jack for the Jetson) and are connected to the PoE switches for networking only.
Raspberry Pi 5 Servers (ai-server1, ai-server2)¶
The two Raspberry Pi 5 units provide CPU-based AI inference with 16 GB of RAM each. Key characteristics:
| Specification | Value |
|---|---|
| Processor | Broadcom BCM2712, quad-core Cortex-A76, 2.4 GHz |
| RAM | 16 GB LPDDR4X |
| Storage | 128 GB microSD (solid-state flash) |
| Inference | CPU-only (no GPU acceleration) |
| OS | Debian GNU/Linux 13 (trixie) / Linux ARM64 |
| Power | USB-C 5V |
With 16 GB of RAM, the Pi 5 units can load and run models up to approximately 8 GB in size, including larger models like qwen3:8b (5.2 GB) and gemma4:e2b (7.2 GB) that cannot fit in the Jetson's 8 GB memory. The trade-off is speed: CPU-only inference on the Pi 5 produces approximately 3--8 tokens per second depending on model size, which is sufficient for demonstration but noticeably slower than the Jetson's GPU-accelerated inference.
Like the Jetson, the Pi 5 servers are configured with OLLAMA_MAX_LOADED_MODELS=1 to ensure only one model is loaded at a time, preventing memory exhaustion. OLLAMA_KEEP_ALIVE=30m keeps the loaded model in memory for 30 minutes of idle time before unloading.
NVIDIA Jetson Orin Nano (ai-server3)¶
The Jetson Orin Nano is the fastest inference device in the Astra system, providing approximately 5--10 times the throughput of the Pi 5 units for the same model.
| Specification | Value |
|---|---|
| Processor | 6-core Arm Cortex-A78AE v8.2 |
| GPU | 1024-core NVIDIA Ampere architecture with 32 Tensor Cores |
| RAM | 8 GB LPDDR5 (unified -- shared between CPU and GPU) |
| Storage | 128 GB eMMC (embedded solid-state, soldered to module) |
| OS | Ubuntu 22.04.5 LTS (Jetson Linux R36.4.7, JetPack 6.2.1) |
| CUDA | JetPack 6.2.1 CUDA toolkit |
| Power | DC barrel jack |
Unified Memory Architecture¶
The Jetson Orin Nano uses a unified memory architecture where the 8 GB of LPDDR5 RAM is shared between the CPU and GPU. Unlike desktop GPUs that have dedicated VRAM separate from system RAM, every byte used by the CUDA GPU reduces the memory available to the operating system and applications.
This has direct implications for model loading:
- The operating system, CUDA runtime, and system services consume approximately 1--1.5 GB at idle.
- Ollama itself requires additional memory for its process.
- The remaining memory is available for loading AI models.
- In practice, models up to approximately 5 GB can be loaded safely. Models larger than this risk exhausting all available memory and freezing the device.
Memory constraints on the Jetson
Loading a model larger than approximately 5 GB on the Jetson can consume all available unified memory. The Linux OOM killer may not recover cleanly because the CUDA driver holds memory in a way that prevents normal reclamation. If this occurs, the device requires a physical power cycle. The Ollama configuration on ai-server3 includes protective settings (OLLAMA_GPU_OVERHEAD=1073741824, OLLAMA_CONTEXT_LENGTH=3072, OLLAMA_MAX_LOADED_MODELS=1) to mitigate this risk.
Jetson-Specific Ollama Configuration¶
The Jetson requires additional environment variables beyond the standard Ollama configuration:
[Service]
Environment="OLLAMA_MAX_LOADED_MODELS=1"
Environment="OLLAMA_CONTEXT_LENGTH=3072"
Environment="OLLAMA_GPU_OVERHEAD=1073741824"
Environment="OLLAMA_LLM_LIBRARY=cuda_jetpack6"
Environment="JETSON_JETPACK=6"
Environment="LD_LIBRARY_PATH=/usr/local/lib/ollama:/usr/local/lib/ollama/cuda_jetpack6:/usr/local/cuda/targets/aarch64-linux/lib"
| Setting | Purpose |
|---|---|
OLLAMA_MAX_LOADED_MODELS=1 |
Ensures only one model is loaded at a time, preventing memory exhaustion |
OLLAMA_CONTEXT_LENGTH=3072 |
Reduces the context window size to save memory (default is 4096) |
OLLAMA_GPU_OVERHEAD=1073741824 |
Reserves 1 GB of memory for the OS and CUDA runtime |
OLLAMA_LLM_LIBRARY=cuda_jetpack6 |
Selects the CUDA-accelerated inference library compiled for JetPack 6 |
Performance Comparison¶
The performance difference between CPU-only inference on the Pi 5 and GPU-accelerated inference on the Jetson is significant:
| Metric | Pi 5 (CPU) | Jetson Orin Nano (GPU) |
|---|---|---|
| Inference speed (4B model) | ~3--8 tokens/second | ~25--40 tokens/second |
| Time to first token | 2--5 seconds | Under 1 second |
| RAM available for models | ~14 GB (of 16 GB) | ~6 GB (of 8 GB shared) |
| Largest model supported | ~8 GB | ~5 GB |
| Power consumption (inference) | ~10W | ~15W |
The Jetson is the preferred server for real-time demonstrations and medical AI queries where response latency matters. The Pi 5 units serve as backup inference capacity and can run larger models that do not fit in the Jetson's constrained memory.
Installed Models¶
All three AI servers have the following models installed and available:
| Model | Size | Purpose | Recommended Server |
|---|---|---|---|
qwen3:4b |
2.5 GB | Fastest general-purpose reasoning model with hybrid thinking mode | Jetson (best quality/speed) |
qwen3:8b |
5.2 GB | Higher-quality general-purpose model | Pi 5 only (exceeds Jetson safe limit) |
phi4-mini |
2.5 GB | Math, reasoning, and coding tasks | All servers |
gemma3:1b |
815 MB | Ultra-fast lightweight model for speed demonstrations | All servers |
gemma4:e2b |
7.2 GB | Multimodal reasoning (vision and audio input) | Pi 5 only (too large for Jetson) |
meditron |
3.8 GB | Medical text Q&A (developed by EPFL and Yale) | All servers |
medgemma-1.5-4b-it |
3.3 GB | Medical imaging and text (Google, quantized to Q4_K_M) | Jetson (fast GPU inference) |
Model selection for demonstrations
For the absolute fastest response times, use gemma3:1b (815 MB) -- it loads almost instantly and produces near-instant responses on any server. For the fastest general-purpose reasoning model with good quality, use qwen3:4b on ai-server3 (the Jetson). For medical demonstrations, use meditron or medgemma-1.5-4b-it.
Why Separate AI Servers¶
Separating AI inference from the K3s application clusters is a deliberate architectural decision based on two principles:
Resource isolation. AI inference is computationally intensive and consumes significant RAM. Running inference on the same nodes that handle the web application, database, and replication would create resource contention. A large model loading into memory could cause PostgreSQL replication lag, web interface timeouts, or K3s scheduling failures.
Independent failure modes. With separate servers, losing a K3s node does not reduce AI inference capacity, and losing an AI server does not affect the web application, database, or failover system. Each layer can fail and recover independently.
This mirrors the design philosophy used in production data centers, where compute, storage, and inference workloads run on dedicated hardware to prevent cross-layer interference.