Skip to content

Ollama Inference Engine

What Is Ollama

Ollama is a lightweight AI inference server that manages the full lifecycle of large language models: downloading, quantization, memory management, and serving via a simple HTTP API. It abstracts the complexity of running LLMs on heterogeneous hardware, providing a consistent interface regardless of whether the underlying compute is a CPU-only ARM device or a CUDA-enabled GPU.

Why Ollama Was Chosen

The Astra system runs on a mix of hardware architectures:

  • Raspberry Pi 5 (ARM64, CPU-only inference)
  • NVIDIA Jetson Orin Nano (ARM64, CUDA GPU inference)

Ollama was selected because it satisfies all of the project's inference engine requirements:

Requirement Ollama Support
ARM64 CPU inference Native support
NVIDIA CUDA GPU inference Native support (including JetPack)
Simple HTTP REST API Built-in on port 11434
Automatic model memory management Loads and unloads models automatically
Quantized model support (GGUF) Full support for Q4, Q5, Q8 quantizations
Single binary deployment One binary, no complex dependencies
Offline operation All models served from local storage

Alternative inference engines such as vLLM offer higher throughput for data center deployments but do not support ARM CPU inference and require more complex configuration. For a student project running on heterogeneous edge hardware, Ollama provides the best balance of simplicity and hardware compatibility.

Version and Deployment

All three AI servers run Ollama v0.20.7.

Server Hardware Ollama User Inference Type
ai-server1 Raspberry Pi 5 (16 GB) ollama CPU
ai-server2 Raspberry Pi 5 (16 GB) ollama CPU
ai-server3 Jetson Orin Nano (8 GB) gateway CUDA GPU

Ollama runs as a systemd service on all three servers. On the Pi 5 units, it runs under the ollama system user. On the Jetson, it runs under the gateway user (configured via a systemd override) to ensure proper access to CUDA libraries and GPU devices.

Configuration via Systemd Override

Ollama's behavior is configured through environment variables set in a systemd override file at /etc/systemd/system/ollama.service.d/override.conf.

Common Settings (All Servers)

[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_ORIGINS=*"
Environment="OLLAMA_MODELS=/home/gateway/.ollama/models"
Environment="OLLAMA_KEEP_ALIVE=30m"
Environment="OLLAMA_NUM_PARALLEL=1"
Environment="OLLAMA_MAX_LOADED_MODELS=1"
Variable Value Purpose
OLLAMA_HOST 0.0.0.0:11434 Listen on all network interfaces so OpenWebUI pods can connect over the LAN
OLLAMA_ORIGINS * Allow cross-origin requests from any source (required for OpenWebUI to connect)
OLLAMA_MODELS /home/gateway/.ollama/models Directory where downloaded model files are stored
OLLAMA_KEEP_ALIVE 30m Keep a model loaded in memory for 30 minutes after the last request, then unload to free resources
OLLAMA_NUM_PARALLEL 1 Process only one inference request at a time to prevent memory exhaustion
OLLAMA_MAX_LOADED_MODELS 1 Allow only one model in memory at a time; switching models triggers automatic unload of the previous model

Jetson-Specific Settings (ai-server3 Only)

The Jetson Orin Nano requires additional configuration for CUDA GPU inference:

Environment="OLLAMA_CONTEXT_LENGTH=3072"
Environment="OLLAMA_GPU_OVERHEAD=1073741824"
Environment="OLLAMA_LLM_LIBRARY=cuda_jetpack6"
Environment="JETSON_JETPACK=6"
Environment="LD_LIBRARY_PATH=/usr/local/lib/ollama:/usr/local/lib/ollama/cuda_jetpack6:/usr/local/cuda/targets/aarch64-linux/lib"
Variable Value Purpose
OLLAMA_CONTEXT_LENGTH 3072 Reduced context window (default is 4096+) to conserve the Jetson's limited 8 GB unified memory
OLLAMA_GPU_OVERHEAD 1073741824 (1 GB) Reserve 1 GB of GPU memory for the operating system, CUDA runtime, and driver overhead
OLLAMA_LLM_LIBRARY cuda_jetpack6 Use the JetPack 6 CUDA compute library for GPU inference
JETSON_JETPACK 6 Identifies the JetPack major version for Ollama's hardware detection
LD_LIBRARY_PATH (see above) Adds CUDA library paths so Ollama can find the GPU compute libraries

Jetson memory constraints

The Jetson Orin Nano has 8 GB of unified RAM shared between the CPU and GPU. Unlike desktop GPUs with dedicated VRAM, the Jetson's GPU and CPU compete for the same physical memory. The OLLAMA_GPU_OVERHEAD and OLLAMA_CONTEXT_LENGTH settings are carefully tuned to prevent out-of-memory conditions that can freeze the device and require a physical power cycle.

API Endpoints

Ollama exposes a REST API on port 11434. The following endpoints are used by OpenWebUI and the status dashboard:

Endpoint Method Purpose Example Response
/ GET Health check Ollama is running
/api/tags GET List all installed models JSON array of model metadata
/api/ps GET Show currently loaded models (in RAM/VRAM) JSON with model name, size, and status
/api/generate POST Generate a text completion Streaming JSON response
/api/chat POST Multi-turn chat completion Streaming JSON response

Health Check

curl http://<server-ip>:11434/
# Response: Ollama is running

List Installed Models

curl http://<server-ip>:11434/api/tags

Returns metadata for every model stored on the server, including name, size, quantization level, and modification date.

Show Loaded Models

curl http://<server-ip>:11434/api/ps

Returns information about models currently loaded in RAM or VRAM. When no model is loaded, returns {"models":[]}. This endpoint is used by the status dashboard to show which model each AI server is currently serving.

Chat Completion

curl http://<server-ip>:11434/api/chat -d '{
  "model": "qwen3:4b",
  "messages": [{"role": "user", "content": "What is K3s?"}]
}'

Model Switching Behavior

Ollama handles model switching automatically. When a user selects a different model in OpenWebUI:

  1. Ollama receives a request for the new model
  2. If a different model is currently loaded, Ollama unloads it from memory
  3. The new model is loaded from disk into RAM (CPU) or VRAM (GPU)
  4. The inference request is processed

This process takes a few seconds depending on model size. The OLLAMA_KEEP_ALIVE=30m setting means a model stays loaded for 30 minutes after the last request before being automatically unloaded. If a new model request arrives while the current model is still loaded, the current model is unloaded immediately to make room.

Single-model constraint

With OLLAMA_MAX_LOADED_MODELS=1, only one model can be in memory at a time on each server. This is a deliberate constraint to prevent out-of-memory conditions on resource-constrained hardware. On the Pi 5 (16 GB), there is theoretically enough RAM for two small models, but keeping one loaded at a time provides a safety margin and consistent behavior.

Operational Commands

Check Ollama Service Status

sudo systemctl status ollama

Restart Ollama

sudo systemctl restart ollama

View Ollama Logs

journalctl -u ollama -f

List Models on a Server

ollama list

Pull a New Model

ollama pull qwen3:4b

Pulling models requires internet

The ollama pull command downloads models from the Ollama model registry over the internet. This must be done before demo day while internet connectivity is available. During offline operation, only pre-pulled models are available.

Deployment Files

Path Purpose
/etc/systemd/system/ollama.service.d/override.conf Ollama systemd environment overrides
/home/gateway/.ollama/models/ Model file storage (Jetson)
~/.ollama/models/ Model file storage (Pi 5 units)