Ollama Inference Engine¶
What Is Ollama¶
Ollama is a lightweight AI inference server that manages the full lifecycle of large language models: downloading, quantization, memory management, and serving via a simple HTTP API. It abstracts the complexity of running LLMs on heterogeneous hardware, providing a consistent interface regardless of whether the underlying compute is a CPU-only ARM device or a CUDA-enabled GPU.
Why Ollama Was Chosen¶
The Astra system runs on a mix of hardware architectures:
- Raspberry Pi 5 (ARM64, CPU-only inference)
- NVIDIA Jetson Orin Nano (ARM64, CUDA GPU inference)
Ollama was selected because it satisfies all of the project's inference engine requirements:
| Requirement | Ollama Support |
|---|---|
| ARM64 CPU inference | Native support |
| NVIDIA CUDA GPU inference | Native support (including JetPack) |
| Simple HTTP REST API | Built-in on port 11434 |
| Automatic model memory management | Loads and unloads models automatically |
| Quantized model support (GGUF) | Full support for Q4, Q5, Q8 quantizations |
| Single binary deployment | One binary, no complex dependencies |
| Offline operation | All models served from local storage |
Alternative inference engines such as vLLM offer higher throughput for data center deployments but do not support ARM CPU inference and require more complex configuration. For a student project running on heterogeneous edge hardware, Ollama provides the best balance of simplicity and hardware compatibility.
Version and Deployment¶
All three AI servers run Ollama v0.20.7.
| Server | Hardware | Ollama User | Inference Type |
|---|---|---|---|
| ai-server1 | Raspberry Pi 5 (16 GB) | ollama |
CPU |
| ai-server2 | Raspberry Pi 5 (16 GB) | ollama |
CPU |
| ai-server3 | Jetson Orin Nano (8 GB) | gateway |
CUDA GPU |
Ollama runs as a systemd service on all three servers. On the Pi 5 units, it runs under the ollama system user. On the Jetson, it runs under the gateway user (configured via a systemd override) to ensure proper access to CUDA libraries and GPU devices.
Configuration via Systemd Override¶
Ollama's behavior is configured through environment variables set in a systemd override file at /etc/systemd/system/ollama.service.d/override.conf.
Common Settings (All Servers)¶
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_ORIGINS=*"
Environment="OLLAMA_MODELS=/home/gateway/.ollama/models"
Environment="OLLAMA_KEEP_ALIVE=30m"
Environment="OLLAMA_NUM_PARALLEL=1"
Environment="OLLAMA_MAX_LOADED_MODELS=1"
| Variable | Value | Purpose |
|---|---|---|
OLLAMA_HOST |
0.0.0.0:11434 |
Listen on all network interfaces so OpenWebUI pods can connect over the LAN |
OLLAMA_ORIGINS |
* |
Allow cross-origin requests from any source (required for OpenWebUI to connect) |
OLLAMA_MODELS |
/home/gateway/.ollama/models |
Directory where downloaded model files are stored |
OLLAMA_KEEP_ALIVE |
30m |
Keep a model loaded in memory for 30 minutes after the last request, then unload to free resources |
OLLAMA_NUM_PARALLEL |
1 |
Process only one inference request at a time to prevent memory exhaustion |
OLLAMA_MAX_LOADED_MODELS |
1 |
Allow only one model in memory at a time; switching models triggers automatic unload of the previous model |
Jetson-Specific Settings (ai-server3 Only)¶
The Jetson Orin Nano requires additional configuration for CUDA GPU inference:
Environment="OLLAMA_CONTEXT_LENGTH=3072"
Environment="OLLAMA_GPU_OVERHEAD=1073741824"
Environment="OLLAMA_LLM_LIBRARY=cuda_jetpack6"
Environment="JETSON_JETPACK=6"
Environment="LD_LIBRARY_PATH=/usr/local/lib/ollama:/usr/local/lib/ollama/cuda_jetpack6:/usr/local/cuda/targets/aarch64-linux/lib"
| Variable | Value | Purpose |
|---|---|---|
OLLAMA_CONTEXT_LENGTH |
3072 |
Reduced context window (default is 4096+) to conserve the Jetson's limited 8 GB unified memory |
OLLAMA_GPU_OVERHEAD |
1073741824 (1 GB) |
Reserve 1 GB of GPU memory for the operating system, CUDA runtime, and driver overhead |
OLLAMA_LLM_LIBRARY |
cuda_jetpack6 |
Use the JetPack 6 CUDA compute library for GPU inference |
JETSON_JETPACK |
6 |
Identifies the JetPack major version for Ollama's hardware detection |
LD_LIBRARY_PATH |
(see above) | Adds CUDA library paths so Ollama can find the GPU compute libraries |
Jetson memory constraints
The Jetson Orin Nano has 8 GB of unified RAM shared between the CPU and GPU. Unlike desktop GPUs with dedicated VRAM, the Jetson's GPU and CPU compete for the same physical memory. The OLLAMA_GPU_OVERHEAD and OLLAMA_CONTEXT_LENGTH settings are carefully tuned to prevent out-of-memory conditions that can freeze the device and require a physical power cycle.
API Endpoints¶
Ollama exposes a REST API on port 11434. The following endpoints are used by OpenWebUI and the status dashboard:
| Endpoint | Method | Purpose | Example Response |
|---|---|---|---|
/ |
GET | Health check | Ollama is running |
/api/tags |
GET | List all installed models | JSON array of model metadata |
/api/ps |
GET | Show currently loaded models (in RAM/VRAM) | JSON with model name, size, and status |
/api/generate |
POST | Generate a text completion | Streaming JSON response |
/api/chat |
POST | Multi-turn chat completion | Streaming JSON response |
Health Check¶
List Installed Models¶
Returns metadata for every model stored on the server, including name, size, quantization level, and modification date.
Show Loaded Models¶
Returns information about models currently loaded in RAM or VRAM. When no model is loaded, returns {"models":[]}. This endpoint is used by the status dashboard to show which model each AI server is currently serving.
Chat Completion¶
curl http://<server-ip>:11434/api/chat -d '{
"model": "qwen3:4b",
"messages": [{"role": "user", "content": "What is K3s?"}]
}'
Model Switching Behavior¶
Ollama handles model switching automatically. When a user selects a different model in OpenWebUI:
- Ollama receives a request for the new model
- If a different model is currently loaded, Ollama unloads it from memory
- The new model is loaded from disk into RAM (CPU) or VRAM (GPU)
- The inference request is processed
This process takes a few seconds depending on model size. The OLLAMA_KEEP_ALIVE=30m setting means a model stays loaded for 30 minutes after the last request before being automatically unloaded. If a new model request arrives while the current model is still loaded, the current model is unloaded immediately to make room.
Single-model constraint
With OLLAMA_MAX_LOADED_MODELS=1, only one model can be in memory at a time on each server. This is a deliberate constraint to prevent out-of-memory conditions on resource-constrained hardware. On the Pi 5 (16 GB), there is theoretically enough RAM for two small models, but keeping one loaded at a time provides a safety margin and consistent behavior.
Operational Commands¶
Check Ollama Service Status¶
Restart Ollama¶
View Ollama Logs¶
List Models on a Server¶
Pull a New Model¶
Pulling models requires internet
The ollama pull command downloads models from the Ollama model registry over the internet. This must be done before demo day while internet connectivity is available. During offline operation, only pre-pulled models are available.
Deployment Files¶
| Path | Purpose |
|---|---|
/etc/systemd/system/ollama.service.d/override.conf |
Ollama systemd environment overrides |
/home/gateway/.ollama/models/ |
Model file storage (Jetson) |
~/.ollama/models/ |
Model file storage (Pi 5 units) |