Skip to content

Model Selection

Installed Models

All three AI servers have the following models installed and available for inference:

Model Size on Disk Parameter Count Purpose Best Run On
qwen3:4b 2.5 GB 4 billion General-purpose chat Jetson (25-40 tok/s), Pis (3-8 tok/s)
qwen3:8b 5.2 GB 8 billion Higher-quality general-purpose chat Pi 5 only (16 GB RAM required)
phi4-mini 2.5 GB ~3.8 billion Math, reasoning, and coding tasks All servers
gemma3:1b 815 MB 1 billion Ultra-fast lightweight responses Pi 5 (near-instant responses)
gemma4:e2b 7.2 GB ~2 billion (multimodal) Vision and audio input analysis Pi 5 only (too large for 8 GB Jetson)
meditron 3.8 GB 7 billion Medical text Q&A (EPFL/Yale) All servers
medgemma-1.5-4b-it 3.3 GB (Q4_K_M) 4 billion Medical imaging and text (Google) Jetson (fast GPU inference)

Why These Models Were Chosen

Model selection was driven by three constraints unique to the Astra project:

  1. Hardware memory limits. The Jetson Orin Nano has 8 GB of unified RAM (shared CPU/GPU), and the Pi 5 units have 16 GB. Models must fit in memory with room for the operating system and CUDA runtime overhead.

  2. Inference speed on edge hardware. Models must produce responses fast enough for a conversational experience. On CPU-only Pi 5 hardware, this means models under approximately 5 GB. On the Jetson with CUDA acceleration, models up to 5 GB run at 25-40 tokens per second.

  3. Mission-relevant capabilities. The medical AI use case requires models specifically trained on biomedical literature, not just general-purpose models adapted for medical questions.

General-Purpose Models

Qwen3 (4B and 8B) -- The Qwen3 family provides strong general-purpose performance at sizes that fit on edge hardware. The 4B variant runs efficiently on the Jetson GPU at 25-40 tokens per second, making it the fastest model that provides strong general-purpose reasoning quality. The 8B variant offers higher quality output but is limited to the Pi 5 units due to its 5.2 GB memory footprint.

Phi4-mini -- A model from Microsoft optimized for mathematical reasoning, logical deduction, and code generation. Its compact 2.5 GB size means it runs on all servers, and its specialized training makes it complementary to the general-purpose Qwen3 models.

Gemma3 (1B) -- The absolute fastest model in the system. At 815 MB and 1 billion parameters, it loads almost instantly and generates responses with minimal latency even on CPU-only hardware. It produces the highest token throughput of any model on any hardware in the system, though with lower reasoning quality than larger models. Its primary purpose is demonstrating the system's responsiveness during live presentations.

Multimodal Model

Gemma4 (E2B, 7.2 GB) -- A multimodal model capable of processing vision and audio inputs alongside text. This model can analyze uploaded images, making it relevant for medical imaging scenarios. However, its 7.2 GB size restricts it to the Pi 5 units.

Medical Models

Meditron (3.8 GB) -- Developed by EPFL and Yale University, Meditron was trained on a curated corpus of medical literature including clinical guidelines, medical textbooks, and peer-reviewed research. It provides medical Q&A capabilities tailored to the needs of non-specialist users (such as astronauts with basic medical training).

MedGemma 1.5 (3.3 GB, quantized) -- A medical AI model from Google designed for both medical imaging analysis and text-based medical Q&A. The model is quantized to Q4_K_M format for efficient inference on edge hardware. It runs best on the Jetson GPU, where it benefits from CUDA acceleration for image processing tasks.

See the Medical AI page for detailed information on the medical use case and these models.

Model Selection Guidance for Demos

Fastest Responses

For the fastest raw token generation, use gemma3:1b on any server -- at 815 MB, it is the absolute fastest model on any hardware. For the fastest responses that also provide strong general-purpose reasoning quality, use ai3.qwen3:4b, which runs on the Jetson Orin Nano with CUDA GPU acceleration at approximately 25-40 tokens per second. The ai3 prefix routes the request to the Jetson specifically.

Medical Demonstrations

Use meditron for text-based medical questions (symptoms, treatment protocols, drug interactions). Use medgemma-1.5-4b-it on the Jetson when medical imaging analysis is needed.

Speed Demonstration

Use gemma3:1b on any server to demonstrate near-instant response times. As the smallest and fastest model in the system, it loads in seconds and generates tokens faster than a user can read them, even on CPU-only hardware.

General Conversation

Use qwen3:4b on the Jetson for fast general chat, or qwen3:8b on a Pi 5 for higher-quality responses at slower speed.

Math and Coding

Use phi4-mini for questions involving mathematical reasoning, logical problems, or code generation.

Performance Characteristics

Inference Speed by Hardware

Hardware Approximate Speed (4B model) Approximate Speed (8B model)
Jetson Orin Nano (CUDA GPU) 25-40 tokens/second Not recommended (memory)
Raspberry Pi 5 (CPU only) 3-8 tokens/second 2-5 tokens/second

What tokens per second means for the user

A typical English word is 1-2 tokens. At 30 tokens/second, the AI produces roughly 15-20 words per second -- faster than a user can comfortably read. At 5 tokens/second, the response appears at roughly 3-4 words per second, which is noticeable but still conversational.

Model Loading Times

Model loading time depends on the model size and storage speed:

Model Size Pi 5 (microSD) Jetson (eMMC)
< 1 GB 2-5 seconds 1-3 seconds
2-4 GB 5-15 seconds 3-8 seconds
5-8 GB 15-30 seconds 10-20 seconds

The first request to a model after it has been unloaded incurs a loading delay. Subsequent requests are served immediately from memory until the model is unloaded (after 30 minutes of inactivity, per the OLLAMA_KEEP_ALIVE setting).

Memory Constraints and Warnings

Do not load gemma4:e2b on the Jetson

The gemma4:e2b model is 7.2 GB. The Jetson Orin Nano has 8 GB of unified RAM shared between CPU and GPU. Loading this model consumes nearly all available memory, causing the Linux OOM killer to activate. Because the CUDA driver holds GPU memory in a way that prevents clean reclamation, the device freezes completely and requires a physical power cycle to recover.

Only use models 5 GB or smaller on the Jetson. Use the Pi 5 units (16 GB RAM) for larger models.

Memory Budget by Server

Server Total RAM OS + Runtime Overhead Available for Models
Pi 5 (ai-server1, ai-server2) 16 GB ~2 GB ~14 GB
Jetson (ai-server3) 8 GB (unified) ~2-3 GB (OS + CUDA) ~5-6 GB

The Jetson's 1 GB OLLAMA_GPU_OVERHEAD reservation and reduced OLLAMA_CONTEXT_LENGTH of 3072 tokens help prevent out-of-memory conditions, but the fundamental constraint is the 8 GB unified memory pool.

Planned Models

The following model is planned for future installation but is not yet available on the servers:

Model Source Size Status
BioMedLM (2.7B) Stanford CRFM TBD Requires HuggingFace-to-GGUF conversion and Ollama import

BioMedLM is a biomedical language model from the Stanford Center for Research on Foundation Models. It was trained exclusively on biomedical text from PubMed and would complement the existing Meditron and MedGemma models. Converting it from HuggingFace format to GGUF (the format Ollama uses) requires additional tooling that has not yet been set up.