# llama.cpp

## Overview

[llama.cpp](https://github.com/ggml-org/llama.cpp) provides an inference path for large language models (LLMs) and vision-language
models (VLMs) in [RVC4](https://docs.luxonis.com/hardware/platform/rvc/rvc4.md) standalone apps. The official
[`oakapp-base`](https://github.com/luxonis/oakapp-base#llamacpp-image) image includes CPU and Hexagon backends, allowing an app to
load GGUF models and offload supported operations to the device's Hexagon Tensor Processor (HTP).

Use this path for tasks such as answering questions about a camera frame, describing a selected region, or generating text locally
on the device. The deployment described here runs entirely inside an RVC4 standalone OAK App and is not supported on RVC2 devices.

Unlike a model deployed through `dai.node.NeuralNetwork`, a GGUF model is loaded by llama.cpp outside the DepthAI graph. DepthAI
handles camera capture and streaming; the application prepares inputs and sends them to llama.cpp. There is no SNPE or NN Archive
conversion step for this path.

## Inference Flow

A typical VLM app uses the following flow:

 1. Package the language model and its matching vision projector with the app.
 2. Start `llama-server` on the device and wait for its `/health` endpoint to report readiness.
 3. Capture a frame with DepthAI, optionally crop a region, and encode the prepared image as JPEG.
 4. Send the image and a text prompt to the local `/v1/chat/completions` endpoint.
 5. The vision projector encodes the image, and the language model processes the image embeddings and prompt to generate an
    answer.
 6. Consume the answer in the application or stream answer tokens to a frontend.

The server exposes an OpenAI-compatible HTTP API on the device; requests to this local endpoint do not use a cloud inference
service. For text-only inference, omit the projector and image input.

| Path | Model artifacts | Execution target | Integration |
| --- | --- | --- | --- |
| llama.cpp CPU | GGUF model, plus a matching projector for vision | RVC4 ARM CPU | Local server API or llama.cpp libraries |
| llama.cpp Hexagon | The same GGUF artifacts, with backend-compatible operations and tensor types | HTP, with CPU execution for
unsupported operations | Local server API or llama.cpp libraries |
| Native DepthAI | Converted RVC4 model, usually supplied through an NN Archive | Native RVC4 inference runtime | `NeuralNetwork`
node in the DepthAI graph |

For applications that need the ONNX Runtime API, see [ONNX Runtime with
QNN](https://docs.luxonis.com/software-v3/ai-inference/inference/onnx-runtime-qnn.md). llama.cpp uses its own Hexagon backend
rather than the ONNX Runtime QNN Execution Provider.

## Requirements

 * An RVC4-based OAK device running Luxonis OS 1.40.0 or newer.
 * A [standalone OAK App](https://docs.luxonis.com/software-v3/oak-apps.md) using `luxonis/oakapp-base:1.5.0-llamacpp` or a
   compatible newer llama.cpp image.
 * A GGUF model supported by the llama.cpp version in that image, with a matching multimodal projector for vision input.
 * Enough device storage and memory for the weights, projector, inference buffers, and chosen context size.
 * For Hexagon acceleration, access to the FastRPC devices and the OS-provided NPU runtime inside the container.

### What the base image includes

The `1.5.0-llamacpp` image contains llama.cpp v0.5.0 for `linux/arm64`, with CPU and Hexagon v73 support. It provides
`llama-server`, `llama-cli`, and `llama-bench` on `PATH`, as well as shared libraries, multimodal support, and C/C++ headers under
`/opt/llama.cpp`. The pinned source revision and build features are recorded in `/opt/llama.cpp/build-info.json`.

The image configures the library search paths for llama.cpp and the device NPU runtime. FastRPC libraries come from Luxonis OS at
`/opt/luxonis/npu-runtime/lib` and must be mounted into an accelerated app. CPU-only apps do not require the NPU mount or device
passthrough.

The base image does not include model weights or start `llama-server` automatically. Built-in HTTPS model downloads are disabled
in this build, so supply local model files or download them using application tooling during the container build.

### OAK App configuration

The following `oakapp.toml` configures a Python app that starts the included server. Keep the standard `/entrypoint.sh` entrypoint
and install your app's Python dependencies from `requirements.txt`.

```toml
identifier = "com.example.llamacpp"
app_version = "1.0.0"
base_image = "luxonis/oakapp-base:1.5.0-llamacpp"

entrypoint = ["/entrypoint.sh", "python3.12", "-u", "/app/main.py"]

prepare_container = [
    { type = "COPY", source = "requirements.txt", target = "requirements.txt" },
    { type = "RUN", command = "python3.12 -m pip install --no-cache-dir -r /app/requirements.txt" },
]

optional_devices = [
    "/dev/fastrpc-cdsp",
    "/dev/adsprpc-smd",
    "/dev/dma_heap/qcom,system",
    "/dev/dma_heap/system",
]
required_mounts = [
    "/opt/luxonis/npu-runtime:/opt/luxonis/npu-runtime:ro,rbind",
]
allowed_devices = [{ allow = true, access = "rw" }]
```

This configuration uses a required NPU runtime mount, as in the VLM example, so deployment fails early if that runtime is
unavailable. An app designed to support CPU-only operation can use `optional_mounts` instead, but must also select the appropriate
CPU server options. See [`oakapp.toml` configuration](https://docs.luxonis.com/software-v3/oak-apps/configuration.md) for the
complete manifest reference.

## Prepare the Model

Use a GGUF export supported by the image's pinned llama.cpp version. GGUF is a file format, not a guarantee that every
architecture, quantization, or operation can execute on Hexagon. Start with the model used in the [llama.cpp VLM
example](https://github.com/luxonis/oak-examples/tree/main/apps/llamacpp-vlm) when validating the setup:

 * Language model: `Qwen3.5-0.8B-Q4_0.gguf`.
 * Vision projector: `mmproj-F16.gguf` from the same model repository and revision.
 * Source:
   [`unsloth/Qwen3.5-0.8B-GGUF`](https://huggingface.co/unsloth/Qwen3.5-0.8B-GGUF/tree/6ab461498e2023f6e3c1baea90a8f0fe38ab64d0).

The projector contains the model-specific image encoding and projection components. Do not substitute a projector from a different
model just because its filename is also `mmproj-F16.gguf`. See the upstream [multimodal
documentation](https://github.com/ggml-org/llama.cpp/blob/7fe450e19305b828c199d602c23a8337aaa1f03b/tools/mtmd/README.md) for model
preparation details.

You can include local weights in the app directory, for example `/app/models/model.gguf` and `/app/models/projector.gguf`, or
download them into the runtime container layer through `prepare_container`. The VLM example's
[`models.py`](https://github.com/luxonis/oak-examples/tree/main/apps/llamacpp-vlm/backend/src/models.py) pins a repository
revision, verifies SHA-256 checksums, and stores both files under `/opt/qwen-models` during the build. This keeps inference
independent of model downloads at runtime.

## Start the Inference Server

Inside the app container, the following command starts a vision-capable server with local model files:

```bash
llama-server \
    --model /app/models/model.gguf \
    --mmproj /app/models/projector.gguf \
    --device HTP0 --mmproj-device HTP0 -ngl 99 \
    --host 127.0.0.1 --port 8081 \
    --ctx-size 2048 --parallel 1 \
    --threads 4 --threads-batch 4 \
    --batch-size 256 --ubatch-size 256 \
    --jinja
```

`--device HTP0` selects the Hexagon backend for language-model offload, `-ngl 99` requests offloading up to 99 layers, and
`--mmproj-device HTP0` requests Hexagon execution for the vision projector. These settings request placement; unsupported
operations may still execute on the CPU. Check the server logs for actual model and projector placement.

| Mode | Language-model options | Vision-projector options |
| --- | --- | --- |
| Hexagon | `--device HTP0 -ngl 99` | `--mmproj /path/to/projector.gguf --mmproj-device HTP0` |
| CPU only | `--device none -ngl 0` | `--mmproj /path/to/projector.gguf --no-mmproj-offload` |

For text-only models, omit all vision-projector options.

In a Python app, launch the command with `subprocess.Popen` and keep the process alive while serving requests. Wait for an HTTP
success response from `http://127.0.0.1:8081/health` before enabling inference, with a startup timeout and a check for early
process exit. Forward server output to app logs; if you pipe stdout or stderr, drain it continuously so a full pipe cannot block
the server. Terminate and wait for the child process when the app shuts down. The example's
[`runtime.py`](https://github.com/luxonis/oak-examples/tree/main/apps/llamacpp-vlm/backend/src/runtime.py) implements this
lifecycle and streamed requests.

## Send a Camera Frame and Prompt

The following Python snippet runs inside the app, after the local server is ready. It requires `depthai`,
`opencv-python-headless`, and `requests` in `requirements.txt`. It captures one frame, encodes it as a JPEG data URL, and requests
a non-streamed answer.

```python
import base64

import cv2
import depthai as dai
import requests

with dai.Pipeline() as pipeline:
    camera = pipeline.create(dai.node.Camera).build(dai.CameraBoardSocket.CAM_A)
    frames = camera.requestOutput(size=(448, 252)).createOutputQueue(
        maxSize=1, blocking=False
    )
    pipeline.start()
    frame = frames.get().getCvFrame()

ok, jpeg = cv2.imencode(".jpg", frame)
if not ok:
    raise RuntimeError("Could not encode the camera frame")
image_url = "data:image/jpeg;base64," + base64.b64encode(jpeg).decode("ascii")

response = requests.post(
    "http://127.0.0.1:8081/v1/chat/completions",
    json={
        "messages": [{
            "role": "user",
            "content": [
                {"type": "image_url", "image_url": {"url": image_url}},
                {"type": "text", "text": "Describe what you see in this image."},
            ],
        }],
        "stream": False,
        "max_tokens": 128,
        "temperature": 0.0,
        # Qwen3.5-specific template setting; adapt for other models.
        "chat_template_kwargs": {"enable_thinking": False},
    },
    timeout=(10, 180),
)
response.raise_for_status()
print(response.json()["choices"][0]["message"]["content"])
```

For a continuous camera app, keep capture running independently of generation and retain the latest frame in a bounded queue. For
region prompts, crop the original frame before resizing so the selected region retains useful detail. The VLM example preserves
the image aspect ratio and downsizes to a maximum edge of 448 pixels by default; this is an application setting, not a fixed
input-shape requirement for all models. llama.cpp handles the selected model's internal vision preprocessing.

To stream answers, set `stream` to `true` and consume the server-sent events from the response, reading answer text from
`choices[].delta.content` until the stream completes. Use request timeouts and surface server errors to the caller.

## Run the Complete VLM Example

The [llama.cpp VLM example](https://github.com/luxonis/oak-examples/tree/main/apps/llamacpp-vlm) combines the server lifecycle,
model download, camera pipeline, and frontend. It uses Qwen3.5-0.8B Q4_0 with its F16 projector, requests Hexagon execution for
both, and processes one inference request at a time. The frontend supports full-frame questions, region selection, streamed
answers, and latency reporting.

From the example directory, use [`oakctl`](https://docs.luxonis.com/software-v3/oak-apps/oakctl.md) to run it in development mode:

```bash
oakctl connect <DEVICE_IP>
oakctl app run .
```

The first container build needs internet access to download dependencies and approximately 712 MB of model assets. Open the
frontend URL printed by `oakctl`, typically `https://<DEVICE_IP>:9000/`, and wait for the model to become ready before submitting
a prompt.

## Performance and Limitations

 * Model compatibility: support depends on the pinned llama.cpp version, model architecture, and tensor types. Validate another
   GGUF model on the device before adopting it.
 * Mixed execution: selecting `HTP0` does not guarantee that every operation runs on the HTP. Inspect logs for language-model and
   projector placement, and compare against explicit CPU-only execution.
 * Memory use: weights are only part of the requirement. Context size, image tokens, batches, caches, and concurrent requests also
   consume memory. Start with a small model and one request at a time.
 * Latency: image encoding, prompt evaluation, and token generation are separate costs. A camera stream's FPS does not represent
   VLM inference throughput. Smaller images or crops and shorter outputs can reduce work, with possible loss of detail or answer
   completeness.
 * Measurement: separate model startup from request latency. The example measures time from the model HTTP request to the first
   and last answer-text events, including vision encoding and prompt evaluation but excluding app-side crop/JPEG preparation and
   frontend transport.
 * Application integration: frame selection, preprocessing, server lifecycle, request scheduling, and result handling belong to
   the app rather than a DepthAI `NeuralNetwork` node.

## Troubleshooting

| Symptom | What to check |
| --- | --- |
| HTP is unavailable or FastRPC initialization fails | Confirm Luxonis OS 1.40.0 or newer, device passthrough, the NPU runtime
mount, and use of `/entrypoint.sh`. Inspect the server's startup logs. |
| Model or projector fails to load | Verify local file paths and checksums, the matching model/projector revision, and support in
the image's llama.cpp version. |
| Server is running but inference is not ready | Wait for `/health` to succeed. Check for model-loading errors or process exit
rather than relying only on a startup delay. |
| Out-of-memory or context-capacity errors | Reduce context size, image-token demand, batch size, or concurrency; use a smaller
supported model or quantization if needed. |
| Unexpectedly slow inference | Check actual HTP placement, image size, prompt length, output-token limit, and other workloads on
the device. |
| Model download options fail | This image disables built-in HTTPS downloads. Package local GGUF files or download and verify them
in the app's build tooling. |
