# infer_utils

Python API: `luxonis_train.core.utils.infer_utils`

Inference over an image, a video, a directory, or a dataset view.

[LuxonisModel.infer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/core.md)
calls these helpers. They run the model, convert the images of its visualizers to OpenCV arrays, and save the arrays or show them
in a window.
[LuxonisModel.annotate](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/core.md)
also reads its images through
[create_loader_from_directory](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md).

`VIDEO_FORMATS` holds the lowercase file extensions that
[LuxonisModel.infer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/core.md)
reads as a video. `IMAGE_FORMATS` holds the lowercase file extensions of the images that
[LuxonisModel.infer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/core.md)
and
[LuxonisModel.annotate](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/core.md)
select in a directory.
[LuxonisLoaderPerlinNoise](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/luxonis_perlin_loader_torch.md)
also selects its anomaly source images with it.

## Functions

### create_loader_from_directory

```python
def create_loader_from_directory(img_paths: Iterable[PathType], model: lxt.LuxonisModel, batch_size: int | None = None, return_sample_metadata: bool = False) -> torch_data.DataLoader:
```

Build a `DataLoader` over image files.

The function creates a local `LuxonisDataset` named `infer_from_directory`, after it deletes a local dataset of that name. It adds
every image with the sample metadata `{"path": <image path>}`, and puts all images in the `test` split. It wraps the dataset in a
[LuxonisLoaderTorch](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/luxonis_loader_torch.md)
on the `test` view. The loader applies these settings of `trainer.preprocessing` of `model`: the train image size, the active
augmentations, the color space, and `keep_aspect_ratio`.

The loader does not keep the order of `img_paths`. Use the `"path"` metadata to match a sample to its file. The function does not
delete the dataset.
[infer_from_directory](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md)
and
[annotate_from_directory](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/annotate_utils.md)
delete it after use.

Parameters

 * `img_paths` (`Iterable[PathType]`): The image files.
 * `model` (`lxt.LuxonisModel`): The model whose preprocessing the loader applies.
 * `batch_size` (`int | None`): The batch size. `None` selects `trainer.batch_size` of the config of `model`.
 * `return_sample_metadata` (`bool`): Also return the metadata of every sample. Each batch is then a tuple of the inputs, the
   labels, and a list with the metadata dictionary of every sample, each with the `"path"` key.

Returns

 * `torch_data.DataLoader`: The loader, with `pin_memory` on and without shuffling. Without `return_sample_metadata`, each batch
   is a list of the inputs and the labels. The inputs are one `Tensor` of shape `[B, C, H, W]`. The labels are an empty
   dictionary, because the dataset has no annotations.

### infer_from_dataset

```python
def infer_from_dataset(model: lxt.LuxonisModel, view: Literal['train', 'val', 'test'], save_dir: PathType | None):
```

Run the model on one view of its dataset.

The function reads `model.pytorch_loaders[view]`. When `trainer.overfit_batches` of the config is above `0` and `view` is
`"train"`, it reads only that many batches and logs a warning. It then runs
[infer_from_loader](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md)
without image paths, so a saved render is named `<node>_<visualizer>_<n>.png`.

Parameters

 * `model` (`lxt.LuxonisModel`): The model to run.
 * `view` (`Literal['train', 'val', 'test']`): The dataset view to read.
 * `save_dir` (`PathType | None`): The directory of the PNG files. `None` shows the renders on screen instead.

### infer_from_directory

```python
def infer_from_directory(model: lxt.LuxonisModel, img_paths: Iterable[PathType], save_dir: Path | None):
```

Run the model on image files.

The function builds a loader with
[create_loader_from_directory](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md)
and runs
[infer_from_loader](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md)
on it. It passes `img_paths` in the given order to name the saved renders. The loader does not keep that order, so with two or
more images a saved render can get the name of a different image. At the end, the function deletes the temporary local dataset
`infer_from_directory`.

Parameters

 * `model` (`lxt.LuxonisModel`): The model to run.
 * `img_paths` (`Iterable[PathType]`): The image files.
 * `save_dir` (`Path | None`): The directory of the PNG files. `None` shows the renders on screen instead.

### infer_from_loader

```python
def infer_from_loader(model: lxt.LuxonisModel, loader: torch_data.DataLoader, save_dir: PathType | None, img_paths: list[PathType] | None = None):
```

Run the prediction loop of the trainer over a loader and render the visualizations.

The function runs `model.pl_trainer.predict` with `model.lightning_module` on `loader`, so each batch goes through
[LuxonisLightningModule.predict_step](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/luxonis_lightning.md).

With `save_dir`, a temporary `BasePredictionWriter` callback saves the renders of each batch as PNG files as they arrive:

 * `<save_dir>/<image stem>_<node>_<visualizer>.png` when `img_paths` is given. The callback takes the next path for each sample,
   in loader order.
 * `<save_dir>/<node>_<visualizer>_<n>.png` otherwise, where `<n>` counts the written files from `0`.

A `/` in a file name becomes `-`.

Without `save_dir`, the function collects all predictions first. It then shows every render of every sample in an OpenCV window
named `<node>/<visualizer>`. It waits for a key press after every sample. `Esc` or `q` stops the loop. The function closes the
windows at the end.

Parameters

 * `model` (`lxt.LuxonisModel`): The model to run.
 * `loader` (`torch_data.DataLoader`): The batches to run on.
 * `save_dir` (`PathType | None`): The directory of the PNG files. `None` shows the renders on screen instead.
 * `img_paths` (`list[PathType] | None`): The source path of every sample, in loader order. It names the saved files. It is not
   read without `save_dir`.

### infer_from_video

```python
def infer_from_video(model: lxt.LuxonisModel, video_path: PathType, save_dir: Path | None):
```

Run the model on every frame of a video, one frame at a time.

The function reads the frames with `cv2.VideoCapture`. When `trainer.preprocessing.color_space` of the config is `"RGB"`, it
converts each frame from BGR to RGB first. It then runs
[prepare_and_infer_image](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md)
on the frame under the input name `"image"`, and renders the visualizations with
[process_visualizations](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md).
When OpenCV cannot open `video_path`, the function reads no frame and writes nothing.

With `save_dir`, the function writes the renders of each visualizer to `<save_dir>/<node>_<visualizer>.mp4`. The video gets the
`mp4v` codec, the frame rate of the source video, and the size of the first render. Without `save_dir`, the function shows the
render of each visualizer in an OpenCV window named `<node>/<visualizer>`. It then waits for a key press after every frame. `Esc`
or `q` stops the loop. At the end, the function releases the capture and the writers, and closes the windows.

Parameters

 * `model` (`lxt.LuxonisModel`): The model to run.
 * `video_path` (`PathType`): The video file.
 * `save_dir` (`Path | None`): The directory of the output videos. `None` shows the renders on screen instead.

### prepare_and_infer_image

```python
def prepare_and_infer_image(model: lxt.LuxonisModel, images: dict[str, Tensor]) -> LuxonisOutput:
```

Preprocess one raw image with the `val` loader and run the model on it.

The function passes `images` to `augment_test_image` of the `val` loader of `model`. That method returns one image of shape `[H,
W, C]`.
[LuxonisLoaderTorch](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/luxonis_loader_torch.md)
applies its augmentations there, such as the resize and the normalization. A
[LuxonisLoaderTorch](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/luxonis_loader_torch.md)
without a height or a width returns the image unchanged. A loader that does not override the method raises `NotImplementedError`.
The function adds the batch dimension, moves the channels first, and casts to `torch.float32`. It then runs
[LuxonisLightningModule.full_forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/luxonis_lightning.md)
on this batch of one, under the `image_source` name of the module. The visualizers draw on the denormalized image. The function
gives no labels, so no loss and no metric runs.

Parameters

 * `model` (`lxt.LuxonisModel`): The model to run.
 * `images` (`dict[str, Tensor]`): The raw image of shape `[H, W, C]`, keyed by the input name of the loader.

Returns

 * `LuxonisOutput`: `outputs` holds the packet of every output node. `visualizations` holds the image batch of every visualizer,
   with a batch size of `1`. `losses` and `metrics` are empty.

### process_visualizations

```python
def process_visualizations(visualizations: dict[str, dict[str, Tensor]]) -> dict[tuple[str, str], list[np.ndarray]]:
```

Convert the images of the visualizers into OpenCV arrays.

The function takes the `visualizations` of a
[LuxonisOutput](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/luxonis_output.md)
and turns every image of every batch into a NumPy array. It detaches each image, moves it to the CPU, and transposes it from `[C,
H, W]` to `[H, W, C]`. It then converts the channels from RGB to BGR with `cv2.cvtColor`, because OpenCV writes and shows BGR. An
image with four channels loses its fourth channel.

> **Example**
> ```pycon
>>> import torch
>>> image = torch.zeros(2, 3, 4, 5, dtype=torch.uint8)
>>> image[:, 0] = 255  # red in RGB
>>> renders = process_visualizations({"head": {"boxes": image}})
>>> sorted(renders)
[('head', 'boxes')]
>>> [render.shape for render in renders["head", "boxes"]]
[(4, 5, 3), (4, 5, 3)]
>>> renders["head", "boxes"][0][0, 0].tolist()
[0, 0, 255]
```

Parameters

 * `visualizations` (`dict[str, dict[str, Tensor]]`): The image batch of every visualizer, of shape `[B, C, H, W]`, keyed by node name and visualizer name.

Returns

 * `dict[tuple[str, str], list[np.ndarray]]`: The `B` images of every visualizer, each of shape `[H, W, 3]` in BGR, keyed by the pair of node name and visualizer name. The result is a `defaultdict`, so a missing key gives an empty list.

### window_closed

```python
def window_closed() -> bool:
```

Wait for a key press in the OpenCV windows and report a stop request.

The function blocks in `cv2.waitKey` until the user presses a key.

Returns

 * `bool`: `True` when the key is `Esc` or `q`.

## Attributes

### IMAGE_FORMATS

### VIDEO_FORMATS
