# base_loader

Python API: `luxonis_train.loaders.base_loader`

The base class every loader inherits, and the type of one sample.

[LuxonisLoaderTorchOutput](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md)
is the type of one sample. It is a pair of the input and the labels. The input is an image of shape `[C, H, W]`, or a dictionary
that maps each input name to its image. The labels map each `"<task_name>/<label>"` key to a tensor.

## Classes

### BaseLoaderTorch

Base class for the loaders of the training pipeline.

A subclass registers in
[luxonis_train.registry.LOADERS](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/registry.md)
under its class name, so the `loader.name` field of a config can name it. A subclass must implement
[input_shapes](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md),
`__len__`, `__getitem__`, and
[get_classes](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md).
A loader with keypoint labels must also override
[get_n_keypoints](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md).

#### Methods

##### init

```python
def __init__(view: list[str], height: int | None = None, width: int | None = None, augmentation_engine: str = 'albumentations', augmentation_config: list[AugmentationConfig] | None = None, image_source: str = 'image', keep_aspect_ratio: bool = True, color_space: Literal['RGB', 'BGR', 'GRAY'] = 'RGB', seed: int | None = None):
```

Store the settings that every loader shares.

Parameters

 * `view` (`list[str]`): The splits that form the view. The list usually holds one split, such as `["train"]`. A dataset can also
   combine splits, such as `["train_synthetic", "train_real"]`.
 * `height` (`int | None`): The height of the output image. With `None`, the
   [height](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md)
   property raises `ValueError`.
 * `width` (`int | None`): The width of the output image. With `None`, the
   [width](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md)
   property raises `ValueError`.
 * `augmentation_engine` (`str`): The name of the augmentation engine, such as `"albumentations"`.
 * `augmentation_config` (`list[AugmentationConfig] | None`): The augmentations. Each item has a `name` and a `params` dictionary.
   With `None`, the
   [augmentation_config](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md)
   property raises `ValueError`.
 * `image_source` (`str`): The name of the main image source. For a dataset with more than one source, such as `"left"` and
   `"right"`, the visualizations use this source.
 * `keep_aspect_ratio` (`bool`): Whether the resize keeps the aspect ratio of the image.
 * `color_space` (`Literal['RGB', 'BGR', 'GRAY']`): The color space of the output image.
 * `seed` (`int | None`): The random seed of the augmentations, or `None`.

##### augment_test_image

```python
def augment_test_image(img: dict[str, Tensor] | Tensor) -> Tensor:
```

Apply the augmentations of the loader to one raw image.

Inference calls this method to prepare an image like the samples of the view. The base implementation only raises. A loader that
supports inference on raw images must override it.
[LuxonisLoaderTorch](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/luxonis_loader_torch.md)
overrides it.

Parameters

 * `img` (`dict[str, Tensor] | Tensor`): The raw image of shape `[H, W, C]`. A dictionary maps each source name to its image.

Returns

 * `Tensor`: An override returns the augmented image of shape `[H, W, C]`.

Raises

 * `NotImplementedError`: Always, in the base implementation.

##### collate_fn

```python
def collate_fn(batch: list[LuxonisLoaderTorchOutput]) -> tuple[dict[str, Tensor] | Tensor, Labels]:
```

Merge a list of samples into one batch.

The method stacks the inputs along a new first dimension. For dictionary inputs, it stacks each input name separately. It merges
each label of the first sample by the label type:

 * `boundingbox` and `keypoints`: The method adds the index of the sample as a new first column, and joins the rows of all
   samples. A `[N, 5]` box label becomes `[N, 6]`.
 * `instance_segmentation`: The method joins the masks of all samples along the first dimension.
 * `metadata/text`: The method pads the character codes of each sample with zeros to the longest text. The result is a
   `torch.int32` tensor of shape `[B, S]`.
 * Other `metadata/<name>` labels: The method joins the values of all samples along the first dimension.
 * Other labels, such as `classification` and `segmentation`: The method stacks them along a new first dimension.

Parameters

 * `batch` (`list[LuxonisLoaderTorchOutput]`): The samples. All samples must have the label keys of the first sample.

Returns

 * `tuple[dict[str, Tensor] | Tensor, Labels]`: The batched input and the batched labels, with the keys of the first sample.

Raises

 * `TypeError`: If the batch mixes tensor inputs and dictionary inputs.

##### dict_numpy_to_torch

```python
def dict_numpy_to_torch(numpy_dictionary: dict[str, np.ndarray]) -> dict[str, Tensor]:
```

Convert a dictionary of NumPy arrays to `torch.float32` tensors.

A string array becomes the character codes of its first string. The method converts the codes to `torch.float32` too.

Parameters

 * `numpy_dictionary` (`dict[str, np.ndarray]`): The arrays, such as the labels of one sample.

Returns

 * `dict[str, Tensor]`: A new dictionary with the same keys and one tensor for each array.

##### get_categorical_encodings

```python
def get_categorical_encodings(self) -> dict[str, dict[str, int]]:
```

Return the integer code of each category of each metadata label.

The base implementation returns an empty dictionary, which means that the loader has no categorical metadata labels.

Returns

 * `dict[str, dict[str, int]]`: The category to code mapping of each categorical metadata label, keyed by the label name.

##### get_classes

```python
def get_classes(self) -> dict[str, dict[str, int]]:
```

Return the class names and the class IDs of each task.

Returns

 * `dict[str, dict[str, int]]`: The class name to class ID mapping of each task, keyed by the task name.

##### get_metadata_types

```python
def get_metadata_types(self) -> dict[str, type[int] | type[Category] | type[float] | type[str]]:
```

Return the Python type of each metadata label.

The base implementation returns an empty dictionary, which means that the loader has no metadata labels.
[DatasetMetadata](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/dataset_metadata.md)
reads the result.

Returns

 * `dict[str, type[int] | type[Category] | type[float] | type[str]]`: The type of each metadata label, keyed by the label name,
   such as `"task_name/metadata/color"`.

##### get_n_keypoints

```python
def get_n_keypoints(self) -> dict[str, int] | None:
```

Return the number of keypoints of each task.

The base implementation returns `None`. A loader with keypoint labels must override it.

Returns

 * `dict[str, int] | None`: The number of keypoints, keyed by the task name, or `None` when the loader has no keypoints.

##### img_numpy_to_torch

```python
def img_numpy_to_torch(img: np.ndarray) -> Tensor:
```

Convert a NumPy image to a `torch.float32` tensor.

The method moves the channels of a 3D image first. A 2D image keeps its shape.

> **Example**
> ```pycon
>>> import numpy as np
>>> image = np.zeros((4, 6, 3), dtype=np.uint8)
>>> BaseLoaderTorch.img_numpy_to_torch(image).shape
torch.Size([3, 4, 6])
>>> gray = np.zeros((4, 6), dtype=np.uint8)
>>> BaseLoaderTorch.img_numpy_to_torch(gray).shape
torch.Size([4, 6])
```

Parameters

 * `img` (`np.ndarray`): The image of shape `[H, W, C]` or `[H, W]`.

Returns

 * `Tensor`: The image of shape `[C, H, W]` or `[H, W]`.

##### read_image

```python
def read_image(path: str) -> npt.NDArray[np.uint8]:
```

Read an image file into an unnormalized NumPy array.

OpenCV reads the file as a BGR color image. The method then converts it to [color_space](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md).

Parameters

 * `path` (`str`): The path to the image file.

Returns

 * `npt.NDArray[np.uint8]`: The image of shape `[H, W, 3]`, or `[H, W]` for the `"GRAY"` color space.

Raises

 * `ValueError`: If OpenCV cannot read the file.

#### Attributes

##### augmentation_config

The augmentations of the loader.

The property raises `ValueError` when the constructor got `None`.

##### augmentation_engine

The name of the augmentation engine.

##### color_space

The color space of the output image.

The value is `"RGB"`, `"BGR"`, or `"GRAY"`.

##### height

The height of the output image.

The property raises `ValueError` when the constructor got `None`.

##### image_source

The name of the main image source, such as `"image"`.

##### input_shape

The shape `[C, H, W]` of the [image_source](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md) input.

The shape comes from [input_shapes](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md) and has no batch dimension.

##### input_shapes

The shape of each input of one sample, keyed by the input name.

An implementation returns one shape for each input, without the batch dimension. An image has the shape `[C, H, W]`. The result must hold the [image_source](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md) key, because [input_shape](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/loaders/base_loader.md) reads it.

> **Examples**
> A loader with one image:

```python
{
    "image": torch.Size([3, 224, 224]),
}
```

A loader with an image and a segmentation input:

```python
{
    "image": torch.Size([3, 224, 224]),
    "segmentation": torch.Size([1, 224, 224]),
}
```

A loader with a left image, a right image, and a disparity map:

```python
{
    "left": torch.Size([3, 224, 224]),
    "right": torch.Size([3, 224, 224]),
    "disparity": torch.Size([1, 224, 224]),
}
```

A loader with an image, keypoints, and a point cloud:

```python
{
    "image": torch.Size([3, 224, 224]),
    "keypoints": torch.Size([17, 2]),
    "point_cloud": torch.Size([20000, 3]),
}
```

##### keep_aspect_ratio

Whether the resize keeps the aspect ratio of the image.

##### seed

The random seed of the augmentations, or `None`.

##### view

The splits that form the view, such as `["train"]`.

##### width

The width of the output image.

The property raises `ValueError` when the constructor got `None`.

## Attributes

### LuxonisLoaderTorchOutput

### MIXED_INPUT_TYPES_ERROR
