# utils

Python API: `luxonis_train.attached_modules.visualizers.utils`

Helpers for the visualizers.

The helpers convert between tensors and images, denormalize the input images, select colors and font sizes, draw labels, and
combine the label image and the prediction image into one image. `Color` is the type of a color: a color name or a hex string such
as `"#FF0000"`, or an RGB tuple.

## Functions

### combine_visualizations

```python
def combine_visualizations(visualization: Tensor | tuple[Tensor, Tensor] | tuple[Tensor, list[Tensor]]) -> Tensor:
```

Combine the output of a visualizer into one image batch.

The trainer calls this function on the result of each
[BaseVisualizer.run](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/base_visualizer.md):

 * A single tensor: the function returns it unchanged.
 * A pair `(labels, predictions)` of tensors, in a tuple or a list: the function resizes both batches to the larger height. Each
   batch keeps its aspect ratio. The function then puts the batches side by side, with the labels on the left.
 * A tensor and a list or a tuple of tensors: the function raises `NotImplementedError`.

> **Example**
> ```pycon
>>> import torch
>>> labels = torch.zeros(1, 3, 4, 4, dtype=torch.uint8)
>>> predictions = torch.zeros(1, 3, 8, 6, dtype=torch.uint8)
>>> combine_visualizations((labels, predictions)).shape
torch.Size([1, 3, 8, 14])
>>> combine_visualizations(labels) is labels
True
```

Parameters

 * `visualization` (`Tensor | tuple[Tensor, Tensor] | tuple[Tensor, list[Tensor]]`): The output of a visualizer. The images have the shape `[B, C, H, W]`, and the two batches of a pair have the same `B` and `C`.

Returns

 * `Tensor`: The images of shape `[B, C, H, W]`. For a pair, `H` is the larger height and `W` is the sum of the resized widths.

Raises

 * `NotImplementedError`: When the second item is a list or a tuple of tensors.
 * `ValueError`: When `visualization` has any other form.

### denormalize

```python
def denormalize(img: Tensor, mean: list[float] | float | None = None, std: list[float] | float | None = None, to_uint8: bool =
False) -> Tensor:
```

Undo the normalization of an image.

For each channel c, the function computes

yc = xcσc + μc

where μ is `mean` and σ is `std`. With `to_uint8`, it then multiplies the result by `255`, clips it to `[0, 255]`, and truncates it to `uint8`.

> **Example**
> ```pycon
>>> import torch
>>> img = torch.tensor([[[0.0]], [[1.0]], [[-1.0]]])
>>> denormalize(img, mean=0.5, std=0.25).flatten().tolist()
[0.5, 0.75, 0.25]
>>> image = denormalize(img, mean=0.5, std=0.25, to_uint8=True)
>>> image.flatten().tolist()
[127, 191, 63]
```

Parameters

 * `img` (`Tensor`): A normalized image of shape `[C, H, W]`.
 * `mean` (`list[float] | float | None`): The mean of the normalization, one value for each channel or one `float` for all
   channels. `None` selects `0`.
 * `std` (`list[float] | float | None`): The standard deviation of the normalization, in the same form as `mean`. `None` selects
   `1`.
 * `to_uint8` (`bool`): Whether to scale the result to a `uint8` image.

Returns

 * `Tensor`: A new image of shape `[C, H, W]`, of dtype `uint8` with `to_uint8`.

### draw_bounding_box_labels

```python
def draw_bounding_box_labels(img: Tensor, label: Tensor, **kwargs) -> Tensor:
```

Draw normalized bounding boxes on an image.

The function converts the boxes from normalized `xywh`, where `x` and `y` are the top-left corner, to pixel `xyxy` with the image
size. It then calls `torchvision.utils.draw_bounding_boxes`. It does not change `label`.

> **Example**
> The box `[0.25, 0.25, 0.5, 0.5]` on an image of width `8` and height `4` becomes the pixel box `[2, 1, 6, 3]`:

```pycon
>>> import torch
>>> img = torch.zeros(3, 4, 8, dtype=torch.uint8)
>>> label = torch.tensor([[0.25, 0.25, 0.5, 0.5]])
>>> draw_bounding_box_labels(img, label, colors="red")[0].tolist()
[[0, 0, 0, 0, 0, 0, 0, 0],
 [0, 0, 255, 255, 255, 255, 255, 0],
 [0, 0, 255, 0, 0, 0, 255, 0],
 [0, 0, 255, 255, 255, 255, 255, 0]]
```

Parameters

 * `img` (`Tensor`): A `uint8` image of shape `[C, H, W]`.
 * `label` (`Tensor`): Boxes of shape `[N, 4]`, with rows `[x, y, w, h]` normalized to `[0, 1]`.
 * `**kwargs`: Keyword arguments forwarded to `draw_bounding_boxes`, such as `labels`, `colors`, and `width`.

Returns

 * `Tensor`: A new image of the same shape as `img`, with the boxes drawn.

### draw_keypoint_labels

```python
def draw_keypoint_labels(img: Tensor, label: Tensor, **kwargs) -> Tensor:
```

Draw normalized keypoints on an image.

The function scales the `x` and `y` values by the image width and height, truncates them to integers, and calls
`torchvision.utils.draw_keypoints`. It draws every keypoint and ignores the visibility.

Side effect: when `label` is contiguous, the function scales the `x` and `y` values of `label` in place.

> **Example**
> The keypoint lands on the pixel in row `1` and column `2`. The call also changes `label` to pixel coordinates:

```pycon
>>> import torch
>>> img = torch.zeros(3, 4, 4, dtype=torch.uint8)
>>> label = torch.tensor([[0.5, 0.25, 2.0]])
>>> out = draw_keypoint_labels(img, label, colors="red", radius=1)
>>> out[0, 1, 2].item(), label.tolist()
(255, [[2.0, 1.0, 2.0]])
```

Parameters

 * `img` (`Tensor`): A `uint8` image of shape `[C, H, W]`.
 * `label` (`Tensor`): Keypoints of shape `[N, 3K]`, with rows `[x_1, y_1, v_1, ..., x_K, y_K, v_K]`. The coordinates are
   normalized to `[0, 1]`.
 * `**kwargs`: Keyword arguments forwarded to `draw_keypoints`, such as `colors`, `radius`, and `connectivity`.

Returns

 * `Tensor`: A new image of the same shape as `img`, with the keypoints drawn. `img` itself when `label` holds no keypoints.

### draw_segmentation_targets

```python
def draw_segmentation_targets(image: Tensor, target: Tensor, alpha: float = 0.4, colors: Color | list[Color] | None = None) -> Tensor:
```

Blend segmentation masks into an image.

The function moves the image and the masks to the CPU and calls `torchvision.utils.draw_segmentation_masks`. A pixel in more than
one mask gets the color black before the blend.

> **Example**
> ```pycon
>>> import torch
>>> image = torch.zeros(3, 1, 3, dtype=torch.uint8)
>>> masks = torch.tensor([[[1, 1, 0]], [[0, 1, 1]]])
>>> out = draw_segmentation_targets(
...     image, masks, alpha=1.0, colors=["red", "blue"]
... )
>>> out[0].tolist(), out[2].tolist()
([[255, 0, 0]], [[0, 0, 255]])
```

Parameters

 * `image` (`Tensor`): An RGB image of shape `[3, H, W]`, of dtype `uint8`, or floating point with values in `[0, 1]`.
 * `target` (`Tensor`): Masks of shape `[N, H, W]`, or one mask of shape `[H, W]`. Every non-zero value marks a pixel of the mask.
 * `alpha` (`float`): The opacity of the masks, from `0` for transparent to `1` for opaque.
 * `colors` (`Color | list[Color] | None`): One color for each mask, or one color for all masks. When `None`, `torchvision` selects the colors.

Returns

 * `Tensor`: A new image on the CPU, of the same shape and dtype as `image`. When `target` holds no masks, `torchvision` issues a warning, and the function returns the image unchanged.

### dynamically_determine_font_scale

```python
def dynamically_determine_font_scale(height: int, width: int, thickness: int, font_scale: float | None = None, scale_factor: float
= 500.0) -> tuple[float, int]:
```

Select a font scale and a line thickness for an image size.

Without `font_scale`, the function derives the scale from a weighted mean of the height H and the width W:

w = min⎛⎝0.4, (W)/(10max(H, 1))⎞⎠, s = ((1 − w)H + w**W)/(scale_factor)

A scale below `1` gets the thickness `1`. A scale of `1` or more keeps `thickness`.

> **Example**
> ```pycon
>>> dynamically_determine_font_scale(500, 500, thickness=2)
(1.0, 2)
>>> dynamically_determine_font_scale(100, 200, thickness=2)
(0.24, 1)
```

Parameters

 * `height` (`int`): The image height, in pixels.
 * `width` (`int`): The image width, in pixels.
 * `thickness` (`int`): The line thickness for a scale of `1` or more.
 * `font_scale` (`float | None`): A fixed font scale. `None` selects the computed scale.
 * `scale_factor` (`float`): The effective image size, in pixels, that gives the scale `1`.

Returns

 * `tuple[float, int]`: The font scale and the line thickness.

### figure_to_torch

```python
def figure_to_torch(fig: Figure, width: int, height: int) -> Tensor:
```

Render a matplotlib figure to an RGB image tensor and close it.

The function saves the figure as a PNG image with a tight bounding box and no padding, and resizes the image to `width` and
`height`. It then closes the figure with `plt.close`.

> **Example**
> ```pycon
>>> import matplotlib.pyplot as plt
>>> fig, ax = plt.subplots()
>>> _ = ax.plot([0, 1], [0, 1])
>>> figure_to_torch(fig, width=64, height=32).shape
torch.Size([3, 32, 64])
```

Parameters

 * `fig` (`Figure`): The matplotlib figure to render.
 * `width` (`int`): The width of the image, in pixels.
 * `height` (`int`): The height of the image, in pixels.

Returns

 * `Tensor`: A `uint8` image of shape `[3, height, width]`.

### get_color

```python
def get_color(seed: int) -> Color:
```

Return a distinct RGB color for an integer.

The color is not random. The same `seed` always gives the same color. The function adds `45` to `seed` and converts the result with [number_to_hsl](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/utils.md) and [hsl_to_rgb](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/utils.md). The visualizers pass the class index as `seed`.

> **Example**
> ```pycon
>>> get_color(0)
(25, 76, 229)
```

Parameters

 * `seed` (`int`): The integer that selects the color.

Returns

 * `Color`: The red, green, and blue values, in `[0, 255]`.

### get_denormalized_images

```python
def get_denormalized_images(cfg: Config, images: Tensor) -> Tensor:
```

Convert a batch of model input images to `uint8` images.

When `trainer.preprocessing.normalize` is active, the function reads `mean` and `std` from its `params`, and
[preprocess_images](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/utils.md)
undoes the normalization. A missing key selects the ImageNet value, `[0.485, 0.456, 0.406]` for `mean` and `[0.229, 0.224, 0.225]`
for `std`. When the normalization is not active, the function only casts the images to `uint8`.

[LuxonisLightningModule](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/luxonis_lightning.md)
and the inference utilities use the result as the canvas of the visualizers.
[GradCamCallback](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/gradcam_visualizer.md)
draws its heat maps on it.

Parameters

 * `cfg` (`Config`): The config of the model.
 * `images` (`Tensor`): The input images of shape `[B, C, H, W]`, as the loader returns them.

Returns

 * `Tensor`: A new `uint8` tensor of shape `[B, C, H, W]`.

### get_prediction_labels

```python
def get_prediction_labels(prediction: Tensor, label_dict: Mapping[int, str] | None, draw_labels: bool, draw_scores: bool) -> list[str] | None:
```

Build the text label of each predicted box.

A label holds the class name when `draw_labels` is set, and the confidence with two decimals when `draw_scores` is set. A space
separates the two parts.

> **Example**
> ```pycon
>>> import torch
>>> prediction = torch.tensor(
...     [
...         [0.0, 0.0, 4.0, 4.0, 0.75, 1.0],
...         [1.0, 1.0, 2.0, 2.0, 0.5, 7.0],
...     ]
... )
>>> get_prediction_labels(prediction, {1: "cat"}, True, True)
['cat 0.75', '7 0.50']
>>> get_prediction_labels(prediction, None, False, True)
['0.75', '0.50']
```

Parameters

 * `prediction` (`Tensor`): Boxes of shape `[M, 6]`, with rows `[x1, y1, x2, y2, conf, class]`.
 * `label_dict` (`Mapping[int, str] | None`): Class names by class index. A class without a name, or every class when `None`, gets its index as the name.
 * `draw_labels` (`bool`): Whether the labels hold the class names.
 * `draw_scores` (`bool`): Whether the labels hold the confidences.

Returns

 * `list[str] | None`: One label for each box, or `None` when both `draw_labels` and `draw_scores` are `False`.

### hsl_to_rgb

```python
def hsl_to_rgb(hsl: tuple[float, float, float]) -> Color:
```

Convert an HSL color to an 8-bit RGB color.

The function truncates each channel to an integer.

> **Example**
> ```pycon
>>> hsl_to_rgb((0, 1.0, 0.5))
(255, 0, 0)
```

Parameters

 * `hsl` (`tuple[float, float, float]`): The hue in degrees, in `[0, 360)`, and the saturation and the lightness, in `[0, 1]`.

Returns

 * `Color`: The red, green, and blue values, in `[0, 255]`.

### number_to_hsl

```python
def number_to_hsl(seed: int) -> tuple[float, float, float]:
```

Map an integer to a hue, with fixed saturation and lightness.

The hue is `(seed * 157) % 360`. The prime factor spreads the hues of consecutive seeds over the color wheel. The saturation is
`0.8` and the lightness is `0.5`.

> **Example**
> ```pycon
>>> number_to_hsl(3)
(111, 0.8, 0.5)
```

Parameters

 * `seed` (`int`): The integer to map.

Returns

 * `tuple[float, float, float]`: The hue in degrees, in `[0, 360)`, the saturation, and the lightness.

### numpy_to_torch_img

```python
def numpy_to_torch_img(img: np.ndarray) -> Tensor:
```

Convert a NumPy image to a torch image.

The result shares its memory with `img` and keeps its dtype.

> **Example**
> ```pycon
>>> import numpy as np
>>> numpy_to_torch_img(np.zeros((4, 6, 3), dtype=np.uint8)).shape
torch.Size([3, 4, 6])
```

Parameters

 * `img` (`np.ndarray`): An image of shape `[H, W, C]`.

Returns

 * `Tensor`: A view of shape `[C, H, W]`.

### potentially_upscale_masks

```python
def potentially_upscale_masks(image_masks: Tensor, scale: float = 1.0) -> Tensor:
```

Resize segmentation masks by a factor with nearest interpolation.

The new height and width are `int(H * scale)` and `int(W * scale)`, so a factor below `1` shrinks the masks. With a factor of `1`,
the function returns `image_masks` unchanged.

> **Example**
> ```pycon
>>> import torch
>>> masks = torch.tensor([[[True, False], [False, True]]])
>>> potentially_upscale_masks(masks, scale=2.0).int().tolist()
[[[1, 1, 0, 0], [1, 1, 0, 0], [0, 0, 1, 1], [0, 0, 1, 1]]]
```

Parameters

 * `image_masks` (`Tensor`): Masks of shape `[N, H, W]`. Every non-zero value marks a pixel of a mask.
 * `scale` (`float`): The resize factor.

Returns

 * `Tensor`: Boolean masks of shape `[N, int(H * scale), int(W * scale)]`, or `image_masks` itself when `scale` is `1`.

### preprocess_images

```python
def preprocess_images(imgs: Tensor, mean: list[float] | float | None = None, std: list[float] | float | None = None) -> Tensor:
```

Convert a batch of model input images to `uint8` images.

When `mean` or `std` is given, [denormalize](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/utils.md) restores each image and scales it to `[0, 255]`. Otherwise the function only casts the images to `uint8`, so they must already have values in `[0, 255]`.

> **Example**
> ```pycon
>>> import torch
>>> imgs = torch.zeros(2, 3, 4, 4)
>>> out = preprocess_images(imgs, mean=[0.5, 0.5, 0.5], std=0.5)
>>> out.dtype, out[0, :, 0, 0].tolist()
(torch.uint8, [127, 127, 127])
```

Parameters

 * `imgs` (`Tensor`): Images of shape `[B, C, H, W]`.
 * `mean` (`list[float] | float | None`): The mean of the normalization, one value for each channel or one value for all channels.
 * `std` (`list[float] | float | None`): The standard deviation of the normalization, in the same form as `mean`.

Returns

 * `Tensor`: A new `uint8` tensor of shape `[B, C, H, W]`.

### torch_img_to_numpy

```python
def torch_img_to_numpy(img: Tensor, reverse_colors: bool = False) -> npt.NDArray[np.uint8]:
```

Convert a torch image to a `uint8` NumPy image.

The function multiplies a floating point image by `255` and truncates it to integers. It clips all values to `[0, 255]`.

> **Example**
> ```pycon
>>> import torch
>>> img = torch.tensor([[[0.5]], [[1.0]], [[2.0]]])
>>> torch_img_to_numpy(img).tolist()
[[[127, 255, 255]]]
>>> torch_img_to_numpy(img, reverse_colors=True).tolist()
[[[255, 255, 127]]]
```

Parameters

 * `img` (`Tensor`): An image of shape `[C, H, W]`. A floating point image has values in `[0, 1]`.
 * `reverse_colors` (`bool`): Whether to swap the first and the third channel, for example from RGB to BGR. The image must then have three channels.

Returns

 * `npt.NDArray[np.uint8]`: A new contiguous image of shape `[H, W, C]`.

## Attributes

### Color
