# spatial_transforms

Python API: `luxonis_train.utils.spatial_transforms`

Maps boxes, keypoints, and masks from the resized image that the model sees back to the original image.

## Functions

### compute_ratio_and_padding

```python
def compute_ratio_and_padding(orig_h: int, orig_w: int, train_size: tuple[int, int], keep_aspect_ratio: bool) -> tuple[float | None, float, float]:
```

Compute the scale and the padding of a letterbox resize.

With `keep_aspect_ratio`, the ratio is r = min(ht ⁄ ho, wt ⁄ wo), where (ht, wt) is `train_size` and (ho, wo) is the original
size. The padding is px = (wt − r**wo) ⁄ 2 on the left and the right, and py = (ht − r**ho) ⁄ 2 on the top and the bottom. Without
`keep_aspect_ratio`, the loader stretches the image, so the function returns no ratio and no padding.

> **Example**
> ```pycon
>>> from luxonis_train.utils.spatial_transforms import (
...     compute_ratio_and_padding,
... )
>>> compute_ratio_and_padding(100, 200, (200, 200), True)
(1.0, 0.0, 50.0)
>>> compute_ratio_and_padding(100, 200, (200, 200), False)
(None, 0, 0)
```

Parameters

 * `orig_h` (`int`): The height of the original image, in pixels.
 * `orig_w` (`int`): The width of the original image, in pixels.
 * `train_size` (`tuple[int, int]`): The height and the width of the model input.
 * `keep_aspect_ratio` (`bool`): Whether the loader letterboxed the image.

Returns

 * `tuple[float | None, float, float]`: The ratio r, the padding px, and the padding py. Without `keep_aspect_ratio`, the result is `(None, 0, 0)`.

### transform_boxes

```python
def transform_boxes(raw_boxes: np.ndarray, orig_h: int, orig_w: int, train_size: tuple[int, int], keep_aspect_ratio: bool) ->
np.ndarray:
```

Convert boxes from the model input to the original image.

With `keep_aspect_ratio`, the function subtracts the padding and divides by the ratio from [compute_ratio_and_padding](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/spatial_transforms.md). Then it divides the coordinates by the original width and height. It does not clip the result, so a box in the padding gets values outside `[0, 1]`.

Warning: Without `keep_aspect_ratio`, the function does not scale the boxes from `train_size`. The result is correct only when the original size equals `train_size`.

> **Example**
> ```pycon
>>> import numpy as np
>>> from luxonis_train.utils import transform_boxes
>>> boxes = np.array([[20.0, 60.0, 120.0, 110.0]])
>>> transform_boxes(boxes, 100, 200, (200, 200), True).tolist()
[[0.1, 0.1, 0.5, 0.5]]
```

Parameters

 * `raw_boxes` (`np.ndarray`): The boxes in the pixels of the model input, of shape `[N, 4]`, in `xyxy` format.
 * `orig_h` (`int`): The height of the original image, in pixels.
 * `orig_w` (`int`): The width of the original image, in pixels.
 * `train_size` (`tuple[int, int]`): The height and the width of the model input.
 * `keep_aspect_ratio` (`bool`): Whether the loader letterboxed the image.

Returns

 * `np.ndarray`: The boxes of shape `[N, 4]` in normalized `xywh` format, where `x` and `y` give the top-left corner. An empty
   `raw_boxes` gives an empty array of shape `[0]`.

### transform_keypoints

```python
def transform_keypoints(raw_kpts: np.ndarray, orig_h: int, orig_w: int, train_size: tuple[int, int], keep_aspect_ratio: bool) -> np.ndarray:
```

Convert keypoints from the model input to the original image.

With `keep_aspect_ratio`, the function subtracts the padding and divides by the ratio from
[compute_ratio_and_padding](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/spatial_transforms.md).
Then it divides `x` by the original width and `y` by the original height. The third value does not change. The function does not
clip the result.

Warning: Without `keep_aspect_ratio`, the function does not scale the keypoints from `train_size`. The result is correct only when
the original size equals `train_size`.

> **Example**
> ```pycon
>>> import numpy as np
>>> from luxonis_train.utils import transform_keypoints
>>> kpts = np.array([[[100.0, 100.0, 2.0]]])
>>> transform_keypoints(kpts, 100, 200, (200, 200), True).tolist()
[[[0.5, 0.5, 2.0]]]
```

Parameters

 * `raw_kpts` (`np.ndarray`): The keypoints in the pixels of the model input, of shape `[N, K, 3]`. The last axis holds `x`, `y`, and a visibility or score value.
 * `orig_h` (`int`): The height of the original image, in pixels.
 * `orig_w` (`int`): The width of the original image, in pixels.
 * `train_size` (`tuple[int, int]`): The height and the width of the model input.
 * `keep_aspect_ratio` (`bool`): Whether the loader letterboxed the image.

Returns

 * `np.ndarray`: The keypoints of shape `[N, K, 3]`, as `float64`, with normalized `x` and `y`.

### transform_masks

```python
def transform_masks(raw_masks: np.ndarray, orig_h: int, orig_w: int, train_size: tuple[int, int], keep_aspect_ratio: bool) ->
np.ndarray:
```

Convert masks from the model input to the original image.

With `keep_aspect_ratio`, the function crops the padding from each mask. It truncates the crop bounds to integers. Then it resizes each mask to the original size with nearest-neighbour interpolation.

> **Example**
> ```pycon
>>> import numpy as np
>>> from luxonis_train.utils import transform_masks
>>> masks = np.zeros((1, 200, 200), dtype=np.float32)
>>> masks[:, 50:150] = 1.0
>>> result = transform_masks(masks, 100, 200, (200, 200), True)
>>> result.shape, float(result.min())
((1, 100, 200), 1.0)
```

Parameters

 * `raw_masks` (`np.ndarray`): The masks at the size of the model input, of shape `[N, H, W]`. `N` must be at least `1`, because
   `np.stack` rejects an empty list. The data type must be one that `cv2.resize` accepts, such as `float32`.
 * `orig_h` (`int`): The height of the original image, in pixels.
 * `orig_w` (`int`): The width of the original image, in pixels.
 * `train_size` (`tuple[int, int]`): The height and the width of the model input.
 * `keep_aspect_ratio` (`bool`): Whether the loader letterboxed the image.

Returns

 * `np.ndarray`: The masks of shape `[N, orig_h, orig_w]`.
