# object_keypoint_similarity

Python API: `luxonis_train.attached_modules.metrics.object_keypoint_similarity`

Object keypoint similarity, the keypoint counterpart of IoU.

The metric pairs the predicted poses of each image with the target poses and averages the similarity of the pairs. The sigma of a
keypoint sets the distance that the keypoint tolerates.

## Classes

### ObjectKeypointSimilarity

Mean object keypoint similarity of the paired poses.

 * `Inputs:`: * `keypoints` (`list[Tensor]`): [Mi, nkeypoints, 3] for each image, `(x, y, conf)` in pixels.
   [FOMOHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/fomo_head.md)
   gives [Mi, 1, 4]. Only `x` and `y` count.
    * `target_boundingbox` (`Tensor`): [N, 6], `[batch, class, x, y, w, h]`, normalized
    * `target_keypoints` (`Tensor | None`): [N, 1 + 3nkeypoints], `[batch, x, y, v, ...]`, normalized. `Tasks.FOMO` does not read
      it.
 * `Outputs:`: * `Tensor`: scalar in `[0, 1]`
 * `Formula:`:
   [compute_pose_oks](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/keypoints.md)
   gives the similarity of each pair of a target pose and a predicted pose of one image. The pose area of a target is
   `area_factor` times the area of its box in pixels. The Hungarian algorithm pairs the poses so that the sum of the similarities
   is the largest. The score of an image is the mean similarity of its pairs. The metric returns the mean score of the images.

> **References**
> * Source: This project.
 * License: Apache-2.0 (this project)

> **Notes**
> * The metric skips an image without targets.
 * A pose without a pair does not lower the score of its image. An image with targets and no predictions scores `0`.
 * For `Tasks.FOMO`, the centers of the target boxes are the target keypoints.

> **Example**
> Attached to a `FOMOHead` in `model.nodes`:

```yaml
- name: FOMOHead
  inputs: [EfficientRep]
  metrics:
    - name: ObjectKeypointSimilarity
```

 * `Compatible with:`: * Used by:
   [KeypointDetectionModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/keypoint_detection/v1/model.md)
    * Nodes: *
      [EfficientKeypointBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_keypoint_bbox_head.md)
       * [FOMOHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/fomo_head.md)

#### Methods

##### init

```python
def __init__(sigmas: list[float] | None = None, area_factor: float | None = None, use_cocoeval_oks: bool = True, **kwargs):
```

Initialize the metric and resolve the sigmas.

The constructor reads
[BaseAttachedModule.n_keypoints](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/base_attached_module.md),
so the module needs a node.
[get_sigmas](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/keypoints.md)
raises `ValueError` when `sigmas` does not have one value for each keypoint.

Parameters

 * `sigmas` (`list[float] | None`): One sigma for each keypoint. A larger sigma tolerates a larger distance. `None` selects the
   COCO person sigmas for `17` keypoints, and `0.04` for each keypoint otherwise.
   [get_sigmas](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/keypoints.md)
   then logs the selection.
 * `area_factor` (`float | None`): The factor that scales the area of a target box to the pose area. `None` selects `0.53` and
   logs an info message.
 * `use_cocoeval_oks` (`bool`): When `True`, use the formula of the COCO evaluation code. When `False`, use the formula of the
   COCO keypoint definition.
   [compute_pose_oks](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/keypoints.md)
   shows both formulas.
 * `**kwargs`: Keyword arguments forwarded to
   [BaseMetric](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/base_metric.md),
   such as `node`.

##### compute

```python
def compute(self) -> Tensor:
```

Pair the stored poses of each image and average the scores.

The class docstring describes the pairing and the score. The method also moves `sigmas` to the device of the metric.

> **Example**
> The first image has two targets and one exact prediction. The target without a pair does not count, so the image scores `1`. The second image has no prediction and scores `0`:

```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import EfficientKeypointBBoxHead
>>> head = EfficientKeypointBBoxHead(
...     n_heads=1,
...     n_classes=1,
...     n_keypoints=1,
...     input_shapes=[{"features": [Size([1, 8, 8, 8])]}],
...     original_in_shape=Size([3, 100, 100]),
... )
>>> metric = ObjectKeypointSimilarity(
...     node=head, sigmas=[0.04], area_factor=0.53
... )
>>> boxes = torch.tensor([[0.0, 0.0, 0.1, 0.1, 0.2, 0.2]] * 3)
>>> boxes[2, 0] = 1.0
>>> keypoints = torch.tensor(
...     [
...         [0.0, 0.2, 0.2, 2.0],
...         [0.0, 0.6, 0.6, 2.0],
...         [1.0, 0.2, 0.2, 2.0],
...     ]
... )
>>> exact = torch.tensor([[[20.0, 20.0, 1.0]]])
>>> metric.update([exact, torch.zeros(0, 1, 3)], boxes, keypoints)
>>> metric.compute().item()
0.5
```

Returns

 * `Tensor`: The mean score of the stored images, a scalar in `[0, 1]`. It is `0` when no image had targets.

##### update

```python
def update(keypoints: list[Tensor], target_boundingbox: Tensor, target_keypoints: Tensor | None = None):
```

Scale the targets of one batch to pixels and store them.

For each image with targets, the method stores the predicted keypoints, the target keypoints in pixels, and the pose area of each
target. The height and the width of
[BaseAttachedModule.original_in_shape](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/base_attached_module.md)
scale the normalized values.

Parameters

 * `keypoints` (`list[Tensor]`): The predicted keypoints of each image, of shape `[M_i, n_keypoints, 3]`, in pixels.
   [FOMOHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/fomo_head.md)
   gives the shape `[M_i, 1, 4]`. Only `x` and `y` count.
 * `target_boundingbox` (`Tensor`): The target boxes of the batch, of shape `[N, 6]`, as `[batch, class, x, y, w, h]` with
   normalized values.
 * `target_keypoints` (`Tensor | None`): The target keypoints of the batch, of shape `[N, 1 + 3 * n_keypoints]`, as `[batch, x, y,
   v, ...]` with normalized coordinates. For `Tasks.FOMO`, the method uses the centers of the target boxes instead.

Raises

 * `ValueError`: When `target_keypoints` is `None` and the task is not `Tasks.FOMO`.

#### Attributes

##### pred_keypoints

##### scales

##### supported_tasks

##### target_keypoints
