# efficient_keypoint_bbox_loss

Python API: `luxonis_train.attached_modules.losses.efficient_keypoint_bbox_loss`

The class, box, and keypoint loss of
[EfficientKeypointBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_keypoint_bbox_head.md).

## Classes

### EfficientKeypointBBoxLoss

Class, box, and keypoint loss of
[EfficientKeypointBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_keypoint_bbox_head.md).

 * `Inputs:`: * `features` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale
    * `class_scores` (`Tensor`): [B, N, nclasses] sigmoid scores
    * `distributions` (`Tensor`): [B, N, 4], `(l, t, r, b)` in stride units
    * `keypoints_raw` (`Tensor`): [B, N, 3**n*keypoints], raw `(x, y, v)` of each keypoint
    * `target_boundingbox` (`Tensor`): [M, 6], `[batch, class, x, y, w, h]`, `xywh` normalized
    * `target_keypoints` (`Tensor`): [M, 1 + 3**n*keypoints], `[batch, x, y, v, ...]`, `x` and `y` normalized
 * `Outputs:`: * `Tensor`: scalar total loss
    * `dict[str, Tensor]`: sub-losses `class`, `iou`, `regression`, and `visibility`, without the loss weights
 * `Formula:`: The class term Lcls and the box term Liou are the two terms of
   [AdaptiveDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md),
   each divided by max(S, 1). S is the sum of the assigned scores. After the warmup,
   [TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md)
   also reads the keypoints. It multiplies its IoU by the object keypoint similarity. The warmup
   [ATSSAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/atss_assigner.md)
   ignores the keypoints. The keypoint terms read only the set P of positive anchors. For anchor a and keypoint i, the predicted
   position (x̂a, i, ŷa, i) comes from
   [dist2kpts_noscale](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/efficient_keypoint_bbox_loss.md).
   The target position (xa, i, ya, i) is the keypoint of the assigned instance, divided by the stride of the anchor. Both
   positions are in grid units. va, i is `1` when the target visibility is above `0`, else `0`: Lkpt = (1)/(|P|)∑a ∈ P(∑iva, i(1 −
   exp( − da, i)))/(∑iva, i), da, i = ((x̂a, i − xa, i)2 + (ŷa, i − ya, i)2)/(2 (2σi)2 Aa) σi is the sigma of keypoint i, and Aa
   is the area of the assigned box in grid units times `area_factor`. The code adds `1e-9` to Aa and to ∑iva, i to prevent a
   division by zero. The visibility term Lvis is the mean binary cross entropy with logits between the raw visibility v̂a, i and
   va, i, over all keypoints of all positive anchors. `viz_pw` is the weight of its positive part. The total loss is: L = λclsLcls
   + λiouLiou + λkptLkpt + λvisLvis λcls, λiou, λkpt, and λvis are `class_loss_weight`, `iou_loss_weight`,
   `regr_kpts_loss_weight`, and `vis_kpts_loss_weight`.

> **References**
> * Source: Reimplemented from [YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications](https://arxiv.org/abs/2209.02976) and [PP-YOLOE: An evolved version of YOLO](https://arxiv.org/abs/2203.16250).
 * License: Apache-2.0 (this project)

> **Notes**
> The loss extends [AdaptiveDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md). The first call builds the anchors, as in that class. A batch without positive anchors, such as a batch without instances, makes both keypoint terms and the total loss `nan`. For a batch without instances, the loss also logs a debug message.

> **Example**
> Attached to a `EfficientKeypointBBoxHead` in `model.nodes`:

```yaml
- name: EfficientKeypointBBoxHead
  inputs: [RepPANNeck]
  losses:
    - name: EfficientKeypointBBoxLoss
```

 * `Compatible with:`: * Used by:
   [KeypointDetectionModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/keypoint_detection/v1/model.md)
    * Nodes:
      [EfficientKeypointBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_keypoint_bbox_head.md)

#### Methods

##### init

```python
def __init__(n_warmup_epochs: int = 0, iou_type: IoUType = 'giou', reduction: Literal['sum', 'mean'] = 'mean', class_loss_weight: float = 0.5, iou_loss_weight: float = 7.5, viz_pw: float = 1.0, regr_kpts_loss_weight: float = 12, vis_kpts_loss_weight: float = 1.0, sigmas: list[float] | None = None, area_factor: float | None = None, **kwargs):
```

Initialize the box loss and the keypoint parameters.

[AdaptiveDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md)
builds the assigners and
[VarifocalLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md).
The loss reads the strides, the grid cell size, and the grid cell offset from the node. It also reads the number of classes, the
number of keypoints, and the input image size. The loss therefore needs a node: without `node`,
[BaseAttachedModule.node](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/base_attached_module.md)
raises `RuntimeError`. The code of the box term adapts
[PPYOLOE_pytorch](https://github.com/Nioolek/PPYOLOE_pytorch/blob/master/ppyoloe/models).

Parameters

 * `n_warmup_epochs` (`int`): The number of epochs, counted from epoch `0`, that use
   [ATSSAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/atss_assigner.md).
   The later epochs use
   [TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md),
   which also reads the keypoints.
 * `iou_type` (`IoUType`): The IoU variant of the box term: `"none"` for the plain IoU, `"giou"`, `"diou"`, `"ciou"`, or `"siou"`.
 * `reduction` (`Literal['sum', 'mean']`): The method stores it as `self.reduction`. The loss does not read it.
 * `class_loss_weight` (`float`): The factor λcls of the class term.
 * `iou_loss_weight` (`float`): The factor λiou of the box term.
 * `viz_pw` (`float`): The weight of the positive part of the binary cross entropy in the visibility term.
 * `regr_kpts_loss_weight` (`float`): The factor λkpt of the keypoint regression term.
 * `vis_kpts_loss_weight` (`float`): The factor λvis of the visibility term.
 * `sigmas` (`list[float] | None`): One sigma for each keypoint. `None` selects the COCO person sigmas for `17` keypoints, and
   `0.04` for each keypoint otherwise.
   [get_sigmas](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/keypoints.md)
   logs the choice. A list with a length other than the number of keypoints raises `ValueError`.
 * `area_factor` (`float | None`): The factor that scales the area of a box to the pose area, in the keypoint regression term and
   in
   [TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md).
   `None` uses `0.53` and logs an info message.
 * `**kwargs`: Keyword arguments forwarded to
   [AdaptiveDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md),
   such as `per_class_weights`, `skip_stal`, `final_loss_weight`, and `node`.

##### dist2kpts_noscale

```python
def dist2kpts_noscale(anchor_points: Tensor, kpts: Tensor) -> Tensor:
```

Decode the raw keypoint offsets around the anchor points.

For the raw values (tx, ty) and the anchor point (ax, ay), the decoded position is (2tx + ax − 0.5, 2ty + ay − 0.5). The method
keeps the third value of each keypoint as it is. Unlike the decoding of
[EfficientKeypointBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_keypoint_bbox_head.md),
it does not multiply the position by the stride. The method does not change `kpts`.

Parameters

 * `anchor_points` (`Tensor`): The anchor points of shape `[N, 2]`, as `(x, y)`. The loss passes them in grid units.
 * `kpts` (`Tensor`): The raw keypoint values of shape `[B, N, n_keypoints, 3]`.

Returns

 * `Tensor`: The decoded keypoints, of the shape of `kpts`, in the units of `anchor_points`.

##### forward

```python
def forward(features: list[Tensor], class_scores: Tensor, distributions: Tensor, keypoints_raw: Tensor, target_boundingbox: Tensor, target_keypoints: Tensor) -> tuple[Tensor, dict[str, Tensor]]:
```

Assign the anchors to the instances and compute the loss.

The method inserts the class of each box into its keypoints. It decodes `distributions` to `xyxy` boxes and `keypoints_raw` to
keypoints, both in grid units around the anchor points. It converts the targets to padded tensors for each image, in pixels of the
input image. The assigner of the current epoch then matches the anchors to the target boxes.
[TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md)
also compares the keypoints in pixels. The method then computes the four terms of the class formula. For a batch without
instances, the method logs a debug message.

Parameters

 * `features` (`list[Tensor]`): The feature maps of the head, one of shape `[B, C_i, H_i, W_i]` for each scale. Only the first
   call uses them, to build the anchors.
 * `class_scores` (`Tensor`): Sigmoid class scores of shape `[B, N, n_classes]`, for the `N` anchors of all scales.
 * `distributions` (`Tensor`): The distances from each anchor point to the left, top, right, and bottom side of its box, of shape
   `[B, N, 4]`, in stride units.
 * `keypoints_raw` (`Tensor`): The raw keypoint values of shape `[B, N, 3 * n_keypoints]`. Each keypoint has the raw `x` and `y`
   offsets and a visibility logit.
 * `target_boundingbox` (`Tensor`): The `boundingbox` label of shape `[M, 6]`. Each row holds the batch index, the class, and the
   normalized `x`, `y`, `w`, and `h` of one box. `x` and `y` give the top-left corner.
 * `target_keypoints` (`Tensor`): The `keypoints` label of shape `[M, 1 + 3 * n_keypoints]`, with the rows in the order of
   `target_boundingbox`. Each row holds the batch index, then the normalized `x` and `y` and the visibility of each keypoint.

Returns

 * `tuple[Tensor, dict[str, Tensor]]`: The total loss as a scalar, and the sub-losses `"class"`, `"iou"`, `"regression"`, and
   `"visibility"`. The sub-losses are the detached terms without their factors. When no anchor is positive, the total loss and the
   two keypoint sub-losses are `nan`.

#### Attributes

##### gt_kpts_scale

##### node

##### supported_tasks
