# adaptive_detection_loss

Python API: `luxonis_train.attached_modules.losses.adaptive_detection_loss`

The YOLOv6 detection loss, and the varifocal loss it uses for classification.

## Classes

### AdaptiveDetectionLoss

Bounding box loss of YOLOv6, with a warmup of the anchor assignment.

 * `Inputs:`: * `features` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale
    * `class_scores` (`Tensor`): [B, N, nclasses] sigmoid scores
    * `distributions` (`Tensor`): [B, N, 4], `(l, t, r, b)` in stride units
    * `target` (`Tensor`): [M, 6], `[batch, class, x, y, w, h]`, `xywh` normalized
 * `Outputs:`: * `Tensor`: scalar total loss
    * `dict[str, Tensor]`: sub-losses `class` and `iou`, without the loss weights
 * `Formula:`: On each call, an assigner matches the N anchors to the target boxes.
   [ATSSAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/atss_assigner.md)
   runs while the epoch is lower than `n_warmup_epochs`, and
   [TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md)
   runs after that. For each anchor a, the assigner gives a class, a box, and a soft score qa, c for each class c. S is the sum of
   all qa, c: L = (λclsLcls + λiou∑a ∈ P(1 − IoUa)∑cqa, c)/(max(S, 1)) Lcls is the
   [VarifocalLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md),
   summed over all anchors and classes. P is the set of positive anchors, and IoUa is the `iou_type` overlap of the predicted and
   the assigned box of anchor a. λcls is `class_loss_weight` and λiou is `iou_loss_weight`. The IoU term is `0` when no anchor is
   positive.

> **References**
> * Source: Reimplemented from [YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications](https://arxiv.org/abs/2209.02976) and [PP-YOLOE: An evolved version of YOLO](https://arxiv.org/abs/2203.16250).
 * License: Apache-2.0 (this project)

> **Notes**
> The first call builds the anchors, the anchor points, and the strides from the shapes of `features`. The loss stores them as non-persistent buffers. All later calls use these buffers, even when the shapes of `features` change. The first call that uses [TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md) logs the change of the assigner, also when `n_warmup_epochs` is `0`.

> **Example**
> Attached to a `EfficientBBoxHead` in `model.nodes`:

```yaml
- name: EfficientBBoxHead
  inputs: [RepPANNeck]
  losses:
    - name: AdaptiveDetectionLoss
```

 * `Compatible with:`: * Used by:
   [DetectionModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/detection/v1/model.md)
    * Nodes:
      [EfficientBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md)

#### Methods

##### init

```python
def __init__(n_warmup_epochs: int = 0, iou_type: IoUType = 'giou', reduction: Literal['sum', 'mean'] = 'mean', class_loss_weight: float = 1.0, iou_loss_weight: float = 2.5, per_class_weights: list[float] | None = None, skip_stal: bool = False, **kwargs):
```

Initialize the loss, its two assigners, and
[VarifocalLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md).

The method reads the strides, the grid cell size, the grid cell offset, the number of classes, and the input image size from the
node. The loss therefore needs a node: without `node`,
[BaseAttachedModule.node](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/base_attached_module.md)
raises `RuntimeError`. The code adapts [PPYOLOE_pytorch](https://github.com/Nioolek/PPYOLOE_pytorch/blob/master/ppyoloe/models).

Parameters

 * `n_warmup_epochs` (`int`): The number of epochs, counted from epoch `0`, that use
   [ATSSAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/atss_assigner.md)
   with `topk=9`. The later epochs use
   [TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md)
   with `topk=13`, `alpha=1.0`, and `beta=6.0`.
 * `iou_type` (`IoUType`): The IoU variant of the box term: `"none"` for the plain IoU, `"giou"`, `"diou"`, `"ciou"`, or `"siou"`.
   [luxonis_train.utils.bbox_iou](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/boundingbox.md)
   describes the variants.
 * `reduction` (`Literal['sum', 'mean']`): The loss does not read it. Both terms always use the normalized sum of the class
   formula.
 * `class_loss_weight` (`float`): The factor λcls of the classification term.
 * `iou_loss_weight` (`float`): The factor λiou of the box term.
 * `per_class_weights` (`list[float] | None`): One factor for each class.
   [VarifocalLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md)
   multiplies the weight of each class by its factor. A list with a length other than the number of classes logs a warning, and
   the loss then uses no factors. `None` uses no factors.
 * `skip_stal` (`bool`): `True` turns off the Small-Target-Aware Label Assignment of
   [TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md).
 * `**kwargs`: Keyword arguments forwarded to
   [BaseLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/base_loss.md),
   such as `final_loss_weight` and `node`.

##### forward

```python
def forward(features: list[Tensor], class_scores: Tensor, distributions: Tensor, target: Tensor) -> tuple[Tensor, dict[str, Tensor]]:
```

Assign the anchors to the target boxes and compute the loss.

The method converts `target` to `xyxy` boxes in pixels of the input image. It pads the boxes of each image to the largest number
of boxes in one image. It decodes `distributions` to predicted boxes around the anchor points. The assigner of the current epoch
then matches the anchors to the target boxes, and the method computes the two terms of the class formula.

Parameters

 * `features` (`list[Tensor]`): The feature maps of the head, one of shape `[B, C_i, H_i, W_i]` for each scale. Only the first
   call uses them, to build the anchors.
 * `class_scores` (`Tensor`): Sigmoid class scores of shape `[B, N, n_classes]`, for the `N` anchors of all scales.
 * `distributions` (`Tensor`): The distances from each anchor point to the left, top, right, and bottom side of its box, of shape
   `[B, N, 4]`, in stride units.
 * `target` (`Tensor`): The `boundingbox` label of shape `[M, 6]`. Each row holds the batch index, the class, and the normalized
   `x`, `y`, `w`, and `h` of one box. `x` and `y` give the top-left corner.

Returns

 * `tuple[Tensor, dict[str, Tensor]]`: The total loss as a scalar, and the sub-losses `"class"` and `"iou"`. The sub-losses are
   the detached terms without the factors `class_loss_weight` and `iou_loss_weight`.

#### Attributes

##### anchor_points

##### anchor_points_strided

##### anchors

##### atss_assigner

##### gt_bboxes_scale

##### n_anchors_list

##### node

##### stride_tensor

##### supported_tasks

##### tal_assigner

##### varifocal_loss

### VarifocalLoss

Varifocal loss, which trains IoU-aware class scores.

The loss trains the predicted class score to match a soft target score, such as the IoU between the predicted box and the assigned
box. It weights the binary cross entropy between the predicted score p and the target score q of each anchor and class. With the
one-hot label y:

L = ∑w⋅BCE(p, q), w = α pγ (1 − y) + q y

A negative element gets the focal weight α**pγ, and a positive element gets the weight q. With `per_class_weights`, the loss also
multiplies the weight of each class by the factor of that class. The loss does not detach w, so the gradient also flows through
pγ.
[AdaptiveDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md)
uses this loss for its classification term.

> **References**
> * [VarifocalNet: An IoU-aware Dense Object Detector](https://arxiv.org/abs/2008.13367)
 * The implementation adapts the code of
   [PPYOLOE_pytorch](https://github.com/Nioolek/PPYOLOE_pytorch/blob/master/ppyoloe/models/losses.py).

> **Example**
> The example has one anchor with two classes. The first class is positive with the target score `0.8`. The binary cross entropy of both elements is ln2, and the weights are `0.8` and 0.75⋅0.52 = 0.1875:

```pycon
>>> import torch
>>> loss = VarifocalLoss()
>>> pred_score = torch.tensor([[[0.5, 0.5]]])
>>> target_score = torch.tensor([[[0.8, 0.0]]])
>>> label = torch.tensor([[[1.0, 0.0]]])
>>> round(loss(pred_score, target_score, label).item(), 4)
0.6845
```

#### Methods

##### init

```python
def __init__(alpha: float = 0.75, gamma: float = 2.0, per_class_weights: Tensor | None = None):
```

Initialize the varifocal loss.

Parameters

 * `alpha` (`float`): The factor α of the weight of a negative element.
 * `gamma` (`float`): The exponent γ of the predicted score in the weight of a negative element.
 * `per_class_weights` (`Tensor | None`): One factor for each class, of shape `[n_classes]`. `None` uses no factors.

##### forward

```python
def forward(pred_score: Tensor, target_score: Tensor, label: Tensor) -> Tensor:
```

Compute the varifocal loss, summed over all elements.

The binary cross entropy runs in `float32`, with autocast off. When `per_class_weights` and `pred_score` are on different devices,
the method replaces `per_class_weights` with a copy on the device of `pred_score`.

Parameters

 * `pred_score` (`Tensor`): Predicted class scores in `[0, 1]`, of shape `[B, N, n_classes]`.
 * `target_score` (`Tensor`): Target scores in `[0, 1]`, of the same shape. An assigner gives a soft score to the assigned class,
   and `0` to the other classes.
 * `label` (`Tensor`): One-hot labels of the same shape. A background anchor has only zeros.

Returns

 * `Tensor`: The loss summed over all elements, as a scalar.
