# precision_dfl_detection_loss

Python API: `luxonis_train.attached_modules.losses.precision_dfl_detection_loss`

The YOLOv8 detection loss: classification, box regression, and distribution focal loss over the regression bins.

## Classes

### BBoxLoss

CIoU and distribution focal loss terms of the positive anchors.

[PrecisionDFLDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
and
[PrecisionDFLSegmentationLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_segmentation_loss.md)
use it for their `iou` and `dfl` terms. P is the set of positive anchors. Anchor a has the predicted box ba, the target box b̂a,
and the weight wa, the sum of its target class scores. S is the normalizer that
[forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
receives as `total_score`:

Liou = (1)/(S)∑a ∈ Pwa(1 − CIoU(ba, b̂a))

Ldfl = (1)/(S)∑a ∈ Pwa DFLa

DFLa is the mean
[DFLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
over the four distances from the anchor point to the sides of the target box.

#### Methods

##### init

```python
def __init__(reg_max: int = 16):
```

Initialize the loss and its DFL part.

Parameters

 * `reg_max` (`int`): Number of distance bins for each side of a box. When `reg_max` is `1` or less, the loss has no DFL part.

##### forward

```python
def forward(pred_dist: Tensor, pred_bboxes: Tensor, anchors: Tensor, targets: Tensor, scores: Tensor, total_score: Tensor, fg_mask: Tensor) -> tuple[Tensor, Tensor]:
```

Compute the CIoU term and the DFL term.

`pred_bboxes`, `anchors`, and `targets` must use one unit, the size of one distance bin. The DFL term clips each target distance
to `[0, reg_max - 1.01]`. The CIoU term does not clip the boxes.

> **Example**
> The predicted box is equal to the target box, so the CIoU term is `0`. Uniform bin logits give the DFL term ln4:

```pycon
>>> import torch
>>> loss = BBoxLoss(reg_max=4)
>>> anchors = torch.tensor([[2.5, 2.5]])
>>> boxes = torch.tensor([[[1.5, 1.5, 3.5, 3.5]]])
>>> pred_dist = torch.zeros(1, 1, 16)
>>> scores, total = torch.ones(1, 1, 1), torch.tensor(1.0)
>>> fg_mask = torch.tensor([[True]])
>>> iou, dfl = loss(
...     pred_dist, boxes, anchors, boxes, scores, total, fg_mask
... )
>>> round(iou.item(), 4), round(dfl.item(), 4)
(0.0, 1.3863)
```

Parameters

 * `pred_dist` (`Tensor`): Distance bin logits of shape `[B, N, 4 * reg_max]`.
 * `pred_bboxes` (`Tensor`): Predicted `xyxy` boxes of shape `[B, N, 4]`, decoded from `pred_dist`.
 * `anchors` (`Tensor`): Anchor centers `(x, y)` of shape `[N, 2]`.
 * `targets` (`Tensor`): Assigned `xyxy` target boxes of shape `[B, N, 4]`.
 * `scores` (`Tensor`): Assigned class scores of shape `[B, N, n_classes]`.
 * `total_score` (`Tensor`): The normalizer S. The detection losses pass the sum of `scores`, at least `1`, as a Python number.
 * `fg_mask` (`Tensor`): Boolean mask of the positive anchors, of shape `[B, N]`.

Returns

 * `tuple[Tensor, Tensor]`: The scalar CIoU term and the scalar DFL term. Without a DFL part, the DFL term is a zero tensor of
   shape `[1]`.

#### Attributes

##### dist_loss

### DFLoss

Distribution focal loss (DFL) over the distance bins of a box side.

A target distance y is a real number between the bins yl = ⌊y⌋ and yr = yl + 1. With the softmax probabilities p over the bins,
the loss is

DFL = − ((yr − y)logpyl + (y − yl)logpyr)

The loss is lowest when bin yl has the probability yr − y and bin yr has the probability y − yl.

#### Methods

##### init

```python
def __init__(reg_max: int = 16):
```

Initialize the loss.

Parameters

 * `reg_max` (`int`): Number of distance bins for each side of a box.

#### Attributes

##### reg_max

### PrecisionDFLDetectionLoss

Bounding box loss for
[PrecisionBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/precision_bbox_head.md),
with a distribution focal loss on the box sides.

The loss decodes a box for each anchor from the distance bins of the head.
[TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md)
then selects the positive anchors and gives each of them a target box and target class scores. The target class scores of the
other anchors are `0`. The loss has three terms:

 * a binary cross-entropy on the class logits of all anchors;
 * a CIoU loss on the boxes of the positive anchors;
 * a distribution focal loss (DFL) on the distance bins of the positive anchors.

[BBoxLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
computes the last two terms.

 * `Inputs:`: * `features` (`list[Tensor]`): [B, 4**reg*max + nclasses, Hi, Wi] per scale, the distance bin logits followed by the
   class logits
    * `target` (`Tensor`): [Ngt, 6], `[batch_index, class, x, y, w, h]`, `xywh` normalized, with `x` and `y` at the top-left
      corner
 * `Outputs:`: * `Tensor`: scalar total loss
    * `dict[str, Tensor]`: scalar sub-losses `class`, `iou`, `dfl`, detached and without the weights
 * `Formula:`: The sums run over the anchors of all scales in all images of the batch. For anchor a and class c, za, c is the
   class logit and t̂a, c the score from the assigner. P is the set of positive anchors. ba is the predicted box and b̂a the
   assigned box, both in units of the stride of the anchor. Each positive anchor has the weight wa = ∑ct̂a, c, and the normalizer
   is S = max(∑awa, 1): Lcls = (1)/(S)∑a∑cBCE(za, c, t̂a, c) Liou = (1)/(S)∑a ∈ Pwa(1 − CIoU(ba, b̂a)) Ldfl = (1)/(S)∑a ∈ Pwa DFLa
   L = λclsLcls + λboxLiou + λdflLdfl DFLa is the mean
   [DFLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
   of the four sides of the box. The weights λ are `class_loss_weight`, `bbox_loss_weight`, and `dfl_loss_weight`. Liou and Ldfl
   are `0` when no anchor is positive.

> **References**
> * Source: Reimplemented from [Real-Time Flying Object Detection with YOLOv8](https://arxiv.org/abs/2305.09972) and [YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications](https://arxiv.org/abs/2209.02976) and [PP-YOLOE: An evolved version of YOLO](https://arxiv.org/abs/2203.16250).
 * License: Apache-2.0 (this project)

> **Notes**
> The first call of `forward` caches the anchor points, their strides, and the image scale. It computes them from the shapes of `features`, and from `stride`, `grid_cell_offset`, and `original_in_shape` of the node. Later calls reuse the cache, so every batch must have the feature map sizes of the first batch. The assigner works in pixels of the input image. The box and DFL terms work in units of the stride. When `reg_max` of the node is `1`, the `dfl` term is always `0`. It then has the shape `[1]` when an anchor is positive, and so does the total loss.

> **Example**
> Attached to a `PrecisionBBoxHead` in `model.nodes`:

```yaml
- name: PrecisionBBoxHead
  inputs: [RepPANNeck]
  losses:
    - name: PrecisionDFLDetectionLoss
```

 * `Compatible with:`: * Nodes:
   [PrecisionBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/precision_bbox_head.md)

#### Methods

##### init

```python
def __init__(tal_topk: int = 10, class_loss_weight: float = 0.5, bbox_loss_weight: float = 7.5, dfl_loss_weight: float = 1.5, skip_stal: bool = False, **kwargs):
```

Initialize the assigner, the box loss, and the class loss.

The loss reads `n_classes`, `original_in_shape`, `stride`, `grid_cell_size`, `grid_cell_offset`, and `reg_max` from the node.
Without a `node`, the constructor raises `RuntimeError`. The code adapts
[PPYOLOE_pytorch](https://github.com/Nioolek/PPYOLOE_pytorch/blob/master/ppyoloe/models). For the best results, multiply each
weight by `trainer.accumulate_grad_batches`.

Parameters

 * `tal_topk` (`int`): The `topk` of
   [TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md),
   the largest number of positive anchors for each target box. The exponents of the assigner are fixed: `alpha` is `0.5` and
   `beta` is `6.0`.
 * `class_loss_weight` (`float`): Weight of the classification term.
 * `bbox_loss_weight` (`float`): Weight of the CIoU box term.
 * `dfl_loss_weight` (`float`): Weight of the DFL term.
 * `skip_stal` (`bool`): Whether to turn off Small-Target-Aware Label Assignment (STAL) in the assigner. When a side of a target
   box is shorter than the smallest stride, STAL gives that side the length of the second smallest stride. The assigner uses the
   enlarged box only to find the anchors inside the box.
 * `**kwargs`: Keyword arguments forwarded to
   [BaseLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/base_loss.md),
   such as `node` and `final_loss_weight`.

##### decode_bbox

```python
def decode_bbox(anchor_points: Tensor, pred_dist: Tensor) -> Tensor:
```

Decode the distance bin logits into boxes.

For each side of each box, the method applies a softmax over the `reg_max` bins. The distance of the side is the expected bin
index under these probabilities.
[dist2bbox](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/boundingbox.md)
then turns the four distances into a box around the anchor point. The method also decodes the bins when `reg_max` is `1`, and then
every distance is `0`.

Parameters

 * `anchor_points` (`Tensor`): Anchor centers `(x, y)` of shape `[N, 2]`.
   [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
   passes them in units of the stride of each anchor.
 * `pred_dist` (`Tensor`): Distance bin logits of shape `[B, N, 4 * reg_max]`, with the sides in the order left, top, right,
   bottom.

Returns

 * `Tensor`: Boxes of shape `[B, N, 4]` in `xyxy` format, in the units of `anchor_points`.

##### forward

```python
def forward(features: list[Tensor], target: Tensor) -> tuple[Tensor, dict[str, Tensor]]:
```

Compute the detection loss of one batch.

The method splits `features` into the distance bin logits and the class logits of all anchors. It pads `target` to the same number
of boxes for each image, and converts the boxes to `xyxy` pixels.
[decode_bbox](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
gives the predicted boxes.
[TaskAlignedAssigner](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/assigners/tal_assigner.md)
reads detached copies of the class probabilities and of the predicted boxes in pixels. A target box with a sum of its `xyxy`
coordinates of `0` or less counts as padding. The first call caches the anchor points and the image scale, as the class notes
describe.

Parameters

 * `features` (`list[Tensor]`): One tensor per scale, of shape `[B, 4 * reg_max + n_classes, H_i, W_i]`. The `features` output of
   the node.
 * `target` (`Tensor`): Target boxes of shape `[N_gt, 6]`, with rows `[batch_index, class, x, y, w, h]`. The coordinates are
   `xywh` normalized to `[0, 1]`, with `x` and `y` at the top-left corner. The `boundingbox` label of the task.

Returns

 * `tuple[Tensor, dict[str, Tensor]]`: The scalar weighted total loss, and a dictionary that maps `"class"`, `"iou"`, and `"dfl"`
   to the detached terms before the weights. The total loss and `"dfl"` have the shape `[1]` when `reg_max` of the node is `1` and
   an anchor is positive.

#### Attributes

##### assigner

##### bbox_loss

##### bce

##### gt_bboxes_scale

##### node

##### supported_tasks
