# precision_bbox_head

Python API: `luxonis_train.nodes.heads.precision_bbox_head`

The YOLOv8 detection head, which regresses a distribution over distance bins instead of a single distance.

## Classes

### PrecisionBBoxHead

Precision bounding box detection head.

 * `Inputs:`: * `inputs` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale, the last `n_heads` outputs of the input node
 * `Outputs:`: * train: * `features` (`list[Tensor]`): [B, 4⋅regmax + nclasses, Hi, Wi] per scale, logits
    * eval: * `features` (`list[Tensor]`): [B, 4⋅regmax + nclasses, Hi, Wi] per scale, logits
       * `boundingbox` (`list[Tensor]`): [Mi, 6] per image, `[x1, y1, x2, y2, score, class]`, pixels
       * `detections_pre_nms` (`Tensor`): [B, N, 5 + nclasses], only when requested
    * export: * `boundingbox` (`list[Tensor]`): [B, 5 + nclasses, Hi, Wi] per scale, DFL-decoded

> **References**
> * Source: Reimplemented from [Real-Time Flying Object Detection with YOLOv8](https://arxiv.org/abs/2305.09972) and [YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications](https://arxiv.org/abs/2209.02976).
 * License: Apache-2.0 (this project)

> **Notes**
> The head runs one [PreciseDecoupledBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md) for each scale. The regression branch predicts a distribution over `reg_max` distance bins for each side of a box. [DFL](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md) converts each distribution to the expected distance, in stride units. In evaluation mode, the head decodes the distances around the anchor points into boxes and runs NMS. Training mode skips this step, and it has priority over export mode. Export mode gives the distances and the sigmoid class scores of each scale, without NMS.

 * `Variants:`: None. Configure the node through `params`.

> **Example**
> A node entry in the `model.nodes` section of a config:

```yaml
- name: PrecisionBBoxHead
  inputs: [RepPANNeck]
```

 * `Compatible with:`: * Required labels: `boundingbox`
    * Losses:
      [PrecisionDFLDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
    * Metrics: *
      [ConfusionMatrix](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/confusion_matrix/confusion_matrix.md)
       * [MeanAveragePrecision](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/mean_average_precision/mean_average_precision.md)
       * [PrecisionRecallCurve](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/precision_recall_curve.md)
    * Visualizers:
      [BBoxVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/bbox_visualizer.md)
    * Export parser: `YOLO`

#### Methods

##### init

```python
def __init__(n_heads: Literal[2, 3, 4] = 3, conf_thres: float = 0.25, iou_thres: float = 0.45, max_det: int = 300, reg_max: int = 16, **kwargs):
```

Initialize one decoupled block for each scale.

The hidden width of every regression branch is the largest of `16`, `in_channels[0] // 4`, and `4 * reg_max`. The hidden width of
every class branch is the larger of `in_channels[0]` and `min(n_classes, 100)`. `in_channels[0]` belongs to the first scale. The
constructor also sets `grid_cell_offset` to `0.5` and `grid_cell_size` to `5.0`.
[PrecisionDFLDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/precision_dfl_detection_loss.md)
reads these two values from the node.

Parameters

 * `n_heads` (`Literal[2, 3, 4]`): Number of scales. The head reads the last `n_heads` outputs of the input node. An
   `attach_index` param replaces this selection. When the head gets fewer outputs, it logs a warning and uses that number. When an
   `attach_index` selects more outputs, the head builds one block for each output but keeps only `n_heads` strides. The
   construction then fails, because
   [initialize_weights](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/precision_bbox_head.md)
   raises `ValueError`.
 * `conf_thres` (`float`): NMS keeps only the boxes whose maximum class score is above this value. The value must be in `[0, 1]`.
   Otherwise, NMS raises `ValueError` in evaluation mode.
 * `iou_thres` (`float`): NMS removes a box when its IoU with a box of the same class and a higher score is above this value. The
   value must be in `[0, 1]`. Otherwise, NMS raises `ValueError` in evaluation mode.
 * `max_det` (`int`): Maximum number of boxes that NMS keeps for each image.
 * `reg_max` (`int`): Number of distance bins for each side of a box. The regression branch gives `4 * reg_max` channels. With a
   value above `1`,
   [DFL](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
   converts the bins to a distance. With `1`, the head uses the regression output as the distance directly.
 * `**kwargs`: Keyword arguments for
   [BaseNode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md).
   They must hold `original_in_shape`, the input sizes through `input_shapes` or `in_sizes`, and the class count through
   `n_classes` or `dataset_metadata`.

##### forward

```python
def forward(inputs: list[Tensor]) -> Packet[Tensor]:
```

Run the block of each scale and build the packet of the mode.

[forward_heads](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/precision_bbox_head.md)
gives the features, the class logits, and the distance bin logits of each scale. Training mode has priority over export mode. The
packet depends on the mode:

 * Training mode: `"features"` holds the distance bin logits and the class logits of shape `[B, 4 * reg_max + n_classes, H_i,
   W_i]` for each scale.
 * Export mode: `"boundingbox"` holds one map of shape `[B, 5 + n_classes, H_i, W_i]` for each scale. Its channels are the
   distances `(l, t, r, b)` in stride units, the maximum class score, and the class scores.
   [DFL](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
   decodes the distances when `reg_max` is above `1`. The scores are sigmoid probabilities.
 * Evaluation mode: `"features"` as in training mode, and the NMS result `"boundingbox"`. The head decodes the distances around
   the anchor points into boxes in pixels and does not clip them to the image. Each image gets a tensor of shape `[M_i, 6]` with
   the rows `[x1, y1, x2, y2, score, class]`. An image without boxes gets a tensor of shape `[0, 5 + n_classes]`. After a call to
   [BaseDetectionHead.request_detections_pre_nms](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/base_detection_head.md),
   the packet also holds `"detections_pre_nms"`. This is the NMS input of shape `[B, N, 5 + n_classes]`: the `xyxy` box in pixels,
   a constant `1`, and the class scores. `N` is the sum of `H_i * W_i`.

> **Example**
> A new head is in training mode:

```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import PrecisionBBoxHead
>>> sizes = [Size([1, 8, 32, 32]), Size([1, 16, 16, 16])]
>>> head = PrecisionBBoxHead(
...     n_heads=2,
...     n_classes=3,
...     input_shapes=[{"features": sizes}],
...     original_in_shape=Size([3, 256, 256]),
... )
>>> inputs = [torch.zeros(size) for size in sizes]
>>> [f.shape for f in head(inputs)["features"]]
[torch.Size([1, 67, 32, 32]), torch.Size([1, 67, 16, 16])]
```

For zero inputs, a new head gives each class a score below `conf_thres`. In evaluation mode, NMS thus keeps no box:

```pycon
>>> out = head.eval()(inputs)
>>> sorted(out), out["boundingbox"][0].shape
(['boundingbox', 'features'], torch.Size([0, 8]))
```

Export mode gives one map for each scale:

```pycon
>>> head.export = True
>>> [b.shape for b in head(inputs)["boundingbox"]]
[torch.Size([1, 8, 32, 32]), torch.Size([1, 8, 16, 16])]
```

Parameters

 * `inputs` (`list[Tensor]`): One feature map for each scale, of shape `[B, C_i, H_i, W_i]`.

Returns

 * `Packet[Tensor]`: The packet of the current mode.

##### forward_heads

```python
def forward_heads(inputs: list[Tensor]) -> tuple[list[Tensor], list[Tensor], list[Tensor]]:
```

Run the decoupled block of each scale.

The method pairs the blocks with `inputs` in order. It applies no sigmoid, so all outputs are logits.
[PrecisionSegmentBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/precision_seg_bbox_head.md)
also calls the method.

> **Example**
> ```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import PrecisionBBoxHead
>>> sizes = [Size([1, 8, 32, 32]), Size([1, 16, 16, 16])]
>>> head = PrecisionBBoxHead(
...     n_heads=2,
...     n_classes=3,
...     reg_max=4,
...     input_shapes=[{"features": sizes}],
...     original_in_shape=Size([3, 256, 256]),
... )
>>> inputs = [torch.zeros(size) for size in sizes]
>>> features, classes, regressions = head.forward_heads(inputs)
>>> features[1].shape
torch.Size([1, 19, 16, 16])
>>> classes[1].shape[1], regressions[1].shape[1]
(3, 16)
```

Parameters

 * `inputs` (`list[Tensor]`): One feature map for each scale, of shape `[B, C_i, H_i, W_i]`. The list must have one map for each block. Otherwise, `zip` raises `ValueError`.

Returns

 * `tuple[list[Tensor], list[Tensor], list[Tensor]]`: Three lists with one tensor for each scale, in the order of `inputs`. * the features, which join the distance bin logits and the class logits along the channel axis, of shape `[B, 4 * reg_max + n_classes, H_i, W_i]`;
    * the class logits, of shape `[B, n_classes, H_i, W_i]`;
    * the distance bin logits, of shape `[B, 4 * reg_max, H_i, W_i]`.

##### get_custom_head_config

```python
def get_custom_head_config(self) -> Params:
```

Return the NMS settings and the strides for the NN Archive.

A subclass adds its own keys to this dictionary, for example `"subtype"`.

Returns

 * `Params`: A dictionary with the keys `"iou_threshold"`, `"conf_threshold"`, `"max_det"`, and `"strides"`. They hold `iou_thres`, `conf_thres`, `max_det`, and `stride` as a list with one integer for each scale.

##### initialize_weights

```python
def initialize_weights(method: str | None = None):
```

Initialize the biases of the last layers of each scale.

The method first calls [BaseNode.initialize_weights](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md) with `method`. Then, for each scale with the stride s, it sets the bias of the last convolution of both branches. Each value of the regression bias becomes `1`. Each value of the class bias becomes:

b = log(5)/(nclasses(H ⁄ s)2)

H is the height in `original_in_shape`. The method does not change the weights. For a zero input, a new head thus gives each class the sigmoid score σ(b). This score is close to 5 ⁄ (nclasses(H ⁄ s)2).

> **Example**
> The first scale has the stride `8`. For `H = 256` and three classes, the class score of a new head is close to 5 ⁄ (3⋅322):

```pycon
>>> from torch import Size
>>> from luxonis_train.nodes import PrecisionBBoxHead
>>> sizes = [Size([1, 8, 32, 32]), Size([1, 16, 16, 16])]
>>> head = PrecisionBBoxHead(
...     n_heads=2,
...     n_classes=3,
...     input_shapes=[{"features": sizes}],
...     original_in_shape=Size([3, 256, 256]),
... )
>>> block = head.heads[0]
>>> block.regression_branch[-1].bias.unique().tolist()
[1.0]
>>> score = block.classification_branch[-1].bias.sigmoid()
>>> round(score[0].item(), 5), round(5 / (3 * 32**2), 5)
(0.00162, 0.00163)
```

Parameters

 * `method` (`str | None`): The method for [BaseNode.initialize_weights](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md). `"yolo"` changes the batch norm and activation settings. Other values skip that step.

#### Attributes

##### dfl

##### export_output_names

The names of the `n_heads` outputs of the exported model.

The default names are `output1_yolov8`, `output2_yolov8`, and so on. The `export_output_names` param replaces them only when it holds exactly `n_heads` names. The head logs a warning each time it gives the default names. The value is never `None`.

##### grid_cell_offset

##### grid_cell_size

##### heads

##### no

##### reg_max
