# efficient_bbox_head

Python API: `luxonis_train.nodes.heads.efficient_bbox_head`

The decoupled detection head of YOLOv6.

## Classes

### EfficientBBoxHead

Efficient object detection head.

 * `Inputs:`: * `inputs` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale
 * `Outputs:`: * train: * `features` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale
       * `class_scores` (`Tensor`): [B, N, nclasses]
       * `distributions` (`Tensor`): [B, N, 4], `(l, t, r, b)` in stride units
    * eval: * `features` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale
       * `class_scores` (`Tensor`): [B, N, nclasses]
       * `distributions` (`Tensor`): [B, N, 4], `(l, t, r, b)` in stride units
       * `boundingbox` (`list[Tensor]`): [Mi, 6] per image, `[x1, y1, x2, y2, conf, class]`, pixels
       * `detections_pre_nms` (`Tensor`): [B, N, 5 + nclasses], only when requested
    * export: * `boundingbox` (`list[Tensor]`): [B, 5 + nclasses, Hi, Wi] per scale, `(l, t, r, b)` then scores

> **References**
> * Source: Reimplemented from [YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications](https://arxiv.org/abs/2209.02976).
 * License: Apache-2.0 (this project)

> **Notes**
> The head runs one [EfficientDecoupledBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md) for each scale. In evaluation mode, it decodes the distances around the anchor points into boxes and runs NMS. Training mode and export mode skip this step.

 * `Variants:`: None. Configure the node through `params`.

> **Example**
> A node entry in the `model.nodes` section of a config:

```yaml
- name: EfficientBBoxHead
  inputs: [RepPANNeck]
```

 * `Compatible with:`: * Required labels: `boundingbox`
    * Used by:
      [DetectionModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/detection/v1/model.md)
    * Losses:
      [AdaptiveDetectionLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/adaptive_detection_loss.md)
    * Metrics: *
      [ConfusionMatrix](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/confusion_matrix/confusion_matrix.md)
       * [MeanAveragePrecision](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/mean_average_precision/mean_average_precision.md)
       * [PrecisionRecallCurve](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/precision_recall_curve.md)
    * Visualizers:
      [BBoxVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/bbox_visualizer.md)
    * Export parser: `YOLO`
    * Pretrained weights: available through `weights: download`

#### Methods

##### init

```python
def __init__(n_heads: Literal[2, 3, 4] = 3, conf_thres: float = 0.25, iou_thres: float = 0.45, max_det: int = 300, bias_init_p: float = 0.01, **kwargs):
```

Initialize one decoupled block for each scale.

Parameters

 * `n_heads` (`Literal[2, 3, 4]`): Number of scales. The head reads the last `n_heads` outputs of the input node. An
   `attach_index` param replaces this selection. The value is usually equal to the number of neck outputs. When the input node
   gives fewer outputs, the head logs a warning and uses that number. Defaults to `3`.
 * `conf_thres` (`float`): NMS keeps only the boxes whose maximum class score is above this value. The value must be in `[0, 1]`.
   Defaults to `0.25`.
 * `iou_thres` (`float`): NMS removes a box when its IoU with a box of the same class and a higher score is above this value. The
   value must be in `[0, 1]`. Defaults to `0.45`.
 * `max_det` (`int`): Maximum number of boxes that NMS keeps for each image. Defaults to `300`.
 * `bias_init_p` (`float`): Initial class score of every anchor point.
   [initialize_weights](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md)
   sets the class bias from it. Defaults to `1e-2`.
 * `**kwargs`: Keyword arguments for
   [BaseNode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md),
   such as `n_classes` and `input_shapes`.

##### forward

```python
def forward(inputs: list[Tensor]) -> Packet[Tensor]:
```

Run the block of each scale and build the packet of the mode.

Each
[EfficientDecoupledBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
returns the decoded features, the class logits, and the distances of one scale. The head applies a sigmoid to the class logits.
Export mode has priority over training mode. The packet depends on the mode:

 * Export mode: `"boundingbox"` holds one map of shape `[B, 5 + n_classes, H_i, W_i]` for each scale. Its channels are the
   distances `(l, t, r, b)`, the maximum class score, and the class scores.
 * Training mode: `"features"` holds the decoded feature map of shape `[B, C_i, H_i, W_i]` for each scale. `"class_scores"` of
   shape `[B, N, n_classes]` and `"distributions"` of shape `[B, N, 4]` hold the class scores and the distances of all `N` anchor
   points. The distances `(l, t, r, b)` are in stride units. `N` is the sum of `H_i * W_i`.
 * Evaluation mode: the training keys and `"boundingbox"`. The head decodes the distances into boxes in pixels and runs NMS. Each
   image gets a tensor of shape `[M_i, 6]` with the rows `[x1, y1, x2, y2, score, class]`. An image without boxes gets a tensor of
   shape `[0, 5 + n_classes]`. After a call to
   [BaseDetectionHead.request_detections_pre_nms](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/base_detection_head.md),
   the packet also holds `"detections_pre_nms"`. This is the NMS input of shape `[B, N, 5 + n_classes]`: the `xyxy` box in pixels,
   a constant `1`, and the class scores.

> **Example**
> ```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import EfficientBBoxHead
>>> sizes = [Size([1, 8, 32, 32]), Size([1, 16, 16, 16])]
>>> head = EfficientBBoxHead(
...     n_heads=2,
...     n_classes=3,
...     input_shapes=[{"features": sizes}],
...     original_in_shape=Size([3, 256, 256]),
... )
>>> out = head([torch.zeros(size) for size in sizes])
>>> out["class_scores"].shape, out["distributions"].shape
(torch.Size([1, 1280, 3]), torch.Size([1, 1280, 4]))
```

A new head gives each class the score `0.01`, which is below `conf_thres`. In evaluation mode, NMS thus keeps no box:

```pycon
>>> out = head.eval()([torch.zeros(size) for size in sizes])
>>> sorted(out)
['boundingbox', 'class_scores', 'distributions', 'features']
>>> out["boundingbox"][0].shape
torch.Size([0, 8])
```

Parameters

 * `inputs` (`list[Tensor]`): One feature map for each scale, of shape `[B, C_i, H_i, W_i]`.

Returns

 * `Packet[Tensor]`: The packet of the current mode.

##### get_custom_head_config

```python
def get_custom_head_config(self) -> Params:
```

Return the NMS settings and the strides for the NN Archive.

A subclass adds its own keys to this dictionary, for example `"subtype"`.

Returns

 * `Params`: A dictionary with the keys `"iou_threshold"`, `"conf_threshold"`, `"max_det"`, and `"strides"`. They hold `iou_thres`, `conf_thres`, `max_det`, and `stride` as a list with one integer for each scale.

##### get_weights_url

```python
def get_weights_url(self) -> str:
```

Select the COCO checkpoint from the input channels.

The head has no variants. The input channels of the `"n"`, `"s"`, and `"l"` variants of [RepPANNeck](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/necks/reppan_neck/reppan_neck.md) select the checkpoint of the same size. The checkpoints hold no class branch, so any `n_classes` can load them.

Raises

 * `NotImplementedError`: When the input channels are not `[32, 64, 128]`, `[64, 128, 256]`, or `[128, 256, 512]`.

##### initialize_weights

```python
def initialize_weights(method: str | None = None):
```

Initialize the weights and the priors of the last layers.

The method first calls [BaseNode.initialize_weights](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md). Then, for each scale, it sets the weights of the last convolution of both branches to zero. The bias of the class branch becomes − log((1 − p) ⁄ p), where p is `bias_init_p`. The bias of the regression branch becomes `1`. A new head thus predicts the class score `bias_init_p` and the distance `1` for every anchor point and every input.

> **Example**
> The node calls the method after construction, so a new head already gives these outputs:

```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import EfficientBBoxHead
>>> sizes = [Size([1, 8, 32, 32]), Size([1, 16, 16, 16])]
>>> head = EfficientBBoxHead(
...     n_heads=2,
...     n_classes=2,
...     input_shapes=[{"features": sizes}],
...     original_in_shape=Size([3, 256, 256]),
... )
>>> out = head([torch.zeros(size) for size in sizes])
>>> torch.allclose(out["class_scores"], torch.tensor(0.01))
True
>>> out["distributions"].unique().tolist()
[1.0]
```

Parameters

 * `method` (`str | None`): Method for [BaseNode.initialize_weights](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md). `"yolo"` changes the batch norm and activation settings. Other values skip that step. Defaults to `None`.

##### load_checkpoint

```python
def load_checkpoint(path: str | None = None, strict: bool = False):
```

Load a checkpoint, with a non-strict key match by default.

The method passes `path` as `ckpt` to [BaseNode.load_checkpoint](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md). The override renames the first parameter and sets the default of `strict` to `False`.

Warning: The override has no `ckpt` parameter. After construction, the node calls `load_checkpoint(ckpt=...)` for a `weights` URL. Thus, a call such as `EfficientBBoxHead(weights="https://...")` raises `TypeError`.

Parameters

 * `path` (`str | None`): Local path or URL of a `.ckpt` file. [LuxonisLightningModule](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/luxonis_lightning.md) also gives a state dictionary, and the base method loads it directly. `None` or `""` takes the URL from [get_weights_url](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md). That call fails for input channels without a checkpoint.
 * `strict` (`bool`): Whether the keys of the checkpoint must match the keys of the head exactly. Defaults to `False`.

#### Attributes

##### export_output_names

The names of the `n_heads` outputs of the exported model.

The default names are `output1_yolov6r2`, `output2_yolov6r2`, and so on. These names are compatible with DepthAI. The `export_output_names` param replaces them only when it holds exactly `n_heads` names. The head logs a warning each time it gives the default names. The value is never `None`.

##### grid_cell_offset

##### grid_cell_size

##### heads
