# efficient_keypoint_bbox_head

Python API: `luxonis_train.nodes.heads.efficient_keypoint_bbox_head`

The YOLOv6 detection head with a keypoint branch for each scale.

## Classes

### EfficientKeypointBBoxHead

Efficient object and keypoint detection head.

 * `Inputs:`: * `inputs` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale
 * `Outputs:`: * train: * `features` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale
       * `class_scores` (`Tensor`): [B, N, nclasses]
       * `distributions` (`Tensor`): [B, N, 4], `(l, t, r, b)` in stride units
       * `keypoints_raw` (`Tensor`): [B, N, 3**n*keypoints]
    * eval: * `features` (`list[Tensor]`): [B, Ci, Hi, Wi] per scale
       * `class_scores` (`Tensor`): [B, N, nclasses]
       * `distributions` (`Tensor`): [B, N, 4], `(l, t, r, b)` in stride units
       * `keypoints_raw` (`Tensor`): [B, N, 3**n*keypoints]
       * `boundingbox` (`list[Tensor]`): [Mi, 6] per image, `[x1, y1, x2, y2, conf, class]`, pixels
       * `keypoints` (`list[Tensor]`): [Mi, nkeypoints, 3] per image, `(x, y, conf)`, pixels
       * `detections_pre_nms` (`Tensor`): [B, N, 5 + nclasses + 3**n*keypoints], only when requested
    * export: * `boundingbox` (`list[Tensor]`): [B, 5 + nclasses, Hi, Wi] per scale
       * `keypoints` (`list[Tensor]`): [B, 3n*keypoints, HiW*i] per scale, `(x, y, conf)`, pixels, `conf` as a logit

> **References**
> * Source: Reimplemented from [YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications](https://arxiv.org/abs/2209.02976).
 * License: Apache-2.0 (this project)

> **Notes**
> The head adds a keypoint branch for each scale to [EfficientBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md). It decodes the keypoints around the anchor points of the scale. NMS keeps the keypoints of each box that it keeps.

 * `Variants:`: None. Configure the node through `params`.

> **Example**
> A node entry in the `model.nodes` section of a config:

```yaml
- name: EfficientKeypointBBoxHead
  inputs: [RepPANNeck]
```

 * `Compatible with:`: * Required labels: * `boundingbox`
       * `keypoints`
    * Used by:
      [KeypointDetectionModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/keypoint_detection/v1/model.md)
    * Losses:
      [EfficientKeypointBBoxLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/efficient_keypoint_bbox_loss.md)
    * Metrics: *
      [ConfusionMatrix](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/confusion_matrix/confusion_matrix.md)
       * [MeanAveragePrecision](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/mean_average_precision/mean_average_precision.md)
       * [ObjectKeypointSimilarity](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/object_keypoint_similarity.md)
       * [PrecisionRecallCurve](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/precision_recall_curve.md)
    * Visualizers:
      [KeypointVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/keypoint_visualizer.md)
    * Export parser: `YOLOExtendedParser`
    * Pretrained weights: the box branches of
      [EfficientBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md),
      through `weights: download`

#### Methods

##### init

```python
def __init__(n_heads: Literal[2, 3, 4] = 3, conf_thres: float = 0.25, iou_thres: float = 0.45, max_det: int = 300, **kwargs):
```

Initialize the box head and the keypoint branches.

Each keypoint branch has two `3x3`
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
layers with batch norm and SiLU, then a `1x1` convolution with `3 * n_keypoints` output channels. The hidden layers of all
branches have the larger of `in_channels[0] // 4` and `3 * n_keypoints` channels. `in_channels[0]` belongs to the first scale.

Parameters

 * `n_heads` (`Literal[2, 3, 4]`): Number of scales. The head reads the last `n_heads` outputs of the input node. An
   `attach_index` param replaces this selection. The value is usually equal to the number of neck outputs. When the input node
   gives fewer outputs, the head logs a warning and uses that number. Defaults to `3`.
 * `conf_thres` (`float`): NMS keeps only the boxes whose maximum class score is above this value. The value must be in `[0, 1]`.
   Defaults to `0.25`.
 * `iou_thres` (`float`): NMS removes a box when its IoU with a box of the same class and a higher score is above this value. The
   value must be in `[0, 1]`. Defaults to `0.45`.
 * `max_det` (`int`): Maximum number of boxes that NMS keeps for each image. Defaults to `300`.
 * `**kwargs`: Keyword arguments for
   [EfficientBBoxHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md),
   such as `bias_init_p`, and for
   [BaseNode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md),
   such as `n_classes`, `n_keypoints`, and `input_shapes`.

##### forward

```python
def forward(inputs: list[Tensor]) -> Packet[Tensor]:
```

Run the box and keypoint branches and build the packet.

The box branches work as in
[EfficientBBoxHead.forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md).
The keypoint branch of each scale reads the input feature map of the scale. It predicts the raw values (tx, ty, tc) of every
keypoint at every anchor point. The head decodes them with the anchor point (ax, ay) in grid units and the stride s of the scale:

x = (2tx + ax − 0.5) s, y = (2ty + ay − 0.5) s, c = σ(tc)

Export mode has priority over training mode. The packet depends on the mode:

 * Export mode: `"boundingbox"` as in
   [EfficientBBoxHead.forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md),
   and `"keypoints"` with one tensor of shape `[B, 3 * n_keypoints, H_i * W_i]` for each scale. The tensor holds x, y, and tc for
   each keypoint. Export mode skips the sigmoid.
 * Training mode: `"features"`, `"class_scores"`, and `"distributions"` as in
   [EfficientBBoxHead.forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/efficient_bbox_head.md),
   and `"keypoints_raw"` with the raw values of shape `[B, N, 3 * n_keypoints]`.
 * Evaluation mode: the training keys and the NMS results. `"boundingbox"` holds a tensor of shape `[M_i, 6]` for each image, with
   the rows `[x1, y1, x2, y2, score, class]`. `"keypoints"` holds the decoded keypoints of these boxes, of shape `[M_i,
   n_keypoints, 3]`. Each keypoint holds (x, y, c). For an image without boxes, `M_i` is `0`. After a call to
   [BaseDetectionHead.request_detections_pre_nms](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/base_detection_head.md),
   the packet also holds the NMS input `"detections_pre_nms"`, of shape `[B, N, 5 + n_classes + 3 * n_keypoints]`.

> **Example**
> ```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import EfficientKeypointBBoxHead
>>> sizes = [Size([1, 8, 32, 32]), Size([1, 16, 16, 16])]
>>> head = EfficientKeypointBBoxHead(
...     n_heads=2,
...     n_classes=3,
...     n_keypoints=5,
...     input_shapes=[{"features": sizes}],
...     original_in_shape=Size([3, 256, 256]),
... )
>>> out = head([torch.zeros(size) for size in sizes])
>>> out["keypoints_raw"].shape
torch.Size([1, 1280, 15])
```

A new head gives each class the score `0.01`, which is below `conf_thres`. In evaluation mode, NMS thus keeps no box and no keypoint:

```pycon
>>> out = head.eval()([torch.zeros(size) for size in sizes])
>>> out["boundingbox"][0].shape, out["keypoints"][0].shape
(torch.Size([0, 6]), torch.Size([0, 5, 3]))
```

Parameters

 * `inputs` (`list[Tensor]`): One feature map for each scale, of shape `[B, C_i, H_i, W_i]`.

Returns

 * `Packet[Tensor]`: The packet of the current mode.

##### get_custom_head_config

```python
def get_custom_head_config(self) -> Params:
```

Return the NMS settings and the strides for the NN Archive.

A subclass adds its own keys to this dictionary, for example `"subtype"`.

Returns

 * `Params`: A dictionary with the keys `"iou_threshold"`, `"conf_threshold"`, `"max_det"`, and `"strides"`. They hold `iou_thres`, `conf_thres`, `max_det`, and `stride` as a list with one integer for each scale.

#### Attributes

##### export_output_names

The names of the outputs of the exported model.

The exported model has `2 * n_heads` outputs. The default names are `output1_yolov6` to `output{n_heads}_yolov6` for the boxes, then `kpt_output1` to `kpt_output{n_heads}` for the keypoints. The `export_output_names` param replaces them only when it holds exactly `n_heads` names. The value then has fewer names than the exported model has outputs. The head logs a warning each time it gives the default names. The value is never `None`.

##### keypoint_heads

##### parser
