# fomo_head

Python API: `luxonis_train.nodes.heads.fomo_head`

The FOMO head, which predicts a heatmap of object centers instead of boxes.

## Classes

### FOMOHead

FOMO heatmap detection head.

 * `Inputs:`: * `inputs` (`Tensor`): [B, C, Hf, Wf]
 * `Outputs:`: * train: * `heatmap` (`Tensor`): [B, nclasses, Hf, Wf] logits
    * eval: * `heatmap` (`Tensor`): [B, nclasses, Hf, Wf] logits
       * `keypoints` (`list[Tensor]`): [Ki, 1, 4] per image, `(x, y, prob, class)`, pixels
    * export: * `outputs` (`list[Tensor]`): [B, nclasses, Hf, Wf], max-pooled when `use_nms`

> **References**
> * Source: This project.
 * License: Apache-2.0 (this project)

> **Notes**
> The head runs a stack of `1x1` convolutions, so each heatmap cell sees only its own feature vector. In evaluation mode, the head keeps the cells with a probability above `0.5` as keypoints. When `use_nms` is `True`, the head also applies a `3x3` max pooling. In evaluation mode, only the local maxima then become keypoints. In export mode, the pooled heatmap replaces the heatmap.

 * `Variants:`: None. Configure the node through `params`.

> **Example**
> A node entry in the `model.nodes` section of a config:

```yaml
- name: FOMOHead
  inputs: [EfficientRep]
```

 * `Compatible with:`: * Attach index: `1`, output 1 of the input node
    * Required labels: `boundingbox`
    * Used by:
      [FOMOModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/fomo/v1/model.md)
    * Losses:
      [FOMOLocalizationLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/fomo_localization_loss.md)
    * Metrics: *
      [ConfusionMatrix](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/confusion_matrix/confusion_matrix.md)
       * [ObjectKeypointSimilarity](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/object_keypoint_similarity.md)
    * Visualizers: *
      [FOMOVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/fomo_visualizer.md)
       * [KeypointVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/keypoint_visualizer.md)

#### Methods

##### init

```python
def __init__(n_conv_layers: int = 3, conv_channels: int = 16, use_nms: bool = True, **kwargs):
```

Initialize the stack of `1x1` convolutions.

The stack has `n_conv_layers - 1` hidden
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
layers with `conv_channels` output channels, a bias, ReLU, and no batch norm. The first hidden layer reads `in_channels` channels.
A last `1x1` convolution maps `conv_channels` channels to `n_classes` channels. This last layer always expects `conv_channels`
input channels. Thus, an `n_conv_layers` below `2` works only when the input has `conv_channels` channels.

Parameters

 * `n_conv_layers` (`int`): Number of convolutions, the last one included. Defaults to `3`.
 * `conv_channels` (`int`): Number of channels of the hidden layers. Defaults to `16`.
 * `use_nms` (`bool`): Whether to apply a `3x3` max pooling with stride `1`. In evaluation mode, only the local maxima of the
   heatmap then become keypoints. In export mode, the pooled heatmap replaces the heatmap. Defaults to `True`.
 * `**kwargs`: Keyword arguments for
   [BaseNode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md),
   such as `n_classes` and `input_shapes`.

##### forward

```python
def forward(inputs: Tensor) -> Packet[Tensor]:
```

Predict the class heatmap and build the packet of the mode.

Training mode has priority over export mode. The packet depends on the mode:

 * Training mode: `"heatmap"` holds the logits, of shape `[B, n_classes, H, W]`.
 * Export mode: `"outputs"` holds a list with one tensor, the heatmap logits of shape `[B, n_classes, H, W]`. When `use_nms` is
   `True`, a `3x3` max pooling with stride `1` first replaces each cell with the maximum of its neighborhood.
 * Evaluation mode: `"heatmap"`, and `"keypoints"` with a tensor of shape `[K_i, 1, 4]` for each image. The head applies a sigmoid
   and keeps the cells with a probability above `0.5`. When `use_nms` is `True`, a cell must also equal the maximum of its `3x3`
   neighborhood. Each of the `K_i` keypoints holds `[x, y, probability, class]`. `x` and `y` are the top-left corner of the cell,
   in the pixels of the original input.

> **Example**
> ```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import FOMOHead
>>> sizes = [Size([1, 8, 64, 64]), Size([1, 16, 32, 32])]
>>> head = FOMOHead(
...     n_classes=2,
...     input_shapes=[{"features": sizes}],
...     original_in_shape=Size([3, 128, 128]),
... )
>>> head.in_channels
16
>>> head(torch.zeros(1, 16, 32, 32))["heatmap"].shape
torch.Size([1, 2, 32, 32])
```

In evaluation mode, the packet also holds one keypoint tensor for each image:

```pycon
>>> out = head.eval()(torch.zeros(1, 16, 32, 32))
>>> sorted(out)
['heatmap', 'keypoints']
>>> len(out["keypoints"]), out["keypoints"][0].shape[1:]
(1, torch.Size([1, 4]))
```

Parameters

 * `inputs` (`Tensor`): Feature map of shape `[B, C, H, W]`. By default, it is output `1` of the input node.

Returns

 * `Packet[Tensor]`: The packet of the current mode.

#### Attributes

##### attach_index

##### conv_layers

##### in_channels

The number of channels of the attached inputs.

It is the third dimension from the end of [in_sizes](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md), so a shape with or without the batch dimension gives the same value. A list of sizes gives a list of channel counts.

Raises

 * `RuntimeError`: When [in_sizes](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md) cannot find the input sizes.
 * `ValueError`: When [attach_index](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md) does not fit the sizes.

##### n_keypoints

The number of keypoints of each object, always `1`.

The head predicts one point for each object. It ignores the `n_keypoints` param and the dataset metadata.
