# sigmoid_focal_loss

Python API: `luxonis_train.attached_modules.losses.sigmoid_focal_loss`

Binary focal loss, which lowers the weight of the easy examples.

## Classes

### SigmoidFocalLoss

Focal loss on independent sigmoid outputs.

The loss treats each element of the logits as one binary decision, so it fits binary and multi-label tasks. The focal factor
lowers the loss of the elements that the model already predicts well.

 * `Inputs:`: * `predictions` (`Tensor`): `[B, C, ...]` logits
    * `target` (`Tensor`): same shape, float values in `[0, 1]`
 * `Outputs:`: * `Tensor`: scalar, or `[B, C, ...]` when `reduction` is `"none"`
 * `Formula:`: For a logit x with the target y, let p = σ(x) and pt = p**y + (1 − p)(1 − y). The loss of one element is ℓ = − αt
   (1 − pt)γ[ylogp + (1 − y)log(1 − p)] Here αt = α**y + (1 − α)(1 − y), and an `alpha` of `-1` gives αt = 1. `reduction` then
   takes the mean or the sum over all elements, or keeps the loss of each element.

> **References**
> * Source: Wraps [torchvision.ops.sigmoid_focal_loss](https://docs.pytorch.org/vision/stable/generated/torchvision.ops.sigmoid_focal_loss.html) (BSD-3-Clause).
 * Paper: [Focal Loss for Dense Object Detection](https://arxiv.org/abs/1708.02002)
 * License: Apache-2.0 (this project)

> **Notes**
> The class wraps `torchvision.ops.sigmoid_focal_loss` and adds no checks of its own.

> **Example**
> Attached to a `DDRNetSegmentationHead` in `model.nodes`:

```yaml
- name: DDRNetSegmentationHead
  inputs: [DDRNet]
  losses:
    - name: SigmoidFocalLoss
```

 * `Compatible with:`: * Nodes: *
   [BiSeNetHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/bisenet_head.md)
       * [ClassificationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/classification_head.md)
       * [DDRNetSegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ddrnet_segmentation_head.md)
       * [SegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/segmentation_head.md)
       * [TransformerClassificationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/transformer_classification_head.md)
       * [TransformerSegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/transformer_segmentation_head.md)

#### Methods

##### init

```python
def __init__(alpha: float = 0.25, gamma: float = 2.0, reduction: Literal['none', 'mean', 'sum'] = 'mean', **kwargs):
```

Initialize the loss and store the focal parameters.

The constructor does not check the values.
[forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/sigmoid_focal_loss.md)
passes them to `torchvision`, which raises `ValueError` for a value that it does not accept.

Parameters

 * `alpha` (`float`): The weight α of the positive elements, in `[0, 1]`. The negative elements get 1 − α. `-1` turns the
   weighting off.
 * `gamma` (`float`): The exponent of the focal factor (1 − pt). A larger value lowers the loss of the well-predicted elements
   more. `0` gives the binary cross entropy, weighted by `alpha`.
 * `reduction` (`Literal['none', 'mean', 'sum']`): How to reduce the loss of the elements: * `"none"`: return the loss of each
   element.
    * `"mean"`: return the mean over all elements.
    * `"sum"`: return the sum over all elements.
 * `**kwargs`: Keyword arguments forwarded to
   [BaseLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/base_loss.md),
   such as `final_loss_weight` and `node`.

##### forward

```python
def forward(predictions: Tensor, target: Tensor) -> Tensor:
```

Compute the sigmoid focal loss between logits and targets.

`torchvision` raises `ValueError` when the shapes differ, and when `alpha` or `reduction` has a value that it does not accept.

> **Example**
> With `alpha=-1` and `gamma=0`, the loss is the binary cross entropy. The default focal factor makes it smaller:

```pycon
>>> import torch
>>> import torch.nn.functional as F
>>> logits = torch.tensor([[2.0, -1.0]])
>>> target = torch.tensor([[1.0, 0.0]])
>>> bce = F.binary_cross_entropy_with_logits(logits, target)
>>> plain = SigmoidFocalLoss(alpha=-1.0, gamma=0.0)
>>> torch.allclose(plain(logits, target), bce)
True
>>> bool(SigmoidFocalLoss()(logits, target) < bce)
True
```

Parameters

 * `predictions` (`Tensor`): Logits of shape `[B, C, ...]`, the main output of the node.
 * `target` (`Tensor`): Float targets in `[0, 1]`, of the same shape as `predictions`.

Returns

 * `Tensor`: A scalar for the `"mean"` and `"sum"` reductions. For `"none"`, the loss of each element, of shape `[B, C, ...]`.

#### Attributes

##### supported_tasks
