# softmax_focal_loss

Python API: `luxonis_train.attached_modules.losses.softmax_focal_loss`

Softmax focal loss, which lowers the weight of the easy examples.

## Classes

### SoftmaxFocalLoss

Focal loss on softmax outputs, for multiclass predictions.

The loss applies a softmax over the class dimension. The focal factor lowers the loss of the elements whose target class already
has a high probability.

 * `Inputs:`: * `predictions` (`Tensor`): `[B, C, ...]` logits, with C ≥ 2
    * `targets` (`Tensor`): same shape, one-hot
 * `Outputs:`: * `Tensor`: scalar, or `[B, ...]` when `reduction` is `"none"`
 * `Formula:`: At one element, pc is the softmax probability of the class c, and yc is its target. For a smoothing factor s > 0,
   the loss first clips each target to [s ⁄ (C − 1), 1 − s]. The loss of the element is then pt = ∑cyc pc + s, ℓ = − αt (1 −
   pt)γlogpt A float `alpha` is αt for every element. A list `alpha` gives αt from the entry of the target class, the `argmax` of
   the targets. `reduction` then takes the mean or the sum over all elements, or keeps the loss of each element.

> **References**
> * Source: This project.
 * License: Apache-2.0 (this project)

> **Notes**
> The loss runs in `float32` with autocast turned off, also in a mixed precision run. With `alpha=1`, `gamma=0`, and `smooth=0`, the loss is the cross entropy.

> **Example**
> Attached to a `DDRNetSegmentationHead` in `model.nodes`:

```yaml
- name: DDRNetSegmentationHead
  inputs: [DDRNet]
  losses:
    - name: SoftmaxFocalLoss
```

 * `Compatible with:`: * Nodes: *
   [BiSeNetHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/bisenet_head.md)
       * [ClassificationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/classification_head.md)
       * [DDRNetSegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ddrnet_segmentation_head.md)
       * [SegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/segmentation_head.md)
       * [TransformerClassificationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/transformer_classification_head.md)
       * [TransformerSegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/transformer_segmentation_head.md)

#### Methods

##### init

```python
def __init__(alpha: float | list[float] = 0.25, gamma: float = 2.0, smooth: float = 0.0, reduction: Literal['none', 'mean', 'sum'] = 'mean', **kwargs):
```

Initialize the loss and check the smoothing factor.

Parameters

 * `alpha` (`float | list[float]`): The class weight αt. A float scales the loss of every element by the same factor, so it does
   not favor a class. A list holds one weight for each class, in class order.
   [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/softmax_focal_loss.md)
   then checks that the list has one entry for each class.
 * `gamma` (`float`): The exponent of the focal factor (1 − pt). `0` turns the focal factor off.
 * `smooth` (`float`): The label smoothing factor s, in `[0, 1]`. The class formula shows how it changes the targets and pt.
 * `reduction` (`Literal['none', 'mean', 'sum']`): How to reduce the loss of the elements: * `"none"`: return the loss of each
   element.
    * `"mean"`: return the mean over all elements.
    * `"sum"`: return the sum over all elements.
      [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/softmax_focal_loss.md)
      treats any other value as `"none"`.
 * `**kwargs`: Keyword arguments forwarded to
   [BaseLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/base_loss.md),
   such as `final_loss_weight` and `node`.

Raises

 * `ValueError`: When `smooth` is outside `[0, 1]`.

##### forward

```python
def forward(predictions: Tensor, targets: Tensor) -> Tensor:
```

Compute the softmax focal loss between logits and targets.

> **Examples**
> With `alpha=1` and `gamma=0`, the loss is the cross entropy:

```pycon
>>> import torch
>>> import torch.nn.functional as F
>>> logits = torch.tensor([[2.0, 0.0, -1.0]])
>>> target = torch.tensor([[1.0, 0.0, 0.0]])
>>> loss = SoftmaxFocalLoss(alpha=1.0, gamma=0.0)
>>> ce = F.cross_entropy(logits, target)
>>> torch.allclose(loss(logits, target), ce)
True
```

One class is not enough:

```pycon
>>> loss(torch.zeros(1, 1), torch.ones(1, 1))
Traceback (most recent call last):
    ...
ValueError: SoftmaxFocalLoss is not suitable for binary tasks. Please use SigmoidFocalLoss instead.
```

Parameters

 * `predictions` (`Tensor`): Logits of shape `[B, C, ...]`, with at least two classes. The main output of the node.
 * `targets` (`Tensor`): One-hot targets of the same shape as `predictions`.

Returns

 * `Tensor`: A `float32` scalar for the `"mean"` and `"sum"` reductions. For any other `reduction`, the loss of each element, of
   shape `[B, ...]`.

Raises

 * `ValueError`: When `predictions` has fewer than two classes, when the shapes differ, or when a list `alpha` does not have one
   entry for each class.

#### Attributes

##### supported_tasks
