# cross_entropy

Python API: `luxonis_train.attached_modules.losses.cross_entropy`

Cross entropy over raw logits.

## Classes

### CrossEntropyLoss

Cross entropy between logits and class targets.

 * `Inputs:`: * `predictions` (`Tensor`): `[B, C, ...]` logits
    * `target` (`Tensor`): `[B, C, ...]` one-hot or `[B, ...]` class indices
 * `Outputs:`: * `Tensor`: scalar, or `[B, ...]` when `reduction` is `"none"`
 * `Formula:`: A one-hot target first becomes class indices through `argmax` over the class dimension. For the logits x of one
   element, its class index t, the number of classes C, and `label_smoothing` ε, the loss is ℓ = − C∑c = 1wc qclog(exc)/(C∑k =
   1exk), qc = (1 − ε) [c = t] + (ε)/(C) Here wc is the `weight` of class c, or `1` without `weight`. An element with the target
   `ignore_index` adds nothing. The `"mean"` reduction divides the sum of ℓ by the sum of wt over all elements that do not have
   this target.

> **References**
> * Source: Wraps [torch.nn.CrossEntropyLoss](https://docs.pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html) (BSD-3-Clause).
 * License: Apache-2.0 (this project)

> **Notes**
> The loss adapts a single channel when `predictions` and `target` have the same number of dimensions. When the class dimension of `predictions` has size `1`, the loss puts a channel of zero logits in front of it. A target y with one channel becomes the two channels 1 − y and y. Without `weight` and `label_smoothing`, a target of `0` or `1` gives the binary cross entropy of the single logit. The loss logs a warning at the first such call. [BCEWithLogitsLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/bce_with_logits.md) handles one class directly.

> **Example**
> Attached to a `DDRNetSegmentationHead` in `model.nodes`:

```yaml
- name: DDRNetSegmentationHead
  inputs: [DDRNet]
  losses:
    - name: CrossEntropyLoss
```

 * `Compatible with:`: * Used by:
   [ClassificationModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/classification/v1/model.md)
    * Nodes: *
      [BiSeNetHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/bisenet_head.md)
       * [ClassificationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/classification_head.md)
       * [DDRNetSegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ddrnet_segmentation_head.md)
       * [SegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/segmentation_head.md)
       * [TransformerClassificationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/transformer_classification_head.md)
       * [TransformerSegmentationHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/transformer_segmentation_head.md)

#### Methods

##### init

```python
def __init__(weight: list[float] | None = None, ignore_index: int = -100, reduction: Literal['none', 'mean', 'sum'] = 'mean', label_smoothing: float = 0.0, **kwargs):
```

Initialize the loss and the wrapped `nn.CrossEntropyLoss`.

Parameters

 * `weight` (`list[float] | None`): The factor wc of each class, one value for each class. `None` gives every class the factor
   `1`.
 * `ignore_index` (`int`): A class index that adds nothing to the loss and to the gradient. A one-hot target becomes indices from
   `0` to `C - 1`. The default `-100` therefore affects only a target of class indices.
 * `reduction` (`Literal['none', 'mean', 'sum']`): How to reduce the loss of the elements: * `"none"`: return the loss of each
   element.
    * `"mean"`: return the weighted mean, as the class formula describes.
    * `"sum"`: return the sum over all elements.
 * `label_smoothing` (`float`): The value ε, in `[0, 1]`. The target keeps 1 − ε of its mass and spreads ε evenly over all
   classes, as in [Rethinking the Inception Architecture for Computer Vision](https://arxiv.org/abs/1512.00567).
 * `**kwargs`: Keyword arguments forwarded to
   [BaseLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/base_loss.md),
   such as `final_loss_weight` and `node`.

##### forward

```python
def forward(predictions: Tensor, target: Tensor) -> Tensor:
```

Compute the cross entropy between logits and class targets.

When both tensors have the same number of dimensions, the method treats `target` as one-hot. The class dimension is `1`, or `0`
for tensors with one dimension. When `predictions` has one channel, the method first adds a second channel, as the class notes
describe. It then converts `target` to class indices with `argmax` over the class dimension. For a multi-hot target, `argmax`
selects the first class with the highest value.

> **Example**
> A one-hot target and the same classes as indices give the same loss:

```pycon
>>> import torch
>>> loss = CrossEntropyLoss()
>>> logits = torch.tensor([[2.0, 0.0], [0.0, 2.0]])
>>> one_hot = torch.tensor([[1.0, 0.0], [0.0, 1.0]])
>>> round(loss(logits, one_hot).item(), 4)
0.1269
>>> round(loss(logits, torch.tensor([0, 1])).item(), 4)
0.1269
```

Parameters

 * `predictions` (`Tensor`): Logits of shape `[B, C, ...]`, the main output of the node.
 * `target` (`Tensor`): One-hot targets of shape `[B, C, ...]`, or class indices of shape `[B, ...]`. The `classification` and
   `segmentation` labels have the shape `[B, C, ...]`.

Returns

 * `Tensor`: A scalar for the `"mean"` and `"sum"` reductions. For `"none"`, the loss of each element, of shape `[B, ...]`.

Raises

 * `RuntimeError`: When `target` has neither the number of dimensions of `predictions` nor one dimension less.

#### Attributes

##### criterion

##### supported_tasks
