# ocr_ctc_head

Python API: `luxonis_train.nodes.heads.ocr_ctc_head`

The OCR head that predicts the character class scores of each sequence step for CTC.

## Classes

### OCRCTCHead

OCR CTC sequence classification head.

 * `Inputs:`: * `inputs` (`Tensor`): [B, C, 1, T]
 * `Outputs:`: * train, eval: * `ocr` (`Tensor`): [B, T, nclasses] logits, class `0` is blank
    * export: * `ocr` (`Tensor`): [B, T, nclasses] softmax probabilities

> **References**
> * Source: Adapted from [PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR) (Apache-2.0).
 * License: Apache-2.0

> **Notes**
> The head removes the height axis of the feature map, so the height must be `1`. Each of the T columns becomes one step of the sequence. One linear layer maps each step to the class logits. With `mid_channels`, two linear layers with a ReLU between them do this.

The classes come from `alphabet`, not from the dataset. nclasses is
[out_channels](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md).
Class `0` is the CTC blank. The unique characters of `alphabet` follow in sorted order. With `ignore_unknown=False`, the last
class is `"<UNK>"`.
[CTCLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/ctc_loss.md)
and
[OCRAccuracy](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/ocr_accuracy.md)
encode the text labels with
[encoder](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md).
[OCRVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/ocr_visualizer.md)
decodes the predictions with
[decoder](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md).

 * `Variants:`: None. Configure the node through `params`.

> **See Also**
> * [The CTC head of PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR/blob/main/ppocr/modeling/heads/rec_ctc_head.py)

> **Example**
> A node entry in the `model.nodes` section of a config:

```yaml
- name: OCRCTCHead
  inputs: [SVTRNeck]
  params:
    alphabet: ["a", "b", "c"]
```

 * `Compatible with:`: * Attach index: `-1`, the last output of the input node
    * Required labels: `metadata/text`
    * Used by:
      [OCRRecognitionModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/ocr_recognition/v1/model.md)
    * Losses:
      [CTCLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/ctc_loss.md)
    * Metrics:
      [OCRAccuracy](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/ocr_accuracy.md)
    * Visualizers:
      [OCRVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/ocr_visualizer.md)
    * Export parser: `ClassificationSequenceParser`

#### Methods

##### init

```python
def __init__(alphabet: list[str], ignore_unknown: bool = True, mid_channels: int | None = None, return_feats: bool = False, **kwargs):
```

Build the encoder, the decoder, and the linear layers.

Parameters

 * `alphabet` (`list[str]`): The characters that the head predicts. The encoder sorts them and puts the blank class before them.
   Each character must occur only once.
 * `ignore_unknown` (`bool`): Whether the encoder drops a label character that is not in `alphabet`. With `False`, the encoder
   maps the character to the extra class `"<UNK>"`, so the head predicts one more class.
 * `mid_channels` (`int | None`): The number of hidden features between two linear layers. `None` gives one linear layer from
   `in_channels` to
   [out_channels](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md).
 * `return_feats` (`bool`): The head stores the value in the `return_feats` attribute.
   [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md)
   does not read it, so the value has no effect.
 * `**kwargs`: Keyword arguments for
   [BaseNode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md).
   They must hold `input_shapes` or `in_sizes`.

Raises

 * `ValueError`: When `alphabet` holds a character more than once.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Compute the class scores of each step of the sequence.

The method removes the height axis and moves the width axis before the channel axis. Each column of the feature map is then one
step. The linear layers map the `C` channels of each step to the class logits. In export mode, the method applies a softmax over
the classes.

> **Example**
> The alphabet `["b", "a"]` gives three classes: the blank, `"a"`, and `"b"`.

```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import OCRCTCHead
>>> head = OCRCTCHead(
...     alphabet=["b", "a"],
...     input_shapes=[{"features": [Size([1, 8, 1, 4])]}],
... )
>>> head(torch.zeros(1, 8, 1, 4)).shape
torch.Size([1, 4, 3])
>>> head.export = True
>>> probabilities = head(torch.zeros(1, 8, 1, 4))
>>> torch.allclose(probabilities.sum(dim=-1), torch.tensor(1.0))
True
```

Parameters

 * `x` (`Tensor`): The feature map of shape `[B, C, 1, T]`. With a height other than `1`, `permute` raises `RuntimeError`.

Returns

 * `Tensor`: The logits of shape `[B, T, out_channels]`. In export mode, the softmax probabilities of the same shape.
   [BaseNode.run](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md)
   puts the result under the `"ocr"` key.

##### get_custom_head_config

```python
def get_custom_head_config(self) -> Params:
```

Return the NN Archive metadata of the head.

[BaseHead.get_head_config](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/base_head.md)
merges the result into the metadata. The `"classes"` and `"n_classes"` keys of the result replace the values from the dataset.
[BaseHead.get_head_config](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/base_head.md)
still reads the class names of the dataset first, so it needs `dataset_metadata`.

> **Example**
> ```pycon
>>> from torch import Size
>>> from luxonis_train.nodes import OCRCTCHead
>>> head = OCRCTCHead(
...     alphabet=["b", "a"],
...     ignore_unknown=False,
...     input_shapes=[{"features": [Size([1, 8, 1, 4])]}],
... )
>>> config = head.get_custom_head_config()
>>> [str(char) for char in config["classes"]]
['', 'a', 'b', '<UNK>']
>>> config["n_classes"], config["ignored_indexes"]
(4, [0])
```

Returns

 * `Params`: A dictionary with these keys. * `"classes"` holds the alphabet of [encoder](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md), with the blank `""` first.
    * `"n_classes"` holds [out_channels](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md).
    * `"is_softmax"` is `True`, because the exported model applies a softmax.
    * `"concatenate_classes"` is `True`.
    * `"ignored_indexes"` is `[0]`, the index of the blank.
    * `"remove_duplicates"` is `True`.

##### initialize_weights

```python
def initialize_weights(method: str | None = None):
```

Initialize the linear layers.

The method first calls [BaseNode.initialize_weights](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md) with `method`. Then it draws the weight and the bias of every `torch.nn.Linear` from the uniform distribution on [ − 1 ⁄ √(nin), 1 ⁄ √(nin)]. nin is the number of input features of the layer. These are the bounds of the PyTorch default initialization of `torch.nn.Linear`.

Parameters

 * `method` (`str | None`): The method for [BaseNode.initialize_weights](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md). With `mid_channels`, `"yolo"` makes the ReLU between the linear layers run in place. Other values change nothing there.

#### Attributes

##### block

##### decoder

The decoder that converts the predictions to text.

It applies a softmax and takes the most probable class at each step. It drops the blank class `0` and each class that repeats the class of the previous step. For each sequence, it returns the text and the mean probability of the kept characters. The mean is `NaN` for an empty text. [OCRVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/ocr_visualizer.md) and [BaseHead.annotate](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/base_head.md) use it.

##### encoder

The encoder that converts text labels to class indices.

[CTCLoss](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/losses/ctc_loss.md) and [OCRAccuracy](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/metrics/ocr_accuracy.md) use it. Its alphabet starts with the blank `""` at index `0`. The sorted unique characters of `alphabet` follow. With `ignore_unknown=False`, `"<UNK>"` is the last entry.

##### export_output_names

The name of the output in the exported model.

The value is always `["output_ocr_ctc"]`. The property ignores the `export_output_names` param.

##### in_channels

The number of channels of the attached inputs.

It is the third dimension from the end of [in_sizes](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md), so a shape with or without the batch dimension gives the same value. A list of sizes gives a list of channel counts.

Raises

 * `RuntimeError`: When [in_sizes](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md) cannot find the input sizes.
 * `ValueError`: When [attach_index](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md) does not fit the sizes.

##### out_channels

The number of classes that the head predicts.

It is the length of the alphabet of [encoder](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md): the blank, the unique characters of `alphabet`, and `"<UNK>"` when `ignore_unknown` is `False`.

##### parser
