# ocr

Python API: `luxonis_train.utils.ocr`

The CTC encoder and decoder between text and the class indices that the OCR head predicts.

## Classes

### OCRDecoder

Greedy CTC decoder that turns class scores into text.

The decoder takes the most probable class at each step of a sequence. It drops the ignored classes and, optionally, the steps that
repeat the class of the previous step. A call of the decoder runs
[decode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/ocr.md).

#### Methods

##### init

```python
def __init__(char_to_int: dict, ignored_tokens: list[int] | None = None, is_remove_duplicate: bool = True):
```

Invert the character mapping and store the options.

Parameters

 * `char_to_int` (`dict`): The class index of each character, as
   [OCREncoder](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/ocr.md)
   builds it.
 * `ignored_tokens` (`list[int] | None`): The class indices to drop. `None` selects `[0]`, the CTC blank. An empty list keeps
   every class.
 * `is_remove_duplicate` (`bool`): Whether to drop a step whose class equals the class of the previous step.

##### decode

```python
def decode(preds: Tensor) -> list[tuple[str, float]]:
```

Decode the class scores of each sequence into text.

The method applies a softmax over the classes and takes the most probable class at each step. It drops each step whose class is in
`ignored_tokens`. With `is_remove_duplicate`, it also drops a step whose class equals the class of the previous step. The
comparison uses the previous step also when the method dropped that step. Thus a blank between two equal characters keeps both
characters.

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.utils import OCRDecoder
>>> decoder = OCRDecoder({"": 0, "a": 1, "b": 2})
>>> classes = torch.tensor([[1, 1, 0, 1, 2]])
>>> logits = torch.nn.functional.one_hot(classes, 3) * 10.0
>>> text, confidence = decoder.decode(logits)[0]
>>> text, round(confidence, 3)
('aab', 1.0)
```

Parameters

 * `preds` (`Tensor`): The logits of shape `[B, T, n_classes]`.

Returns

 * `list[tuple[str, float]]`: One `(text, confidence)` pair for each sequence. The confidence is the mean probability of the kept steps, and `nan` for an empty text.

### OCREncoder

CTC encoder that turns text labels into class indices.

Class `0` is the CTC blank `""`. The sorted unique characters of the alphabet follow. With `ignore_unknown=False`, `"<UNK>"` is the last class. A call of the encoder runs [encode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/ocr.md).

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.utils import OCREncoder
>>> encoder = OCREncoder(["b", "a"])
>>> [str(char) for char in encoder.alphabet], encoder.n_classes
(['', 'a', 'b'], 3)
>>> codes = torch.tensor([[ord("b"), ord("x"), ord("a"), 0]])
>>> encoder.encode(codes).tolist()
[[2, 1, 0, 0]]
>>> strict_encoder = OCREncoder(["b", "a"], ignore_unknown=False)
>>> strict_encoder.encode(codes).tolist()
[[2, 3, 1, 0]]
```

#### Methods

##### init

```python
def __init__(alphabet: list[str], ignore_unknown: bool = True):
```

Build the alphabet and the class index of each character.

Parameters

 * `alphabet` (`list[str]`): The characters of the labels. The encoder sorts them and drops the duplicates.
 * `ignore_unknown` (`bool`): Whether
   [encode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/ocr.md)
   drops a character that is not in `alphabet`. With `False`, the encoder adds the class `"<UNK>"` and maps such a character to
   it.

##### encode

```python
def encode(targets: Tensor) -> Tensor:
```

Convert the character codes of the labels into class indices.

The value `0` is padding and gives the blank class `0`. With `ignore_unknown`, the method drops a character that is not in the
alphabet. The later characters move to the left, and `0` fills the end of the row. Otherwise the character gets the `"<UNK>"`
class.

Parameters

 * `targets` (`Tensor`): The Unicode code points of the labels, of shape `[N, L]`.

Returns

 * `Tensor`: The class indices of shape `[N, L]`, as `int64`.

#### Attributes

##### alphabet

The character of each class, in the order of the indices.

The blank `""` comes first. The sorted unique characters follow, then `"<UNK>"` when `ignore_unknown` is `False`. The sorted
characters are `np.str_` values.

##### char_to_int

The class index of each character.

##### n_classes

The number of classes, the length of
[alphabet](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/utils/ocr.md).
