# ocr_visualizer

Python API: `luxonis_train.attached_modules.visualizers.ocr_visualizer`

The visualizer that writes the recognized text on a white panel.

## Classes

### OCRVisualizer

Visualizer for the text predictions of an OCR head.

The visualizer does not draw on the input images. It returns them with a white panel for each image. The panel shows the target
text and the predicted text with its mean probability.

An input image and its text panel.

 * `Inputs:`: * `prediction_canvas`, `target_canvas` (`Tensor`): [B, 3, H, W]
    * `predictions` (`Tensor`): [B, T, C] logits
    * `targets` (`Tensor | None`): [B, Tmax] Unicode code points, padded with `0`
 * `Outputs:`: * `tuple[Tensor, Tensor]`: [B, 3, H, W], the input images and the text panels

> **References**
> * Source: This project.
 * License: Apache-2.0 (this project)

> **Notes**
> The `decoder` of the attached [OCRCTCHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md) converts the predictions to text. OpenCV writes the text on the panels.

> **Example**
> Attached to a `OCRCTCHead` in `model.nodes`:

```yaml
- name: OCRCTCHead
  inputs: [SVTRNeck]
  visualizers:
    - name: OCRVisualizer
```

 * `Compatible with:`: * Used by:
   [OCRRecognitionModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/predefined_models/ocr_recognition/v1/model.md)
    * Nodes:
      [OCRCTCHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md)

#### Methods

##### init

```python
def __init__(font_scale: float = 0.5, color: tuple[int, int, int] = (0, 0, 0), thickness: int = 1, **kwargs):
```

Initialize the visualizer and store the text options.

Parameters

 * `font_scale` (`float`): The OpenCV font scale of the text.
 * `color` (`tuple[int, int, int]`): The color of the text, one value in `[0, 255]` for each channel, in the channel order of the
   canvas. The default is black.
 * `thickness` (`int`): The line thickness of the text, in pixels.
 * `**kwargs`: Keyword arguments forwarded to
   [BaseVisualizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/base_visualizer.md),
   such as `scale` and `node`. The `node` must be an
   [OCRCTCHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/ocr_ctc_head.md),
   because
   [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/attached_modules/visualizers/ocr_visualizer.md)
   uses its `decoder`. Another node type raises `IncompatibleError`.

##### forward

```python
def forward(prediction_canvas: Tensor, target_canvas: Tensor, predictions: Tensor, targets: Tensor | None) -> tuple[Tensor, Tensor]:
```

Write the target text and the predicted text on white panels.

The `decoder` of the node converts `predictions` to a text and a mean probability for each image. For each image, the method fills
a panel of the canvas size with white. When `targets` is given, it writes `"GT: <target text>"` with the bottom-left corner of the
text at the pixel `(5, 20)`. It writes `"Pred: <predicted text> <probability>"` at `(5, 40)`. The probability has two decimals and
is `nan` for an empty text.

> **Example**
> ```pycon
>>> import torch
>>> from torch import Size
>>> from luxonis_train.nodes import OCRCTCHead
>>> shapes = [{"features": [Size([1, 8, 1, 4])]}]
>>> head = OCRCTCHead(alphabet=["a", "b"], input_shapes=shapes)
>>> visualizer = OCRVisualizer(node=head)
>>> canvas = torch.zeros(1, 3, 48, 96, dtype=torch.uint8)
>>> targets = torch.tensor([[97, 98, 0]])
>>> images, panels = visualizer(
...     canvas, canvas, torch.zeros(1, 4, 3), targets
... )
>>> bool((images == canvas).all()), panels[0, :, 0, 0].tolist()
(True, [255, 255, 255])
```

Parameters

 * `prediction_canvas` (`Tensor`): Images of shape `[B, 3, H, W]`. The method uses only the shape, the dtype, and the device of this tensor.
 * `target_canvas` (`Tensor`): `uint8` images of shape `[B, 3, H, W]`. The panels get the size of these images.
 * `predictions` (`Tensor`): The logits of the node, of shape `[B, T, n_classes]`.
 * `targets` (`Tensor | None`): The `metadata/text` label, of shape `[B, T_max]`. Each row holds the Unicode code points of one text, padded with `0`. `None` when the batch has no text labels.

Returns

 * `tuple[Tensor, Tensor]`: A copy of `target_canvas`, and the panels in a tensor of the same shape.

#### Attributes

##### node
