# blocks

Python API: `luxonis_train.nodes.necks.svtr_neck.blocks`

The token mixers and the transformer block of the SVTR neck.

## Classes

### Attention

Multi-head self-attention over a sequence of tokens.

One linear layer computes the queries Q, the keys K, and the values V of all heads. Each head computes

softmax(s Q**KT + M)V

where s is the scale and M is the mask. A linear layer then merges the heads. Dropout follows the softmax and the merge.

The `"global"` mixer has no mask, so each token attends to every token. The `"local"` mixer lets a token attend only to the tokens
in a window of `kernel_size` around it. For odd sizes, the window is centered on the token. The mask is `0` in the window and − ∞
outside. It has the shape `[1, 1, N, N]`, where `N` is `height * width`.

Warning: The mask is a plain attribute, not a buffer. `torch.nn.Module.to` does not move it, and the state dictionary does not
hold it. The mask stays on the CPU, so the `"local"` mixer works only for an input on the CPU.

> **Example**
> With a `3x3` window, token `7` in row `1` and column `2` sees three columns of the `3x5` map. Only the tokens that it sees get a gradient from its output:

```pycon
>>> import torch
>>> from luxonis_train.nodes.necks.svtr_neck.blocks import Attention
>>> attention = Attention(
...     8, height=3, width=5, n_heads=2, mixer="local", kernel_size=3
... )
>>> tokens = torch.randn(1, 15, 8, requires_grad=True)
>>> attention(tokens)[0, 7].sum().backward()
>>> seen = tokens.grad[0].abs().sum(dim=1) > 0
>>> seen.view(3, 5).int().tolist()
[[0, 1, 1, 1, 0], [0, 1, 1, 1, 0], [0, 1, 1, 1, 0]]
```

#### Methods

##### init

```python
def __init__(dim: int, height: int | None = None, width: int | None = None, n_heads: int = 8, mixer: Literal['global', 'local'] = 'global', kernel_size: tuple[int, int] | int = (7, 11), qk_scale: float | None = None, attn_drop: float = 0.0, proj_drop: float = 0.0):
```

Build the linear layers and the mask of the local mixer.

Parameters

 * `dim` (`int`): The number of channels of each token. It must be a multiple of `n_heads`. When `qk_scale` is `None` or `0`, a
   value below `n_heads` makes the constructor raise `ZeroDivisionError`.
 * `height` (`int | None`): The height of the token map. The `"local"` mixer needs it. The `"global"` mixer ignores it.
 * `width` (`int | None`): The width of the token map. The `"local"` mixer needs it. The `"global"` mixer ignores it.
 * `n_heads` (`int`): The number of attention heads. Each head gets `dim // n_heads` channels.
 * `mixer` (`Literal['global', 'local']`): The attention type. `"global"` has no mask. `"local"` builds the window mask.
 * `kernel_size` (`tuple[int, int] | int`): The height and the width of the window of the `"local"` mixer. An integer gives a
   square window. The `"global"` mixer ignores it.
 * `qk_scale` (`float | None`): The scale s of the queries. `None` or `0` gives 1 ⁄ √(d), where d is `dim // n_heads`.
 * `attn_drop` (`float`): The dropout probability of the attention weights.
 * `proj_drop` (`float`): The dropout probability after the linear layer that merges the heads.

Raises

 * `ValueError`: When `mixer` is `"local"` and `height` or `width` is `None`.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Apply the self-attention to a sequence of tokens.

Parameters

 * `x` (`Tensor`): The tokens, of shape `[B, N, dim]`. With the `"local"` mixer, `N` must be `height * width`, and `x` must be on
   the CPU. On another device, the addition of the CPU mask raises `RuntimeError`.

Returns

 * `Tensor`: The attended tokens, of shape `[B, N, dim]`.

#### Attributes

##### attn_drop

##### proj

##### proj_drop

##### qkv

### ConvMixer

Token mixer of
[SVTRBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/necks/svtr_neck/blocks.md)
that uses a grouped convolution.

The mixer reshapes a sequence of tokens into a map of `height` by `width` tokens. A convolution with `n_heads` groups and a bias
mixes each token with its neighbors. The mixer then flattens the map back into a sequence, row by row.

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.nodes.necks.svtr_neck.blocks import ConvMixer
>>> mixer = ConvMixer(8, height=2, width=4, n_heads=2)
>>> mixer(torch.zeros(1, 8, 8)).shape
torch.Size([1, 8, 8])
```

#### Methods

##### init

```python
def __init__(dim: int, height: int, width: int, n_heads: int, kernel_size: tuple[int, int] = (3, 3)):
```

Build the grouped convolution.

Parameters

 * `dim` (`int`): The number of channels of each token. It must be a multiple of `n_heads`.
 * `height` (`int`): The height of the token map.
 * `width` (`int`): The width of the token map.
 * `n_heads` (`int`): The number of groups of the convolution.
 * `kernel_size` (`tuple[int, int]`): The height and the width of the kernel. The padding is half of each value, rounded down. Thus only odd values keep the size of the map.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Mix each token with its neighbors in the token map.

Parameters

 * `x` (`Tensor`): The tokens, of shape `[B, height * width, dim]`. With another number of tokens, `reshape` raises `RuntimeError`.

Returns

 * `Tensor`: The mixed tokens, of the shape of `x` for an odd kernel size.

#### Attributes

##### local_mixer

### SVTRBlock

Transformer block of SVTR with a token mixer and an MLP.

The block has two residual branches. The first branch runs the mixer:

 * `"global"` and `"local"` build an [Attention](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/necks/svtr_neck/blocks.md).
 * `"conv"` builds a [ConvMixer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/necks/svtr_neck/blocks.md).

The second branch runs an MLP: a linear layer to `int(dim * mlp_ratio)` features, `act_layer`, dropout, a linear layer back to `dim`, and dropout. When `drop_path` is above `0`, [DropPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md) drops each branch for random samples in training mode.

`prenorm` sets the place of the two norm layers. f is a branch:

 * `True`: after each residual sum, x ← norm(x + f(x)).
 * `False`: before each branch, x ← x + f(norm(x)).

The flag name does not match the usual term: `prenorm=True` puts the norm layers after the sums.

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.nodes.necks.svtr_neck.blocks import SVTRBlock
>>> block = SVTRBlock(8, n_heads=2, mixer="conv", height=2, width=4)
>>> block(torch.zeros(1, 8, 8)).shape
torch.Size([1, 8, 8])
```

#### Methods

##### init

```python
def __init__(dim: int, n_heads: int, height: int | None = None, width: int | None = None, mixer: Literal['global', 'local', 'conv'] = 'global', mixer_kernel_size: tuple[int, int] = (7, 11), mlp_ratio: float = 4.0, qk_scale: float | None = None, dropout: float = 0.0, attn_drop: float = 0.0, drop_path: float = 0.0, act_layer: type[nn.Module] = nn.GELU, norm_layer: type[nn.Module] = nn.LayerNorm, epsilon: float = 1e-06, prenorm: bool = True):
```

Build the mixer, the MLP, and the two norm layers.

Parameters

 * `dim` (`int`): The number of channels of each token. It must be a multiple of `n_heads`.
 * `n_heads` (`int`): The number of attention heads, or the number of convolution groups of the `"conv"` mixer.
 * `height` (`int | None`): The height of the token map. The `"local"` and `"conv"` mixers need it.
 * `width` (`int | None`): The width of the token map. The `"local"` and `"conv"` mixers need it.
 * `mixer` (`Literal['global', 'local', 'conv']`): The token mixer.
   [Attention](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/necks/svtr_neck/blocks.md)
   raises `ValueError` for `"local"` without `height` and `width`.
 * `mixer_kernel_size` (`tuple[int, int]`): The window of the `"local"` mixer, or the kernel of the `"conv"` mixer. The `"global"`
   mixer ignores it.
 * `mlp_ratio` (`float`): The number of hidden features of the MLP, as a multiple of `dim`.
 * `qk_scale` (`float | None`): The scale of the attention queries. `None` or `0` gives 1 ⁄ √(d), where d is `dim // n_heads`. The
   `"conv"` mixer ignores it.
 * `dropout` (`float`): The dropout probability of the two dropout layers of the MLP. The `"global"` and `"local"` mixers also
   apply it to their output.
 * `attn_drop` (`float`): The dropout probability of the attention weights. The `"conv"` mixer ignores it.
 * `drop_path` (`float`): The probability that
   [DropPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
   drops a branch for a sample in training mode. `0.0` adds no
   [DropPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md).
 * `act_layer` (`type[nn.Module]`): The activation class of the MLP. The MLP calls it without arguments.
 * `norm_layer` (`type[nn.Module]`): The norm class. The block calls it as `norm_layer(dim, eps=epsilon)`.
 * `epsilon` (`float`): The `eps` of both norm layers.
 * `prenorm` (`bool`): `True` puts the norm layers after the residual sums. `False` puts them before the branches.

Raises

 * `ValueError`: When `mixer` is `"conv"` and `height` or `width` is `None`.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Run the mixer branch and then the MLP branch.

Parameters

 * `x` (`Tensor`): The tokens, of shape `[B, N, dim]`. The `"local"` and `"conv"` mixers need `N` to be `height * width`.

Returns

 * `Tensor`: The tokens after both residual branches, of shape `[B, N, dim]`.

#### Attributes

##### drop_path

##### mixer

##### mlp

##### norm1

##### norm2
