# blocks

Python API: `luxonis_train.nodes.backbones.efficientvit.blocks`

The convolution and attention blocks of the EfficientViT backbone.

## Classes

### DepthWiseSeparableConv

Depthwise separable convolution with an optional residual connection.

The block runs a depthwise
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
and then a `1x1` pointwise
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md).
The depthwise convolution has one group for each input channel. Both convolutions have a batch norm.

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.nodes.backbones.efficientvit.blocks import (
...     DepthWiseSeparableConv,
... )
>>> block = DepthWiseSeparableConv(8, 16, stride=2)
>>> block(torch.zeros(1, 8, 15, 15)).shape
torch.Size([1, 16, 8, 8])
```

#### Methods

##### init

```python
def __init__(in_channels: int, out_channels: int, kernel_size: int = 3, stride: int = 1, depthwise_bias: bool = False,
pointwise_bias: bool = False, depthwise_activation: nn.Module | None = None, pointwise_activation: nn.Module | None = None,
padding: int | str | None = None, dilation: int | tuple[int, int] = 1, use_residual: bool = False):
```

Build the depthwise and the pointwise convolutions.

Parameters

 * `in_channels` (`int`): Number of input channels. The depthwise convolution keeps this number of channels.
 * `out_channels` (`int`): Number of output channels of the pointwise convolution.
 * `kernel_size` (`int`): Kernel size of the depthwise convolution. Defaults to `3`.
 * `stride` (`int`): Stride of the depthwise convolution. Defaults to `1`.
 * `depthwise_bias` (`bool`): Whether the depthwise convolution has a bias term. Defaults to `False`.
 * `pointwise_bias` (`bool`): Whether the pointwise convolution has a bias term. Defaults to `False`.
 * `depthwise_activation` (`nn.Module | None`): The activation after the depthwise convolution. `None` selects `torch.nn.ReLU6`.
 * `pointwise_activation` (`nn.Module | None`): The activation after the pointwise convolution. `None` selects no activation.
 * `padding` (`int | str | None`): Padding of the depthwise convolution, or the string `"same"` or `"valid"`. `None` selects `kernel_size // 2`. This value keeps the size only for an odd kernel size, a stride of `1`, and a dilation of `1`.
 * `dilation` (`int | tuple[int, int]`): Dilation of the depthwise convolution. Defaults to `1`.
 * `use_residual` (`bool`): Whether [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) adds the input to the output. The input and the output must then have the same shape. Defaults to `False`.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Apply the depthwise and the pointwise convolutions.

Parameters

 * `x` (`Tensor`): Input of shape `[B, in_channels, H, W]`.

Returns

 * `Tensor`: Output of shape `[B, out_channels, H', W']`. With the default padding, an odd kernel size, and a dilation of `1`, H’ = ⌈H ⁄ s⌉ and W’ = ⌈W ⁄ s⌉, where s is `stride`. When `use_residual` is `True`, the method adds the input to the result.

#### Attributes

##### depthwise_conv

##### pointwise_conv

### EfficientViTBlock

EfficientViT block of an attention part and a convolution part.

A [LightweightMLABlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) mixes the features of all positions. A [MobileBottleneckBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) then mixes the features of adjacent positions. Both parts add their input to their output, so the block keeps the shape of its input.

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.nodes.backbones.efficientvit.blocks import (
...     EfficientViTBlock,
... )
>>> block = EfficientViTBlock(32, head_dim=8)
>>> block(torch.zeros(1, 32, 8, 8)).shape
torch.Size([1, 32, 8, 8])
```

#### Methods

##### init

```python
def __init__(n_channels: int, attention_ratio: float = 1.0, head_dim: int = 32, expansion_factor: float = 4.0, aggregation_scales: tuple[int, ...] = (5)):
```

Build the attention part and the convolution part.

The attention part is a
[LightweightMLABlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md)
with a batch norm only after its projection. The convolution part is a
[MobileBottleneckBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md)
with a `3x3` kernel. Its expansion and depthwise layers have a bias term and a `torch.nn.Hardswish` activation. Only its
projection layer has a batch norm.

Parameters

 * `n_channels` (`int`): Number of input and output channels.
 * `attention_ratio` (`float`): Factor for the number of attention heads. The attention part has `int(n_channels // head_dim *
   attention_ratio)` heads. The number of heads must be at least `1`. With `0` heads and at least one aggregation scale,
   `torch.nn.Conv2d` raises `ValueError`. Defaults to `1.0`.
 * `head_dim` (`int`): Number of channels of the query, the key, and the value of each attention head. Defaults to `32`.
 * `expansion_factor` (`float`): Channel expansion factor of the convolution part. Defaults to `4.0`.
 * `aggregation_scales` (`tuple[int, ...]`): Kernel size of the depthwise convolution of each multi-scale aggregation branch of
   the attention part. The values must be odd. Defaults to `(5,)`.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Apply the attention part and then the convolution part.

Parameters

 * `x` (`Tensor`): Input of shape `[B, n_channels, H, W]`.

Returns

 * `Tensor`: Output of shape `[B, n_channels, H, W]`.

#### Attributes

##### attention_module

##### feature_module

### LightweightMLABlock

Lightweight multi-scale linear attention block of EfficientViT.

A `1x1`
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
computes the queries, the keys, and the values of all heads as one `qkv` tensor. Each aggregation branch runs a depthwise
convolution with a kernel size from `scale_factors` and a grouped `1x1` convolution on that tensor. The block concatenates the
`qkv` tensor and the branch outputs. Each head of each scale then computes its own attention. For the query Q, the key K, and the
value V of one head, the output at the position j is:

Oj = (∑iVi ϕ(Ki)⊤ϕ(Qj))/(∑iϕ(Ki)⊤ϕ(Qj) + ϵ)

The sums run over all positions i. ϕ is `kernel_activation`. A `1x1`
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
projects the attention outputs of all scales to `output_channels`. When `use_residual` is `True`, the block adds the input to the
projection output.

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.nodes.backbones.efficientvit.blocks import (
...     LightweightMLABlock,
... )
>>> block = LightweightMLABlock(16, 24, use_residual=False)
>>> block.qkv_layer.conv.out_channels
48
>>> block(torch.zeros(1, 16, 4, 4)).shape
torch.Size([1, 24, 4, 4])
```

#### Methods

##### init

```python
def __init__(input_channels: int, output_channels: int, n_heads: int | None = None, head_ratio: float = 1.0, dimension: int = 8,
use_bias: list[bool] | None = None, use_norm: list[bool] | None = None, activations: list[nn.Module] | None = None, scale_factors:
tuple[int, ...] = (5), epsilon: float = 1e-15, use_residual: bool = True, kernel_activation: nn.Module | None = None):
```

Build the `qkv` layer, the aggregation branches, and the projection.

Parameters

 * `input_channels` (`int`): Number of input channels.
 * `output_channels` (`int`): Number of output channels. It must be equal to `input_channels` when `use_residual` is `True`.
 * `n_heads` (`int | None`): Number of attention heads. `None` or `0` selects `int(input_channels // dimension * head_ratio)`. The number of heads must be at least `1`. With `0` heads and at least one aggregation branch, `torch.nn.Conv2d` raises `ValueError`.
 * `head_ratio` (`float`): Factor for the number of heads when `n_heads` is `None` or `0`. Defaults to `1.0`.
 * `dimension` (`int`): Number of channels of the query, the key, and the value of each head. Defaults to `8`.
 * `use_bias` (`list[bool] | None`): Whether the layers have a bias term, as two values. The first value applies to the `qkv` layer and to the convolutions of the aggregation branches. The second value applies to the projection. `None` selects `[False, False]`.
 * `use_norm` (`list[bool] | None`): Whether the `qkv` layer and the projection have a batch norm, as two values. `None` selects `[False, True]`.
 * `activations` (`list[nn.Module] | None`): The activations after the `qkv` layer and after the projection. `None` selects two `torch.nn.Identity` modules.
 * `scale_factors` (`tuple[int, ...]`): Kernel size of the depthwise convolution of each aggregation branch. The block has one branch for each value. The values must be odd. An even value changes the height and the width of the branch output, and [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) fails. Defaults to `(5,)`.
 * `epsilon` (`float`): Value that the attention adds to its denominator. Defaults to `1e-15`.
 * `use_residual` (`bool`): Whether [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) adds the input to the output. Defaults to `True`.
 * `kernel_activation` (`nn.Module | None`): The kernel function ϕ that runs on the queries and the keys. `None` selects `torch.nn.ReLU`.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Compute the multi-scale attention and project the result.

The method runs the `qkv` layer and the aggregation branches, and concatenates their outputs. It calls [linear_attention](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) when `H * W` is larger than `dimension`. Otherwise, it calls [quadratic_attention](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md). It converts the result of [linear_attention](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) back to the type of the `qkv` tensor. The projection then maps the result to `output_channels`. When `use_residual` is `True`, the method adds the input to the projection output in place.

Parameters

 * `x` (`Tensor`): Input of shape `[B, input_channels, H, W]`.

Returns

 * `Tensor`: Output of shape `[B, output_channels, H, W]`.

##### linear_attention

```python
def linear_attention(qkv_tensor: Tensor) -> Tensor:
```

Compute the attention of each head at a linear cost.

The method splits the channels of `qkv_tensor` into groups of `3 * dimension` channels. Each group holds the query, the key, and the value of one head. The method computes the attention formula that [LightweightMLABlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) shows. It pads V with a row of ones and multiplies the result with ϕ(K)⊤ first. The extra row then gives the denominator. Thus the cost grows linearly with `H * W`.

The method runs with CUDA autocast disabled. A `float16` input runs in `float32`. A `bfloat16` input runs in `bfloat16`, and the method converts the result to `float32` before the division.

> **Example**
> The linear and the quadratic attention give the same values.

```pycon
>>> import torch
>>> from luxonis_train.nodes.backbones.efficientvit.blocks import (
...     LightweightMLABlock,
... )
>>> block = LightweightMLABlock(8, 8, dimension=4)
>>> qkv = torch.arange(96.0).reshape(1, 24, 2, 2) / 96
>>> linear = block.linear_attention(qkv)
>>> linear.shape
torch.Size([1, 8, 2, 2])
>>> torch.allclose(linear, block.quadratic_attention(qkv))
True
```

Parameters

 * `qkv_tensor` (`Tensor`): Queries, keys, and values of shape `[B, G * 3 * dimension, H, W]`, where `G` is the number of groups.

Returns

 * `Tensor`: Attention output of shape `[B, G * dimension, H, W]`. It is `float32` when the input is `float16` or `bfloat16`.

##### quadratic_attention

```python
def quadratic_attention(qkv_tensor: Tensor) -> Tensor:
```

Compute the attention of each head with a full attention map.

The method splits the channels of `qkv_tensor` into groups as [linear_attention](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md) does. It builds the attention map Aij = ϕ(Ki)⊤ϕ(Qj) of shape `[B, G, H * W, H * W]`. It divides each column j of the map by ∑iAij + ϵ. Then it returns Oj = ∑iViAij with the divided map. The values are the same as the values of [linear_attention](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md), but the cost grows with the square of `H * W`.

The method runs with CUDA autocast disabled. It normalizes a `float16` or `bfloat16` attention map in `float32` and then converts the map back to the input type.

Parameters

 * `qkv_tensor` (`Tensor`): Queries, keys, and values of shape `[B, G * 3 * dimension, H, W]`, where `G` is the number of groups.

Returns

 * `Tensor`: Attention output of shape `[B, G * dimension, H, W]`, with the type of the input.

#### Attributes

##### kernel_activation

##### multi_scale_aggregators

##### projection_layer

##### qkv_layer

### MobileBottleneckBlock

Mobile inverted bottleneck block.

The block runs three [ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md) layers. A `1x1` convolution expands the channels. A depthwise convolution with one group for each hidden channel follows. A second `1x1` convolution projects the channels to `out_channels`. Each of the three layers has its own bias, batch norm, and activation settings.

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.nodes.backbones.efficientvit.blocks import (
...     MobileBottleneckBlock,
... )
>>> block = MobileBottleneckBlock(8, 16, stride=2, expand_ratio=4)
>>> block.depthwise_conv.conv.in_channels
32
>>> block(torch.zeros(1, 8, 16, 16)).shape
torch.Size([1, 16, 8, 8])
```

#### Methods

##### init

```python
def __init__(in_channels: int, out_channels: int, kernel_size: int = 3, stride: int = 1, expand_ratio: float = 6, use_bias: list[bool] | None = None, use_norm: list[bool] | None = None, activation: list[nn.Module] | None = None, use_residual: bool = False):
```

Build the expansion, the depthwise, and the projection layers.

Parameters

 * `in_channels` (`int`): Number of input channels.
 * `out_channels` (`int`): Number of output channels.
 * `kernel_size` (`int`): Kernel size of the depthwise convolution. The padding is `kernel_size // 2`. Defaults to `3`.
 * `stride` (`int`): Stride of the depthwise convolution. Defaults to `1`.
 * `expand_ratio` (`float`): Channel expansion factor. The hidden layers have `round(in_channels * expand_ratio)` channels.
   Defaults to `6`.
 * `use_bias` (`list[bool] | None`): Whether each layer has a bias term, as three values for the expansion, the depthwise, and the
   projection layers. `None` selects `[False, False, False]`.
 * `use_norm` (`list[bool] | None`): Whether each layer has a batch norm, in the same order. `None` selects `[True, True, True]`.
 * `activation` (`list[nn.Module] | None`): The activation after each layer, in the same order. `None` selects `torch.nn.ReLU6`,
   `torch.nn.ReLU6`, and `torch.nn.Identity`.
 * `use_residual` (`bool`): Whether
   [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/efficientvit/blocks.md)
   adds the input to the output. The input and the output must then have the same shape. Defaults to `False`.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Apply the expansion, the depthwise, and the projection layers.

Parameters

 * `x` (`Tensor`): Input of shape `[B, in_channels, H, W]`.

Returns

 * `Tensor`: Output of shape `[B, out_channels, H', W']`. For an odd kernel size, H’ = ⌈H ⁄ s⌉ and W’ = ⌈W ⁄ s⌉, where s is
   `stride`. When `use_residual` is `True`, the method adds the input to the result.

#### Attributes

##### depthwise_conv

##### expand_conv

##### project_conv
