# blocks

Python API: `luxonis_train.nodes.backbones.micronet.blocks`

The blocks of the MicroNet backbone.

## Classes

### ChannelShuffle

Channel shuffle that interleaves the channels of the groups.

The module splits the channels into `groups` groups of equal size. The output takes the first channel of each group, then the
second channel of each group, and so on. A grouped convolution after the shuffle then reads channels from all groups. With
`groups` equal to `1` or to the number of channels, the order does not change.

> **Example**
> ```pycon
>>> import torch
>>> x = torch.arange(6.0).view(1, 6, 1, 1)
>>> ChannelShuffle(3)(x).flatten().tolist()
[0.0, 2.0, 4.0, 1.0, 3.0, 5.0]
```

#### Methods

##### init

```python
def __init__(groups: int):
```

Store the number of groups.

Parameters

 * `groups` (`int`): The number of groups. The number of input channels must be a multiple of it.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Interleave the channels of the groups.

Parameters

 * `x` (`Tensor`): The input of shape `[B, C, H, W]`. `C` must be a multiple of `groups`. Otherwise, the reshape raises `RuntimeError`.

Returns

 * `Tensor`: The input with its channels in the new order, of the same shape.

### DepthSpatialSepConv

Factorized depthwise convolution that expands the channels.

A k×1 convolution with `in_channels` groups multiplies the channels by `expand[0]` and applies the stride along the height. A 1×k convolution with one group for each of its input channels multiplies the channels by `expand[1]` and applies the stride along the width. A batch norm follows each convolution. The convolutions have no bias, and the block has no activation.

> **Example**
> ```pycon
>>> import torch
>>> conv = DepthSpatialSepConv(4, (2, 3), kernel_size=3, stride=2)
>>> conv(torch.zeros(1, 4, 9, 9)).shape
torch.Size([1, 24, 5, 5])
```

#### Methods

##### init

```python
def __init__(in_channels: int, expand: tuple[int, int], kernel_size: int, stride: int):
```

Initialize the two depthwise convolutions.

Parameters

 * `in_channels` (`int`): The number of input channels.
 * `expand` (`tuple[int, int]`): The channel multipliers of the k×1 and the 1×k convolution. The block has `in_channels *
   expand[0] * expand[1]` output channels.
 * `kernel_size` (`int`): The kernel size k. The padding is `kernel_size // 2`.
 * `stride` (`int`): The stride along the height and the width.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Apply the two convolutions and their batch norms.

Parameters

 * `x` (`Tensor`): The input of shape `[B, in_channels, H, W]`.

Returns

 * `Tensor`: The output of shape `[B, in_channels * expand[0] * expand[1], H', W']`. For an odd `kernel_size`, `H'` is `ceil(H /
   stride)` and `W'` is `ceil(W / stride)`.

#### Attributes

##### conv

### DYShiftMax

Dynamic Shift-Max activation of MicroNet.

The activation mixes each channel with one channel of the next group. The shifted input x̃ takes channel `c + 1` of group `g + 1`
for channel `c` of group `g`. Both indices wrap around. With 8 channels in 2 groups, channel `0` of x̃ takes channel `5`, the
second channel of the second group. With two branches, the output is

y = max(a1x + b1x̃, a2x + b2x̃)

With one branch, the output is y = a1x + b1x̃.

A squeeze network computes the coefficients from the input. It averages each channel over the height and the width. Two linear
layers with a `ReLU` between them and a hard sigmoid follow. The module maps the result from `[0, 1]` to `[-2, 2]` and adds the
offsets `init_a` and `init_b`. Each coefficient has `out_channels` values for each sample.

> **Example**
> ```pycon
>>> import torch
>>> act = DYShiftMax(8, 8, groups=2)
>>> act(torch.ones(2, 8, 4, 4)).shape
torch.Size([2, 8, 4, 4])
```

#### Methods

##### init

```python
def __init__(in_channels: int, out_channels: int, init_a: tuple[float, float] = (0.0, 0.0), init_b: tuple[float, float] = (0.0,
0.0), use_relu: bool = True, groups: int = 6, reduction: int = 4, expansion: bool = False):
```

Initialize the squeeze network and the channel shift.

Parameters

 * `in_channels` (`int`): The number of input channels. It must be a multiple of the number of groups.
 * `out_channels` (`int`): The number of channels of each coefficient. [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md) needs it equal to `in_channels` or to `1`. With `1`, all channels of a sample share each coefficient.
 * `init_a` (`tuple[float, float]`): The offsets for a1 and a2. [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md) adds `init_a[0]` to a1. It does not read `init_a[1]`, because it adds `init_b[1]` to a2.
 * `init_b` (`tuple[float, float]`): The offsets for b1 and b2. `init_b[1]` also goes to a2.
 * `use_relu` (`bool`): `True` selects two branches and their maximum, a dynamic form of `ReLU`. `False` selects one branch without a maximum.
 * `groups` (`int`): The number of channel groups for the shift. With `1`, the shift moves the channels by one position.
 * `reduction` (`int`): The divisor of `in_channels` for the hidden layer of the squeeze network. The module rounds `in_channels // reduction` to the nearest multiple of `4`, with a minimum of `4`. When the rounding goes more than 10% down, it adds `4`.
 * `expansion` (`bool`): When `True` and `groups` is not `1`, the number of groups is `in_channels // groups`. Then `groups` is the number of channels in each group.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Apply the activation with the coefficients of the input.

Parameters

 * `x` (`Tensor`): The input of shape `[B, in_channels, H, W]`.

Returns

 * `Tensor`: The activated input, of the same shape.

#### Attributes

##### avg_pool

##### fc

### MicroBlock

The basic block of MicroNet.

The block expands the input to `in_channels * expand_ratio[0] * expand_ratio[1]` channels. `groups_1` and `groups_2` select one of three layouts:

 * A lite block, when `groups_1[0]` is `0`. A [DepthSpatialSepConv](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md) expands the channels and applies the stride. A grouped 1×1 convolution projects them to `out_channels`.
 * A transition block, when `groups_2[1]` is `0` and `groups_1[0]` is not `0`. A grouped 1×1 convolution expands the channels. The block has no depthwise convolution and no projection. Its output keeps the expanded channels.
 * A full block, in all other cases. A grouped 1×1 convolution expands the channels. A [DepthSpatialSepConv](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md) without expansion applies the stride. A second grouped 1×1 convolution projects the channels to `out_channels`.

Each convolution has a batch norm and no bias. A [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md), a `ReLU6`, or no activation follows each 1×1 convolution and each [DepthSpatialSepConv](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md), as `dy_shift` selects. [ChannelShuffle](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md) layers mix the channels of the groups. The block adds its input to its output when `stride` is `1` and `out_channels` is equal to `in_channels`.

> **Example**
> ```pycon
>>> import torch
>>> x = torch.zeros(2, 8, 16, 16)
>>> lite = MicroBlock(
...     8, 16, stride=2, groups_1=(0, 8), groups_2=(4, 4)
... )
>>> lite(x).shape
torch.Size([2, 16, 8, 8])
```

A transition block ignores `out_channels` and `stride` in its layers:

```pycon
>>> transition = MicroBlock(
...     8,
...     16,
...     stride=2,
...     expand_ratio=(1, 6),
...     groups_1=(4, 4),
...     groups_2=(0, 0),
... )
>>> transition(x).shape
torch.Size([2, 48, 16, 16])
```

#### Methods

##### init

```python
def __init__(in_channels: int, out_channels: int, kernel_size: int = 3, stride: int = 1, expand_ratio: tuple[int, int] = (2, 2), groups_1: tuple[int, int] = (0, 6), groups_2: tuple[int, int] = (1, 1), dy_shift: tuple[int, int, int] = (2, 0, 1), reduction_factor: int = 1, init_a: tuple[float, float] = (1.0, 1.0), init_b: tuple[float, float] = (0.0, 0.0)):
```

Build the layers of the lite, transition, or full layout.

Parameters

 * `in_channels` (`int`): The number of input channels.
 * `out_channels` (`int`): The number of output channels of a lite or a full block. The layers of a transition block do not read
   it. The residual check reads it in all layouts.
 * `kernel_size` (`int`): The kernel size k of the
   [DepthSpatialSepConv](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md),
   which is a k×1 and a 1×k convolution. A transition block ignores it.
 * `stride` (`int`): The stride of the
   [DepthSpatialSepConv](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md).
   The layers of a transition block ignore it.
 * `expand_ratio` (`tuple[int, int]`): The two channel multipliers of the expansion. A lite block applies one multiplier in each
   half of its
   [DepthSpatialSepConv](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md).
   The other layouts apply the product in the expansion 1×1 convolution.
 * `groups_1` (`tuple[int, int]`): The first value is the number of groups of the expansion 1×1 convolution. `0` selects a lite
   block. The second value sets the groups of the
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md)
   layers before the projection and of the
   [ChannelShuffle](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md)
   after the first activation. In the
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md)
   after the depthwise convolution of a full block, the groups are the expanded channels divided by the value, when the value is
   not `1`. In a transition block, the value sets the groups of its only
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md).
 * `groups_2` (`tuple[int, int]`): The first value is the number of groups of the projection 1×1 convolution. The second value
   sets the groups of the last
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md)
   and of the
   [ChannelShuffle](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md)
   after it. `0` selects a transition block when `groups_1[0]` is not `0`.
 * `dy_shift` (`tuple[int, int, int]`): The activations after the expansion convolution, after the depthwise convolution, and
   after the projection. In the first two positions, a positive value selects
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md),
   and other values select `ReLU6`. `2` selects two branches, and another positive value selects one branch. In the last position,
   a positive value selects a one-branch
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md),
   and other values select no activation. A lite block ignores the first value. A transition block reads only the last value, for
   the activation after its expansion. In a lite or a full block, values other than `0` also add a
   [ChannelShuffle](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md)
   after an activation. In a lite block, the second value adds one with `C // 2` groups, where `C` is the number of expanded
   channels. In a full block, the first two values share one after the depthwise activation. It has `C // 4` groups when both
   values are not `0`, and `C // 2` groups when only one is not `0`. The last value adds one with `out_channels // 2` groups. In a
   lite block, it does so only when `out_channels` is even.
 * `reduction_factor` (`int`): The reduction of the squeeze network of
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md).
   The activations use a `reduction` of `8 * reduction_factor`. The last activation of a lite block uses `4 * reduction_factor`.
   The last activation of a full block also does, when `out_channels` is smaller than the expanded channels.
 * `init_a` (`tuple[float, float]`): The offsets for the weights of the input in
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md).
   The activations after the expansion and after the depthwise convolution use them. The last activation, and the activation of a
   transition block, use `(1.0, 0.0)` instead.
 * `init_b` (`tuple[float, float]`): The offsets for the weights of the shifted input in
   [DYShiftMax](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md).
   The same activations as for `init_a` use them. The others use `(0.0, 0.0)`.

##### forward

```python
def forward(inputs: Tensor) -> Tensor:
```

Run the layers, and add the input for a residual connection.

Parameters

 * `inputs` (`Tensor`): The input of shape `[B, in_channels, H, W]`.

Returns

 * `Tensor`: The output of shape `[B, C, H', W']`. `C` is `out_channels` for a lite or a full block, and the expanded number of
   channels for a transition block. For an odd `kernel_size`, `H'` is `ceil(H / stride)` and `W'` is `ceil(W / stride)`. A
   transition block keeps `H` and `W`. With the residual connection, the output is the sum of the layer output and `inputs`. For a
   transition block, this sum raises `RuntimeError` unless the expanded number of channels is equal to `in_channels`.

#### Attributes

##### layers

### SpatialSepConvSF

Spatially separable convolution with a channel shuffle.

A k×1 convolution maps the input to `outs[0]` channels and applies the stride along the height. A 1×k convolution with `outs[0]`
groups multiplies the channels by `outs[1]` and applies the stride along the width. A batch norm follows each convolution. A
[ChannelShuffle](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md)
with `outs[0]` groups ends the block. The convolutions have no bias, and the block has no activation.

> **Example**
> ```pycon
>>> import torch
>>> conv = SpatialSepConvSF(3, (4, 2), kernel_size=3, stride=2)
>>> conv(torch.zeros(1, 3, 32, 32)).shape
torch.Size([1, 8, 16, 16])
```

#### Methods

##### init

```python
def __init__(in_channels: int, outs: tuple[int, int], kernel_size: int, stride: int):
```

Initialize the two convolutions and the channel shuffle.

Parameters

 * `in_channels` (`int`): The number of input channels.
 * `outs` (`tuple[int, int]`): The channel layout. The first value is the number of output channels of the k×1 convolution. It is also the number of groups of the 1×k convolution and of the shuffle. The second value is the channel multiplier of the 1×k convolution.
 * `kernel_size` (`int`): The kernel size k. The padding is `kernel_size // 2`.
 * `stride` (`int`): The stride along the height and the width.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Apply the convolutions, the batch norms, and the shuffle.

Parameters

 * `x` (`Tensor`): The input of shape `[B, in_channels, H, W]`.

Returns

 * `Tensor`: The output of shape `[B, outs[0] * outs[1], H', W']`. For an odd `kernel_size`, `H'` is `ceil(H / stride)` and `W'` is `ceil(W / stride)`.

#### Attributes

##### conv

### Stem

Stem of MicroNet.

The stem is a [SpatialSepConvSF](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md) with a kernel size of `3`, followed by an in-place `ReLU6`.

> **Example**
> ```pycon
>>> import torch
>>> stem = Stem(3, stride=2, outs=(3, 2))
>>> stem(torch.zeros(1, 3, 32, 32)).shape
torch.Size([1, 6, 16, 16])
```

#### Methods

##### init

```python
def __init__(in_channels: int, stride: int, outs: tuple[int, int] = (4, 4)):
```

Initialize the convolution and the activation.

Parameters

 * `in_channels` (`int`): The number of input channels.
 * `stride` (`int`): The stride along the height and the width.
 * `outs` (`tuple[int, int]`): The channel layout of the
   [SpatialSepConvSF](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/micronet/blocks.md).
   The first value is the number of output channels of the 3×1 convolution and the number of groups after it. The second value is
   the channel multiplier. The stem has `outs[0] * outs[1]` output channels.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Apply the convolution and `ReLU6`.

Parameters

 * `x` (`Tensor`): The input of shape `[B, in_channels, H, W]`.

Returns

 * `Tensor`: The output of shape `[B, outs[0] * outs[1], ceil(H / stride), ceil(W / stride)]`, with values in `[0, 6]`.

#### Attributes

##### stem
