# contextspatial

Python API: `luxonis_train.nodes.backbones.contextspatial`

The BiSeNet V1 context-spatial backbone and its two paths.

[SpatialPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md)
keeps the spatial detail of the image.
[ContextPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md)
runs another backbone for a large receptive field.
[ContextSpatial](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md)
fuses the outputs of the two paths.

## Classes

### ContextPath

Context path of
[ContextSpatial](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md)
on top of a backbone.

The path reads the last two feature maps of the backbone. Below, `f16` and `f32` are these maps, at the strides 16 and 32. The
path runs these steps:

 1. An
    [AttentionRefinementBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
    maps each of `f16` and `f32` to 128 channels.
 2. A global average pool and a `1x1`
    [ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
    turn `f32` into one context vector for each image. The path resizes the vector to the size of `f32` and adds it to the refined
    `f32`.
 3. The path upsamples the sum by `2`, applies a `3x3`
    [ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md),
    and adds the result to the refined `f16`.
 4. The path upsamples the second sum by `2` and applies a second `3x3`
    [ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md).

The
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
layers of the steps 2 to 4 use batch norm and ReLU. All resizes are bilinear with aligned corners.

The attention blocks and the global context layer depend on the channel counts of the backbone. Thus the path creates them in the
first
[forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md)
call. Before that call, the state dictionary of the path does not hold them.

#### Methods

##### init

```python
def __init__(backbone: nn.Module):
```

Store the backbone and build the fixed layers.

The fixed layers are the two upsampling layers and the two `3x3` refinement blocks.
[forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md)
builds the other layers.

Parameters

 * `backbone` (`nn.Module`): Module that returns a sequence of feature maps. The path reads the last two entries. They must have
   the strides 16 and 32.

##### forward

```python
def forward(x: Tensor) -> tuple[Tensor, Tensor]:
```

Refine the last two backbone maps and merge them.

The first call also creates the two attention blocks and the global context layer. It reads their input channels from the backbone
outputs. The new layers use the default device and dtype of PyTorch, not those of `x`. They start in the training state, also when
the path is in the eval state.

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.nodes import MobileNetV2
>>> path = ContextPath(MobileNetV2())
>>> hasattr(path, "arm16")
False
>>> [tuple(t.shape) for t in path(torch.zeros(2, 3, 64, 64))]
[(2, 128, 8, 8), (2, 128, 4, 4)]
>>> hasattr(path, "arm16")
True
```

Parameters

 * `x` (`Tensor`): Input of the backbone, of shape `[B, C, H, W]`. The stride-32 map, upsampled by `2`, must have the size of the stride-16 map. All multiples of `32` for `H` and `W` meet this condition. In the training state, `B` must be larger than `1`. The reason is that each batch norm after a global pooling sees one value per image.

Returns

 * `tuple[Tensor, Tensor]`: Two maps with 128 channels. The first is the final merged map, at twice the size of the stride-16 map. The second is the refined stride-32 branch after its upsampling and its `3x3` [ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md), at the size of the stride-16 map. When `H` and `W` are multiples of `32`, the shapes are `[B, 128, H / 8, W / 8]` and `[B, 128, H / 16, W / 16]`.

#### Attributes

##### arm16

##### arm32

##### backbone

##### global_context

##### refine16

##### refine32

##### up16

##### up32

### ContextSpatial

BiSeNet V1 backbone that fuses a spatial and a context path.

The node runs two paths on the same image and fuses their outputs:

 * [SpatialPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md) keeps the detail. It gives 128 channels at 1/8 of the input size.
 * [ContextPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md) runs a context backbone and refines its last two feature maps. It also gives 128 channels at 1/8 of the input size.
 * A [FeatureFusionBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md) fuses the two maps into 256 channels.

[BiSeNetHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/bisenet_head.md) can read the fused map.

 * `Inputs:`: * `inputs` (`Tensor`): [B, 3, H, W]
 * `Outputs:`: * `features` (`list[Tensor]`): [B, 256, H ⁄ 8, W ⁄ 8]

> **References**
> * Source: Reimplemented from [BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation](https://arxiv.org/abs/1808.00897).
 * License: Apache-2.0 (this project)

> **Notes**
> The node builds [SpatialPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md) for 3 input channels. The last two feature maps of the context backbone must have the strides 16 and 32. Then all multiples of 32 for H and W work. Other sizes can give feature maps of different sizes, and PyTorch then raises `RuntimeError`. In the training state, a batch must hold more than one image. The reason is that each batch norm after a global pooling sees one value per image. [ContextPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md) creates its attention refinement blocks and its global context layer in the first forward call. Before that call, the state dictionary of the node does not hold them.

 * `Variants:`: None. Configure the node through `params`.

> **See Also**
> [The BiseNetv1 repository](https://github.com/taveraantonio/BiseNetv1)

> **Example**
> A node entry in the `model.nodes` section of a config:

```yaml
- name: ContextSpatial
```

 * `Compatible with:`: * Attach index: `-1`, the last output of the input node

#### Methods

##### init

```python
def __init__(context_backbone: str | nn.Module = 'MobileNetV2', backbone_kwargs: Kwargs | None = None, **kwargs):
```

Build the two paths and the feature fusion block.

Parameters

 * `context_backbone` (`str | nn.Module`): The backbone of [ContextPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md). A string names a node in [luxonis_train.registry.NODES](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/registry.md). The constructor builds that node with `backbone_kwargs` and `kwargs`. An unknown name makes the registry raise `KeyError`. A module goes to [ContextPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md) as it is. The backbone must return a sequence of feature maps, and the last two entries must have the strides 16 and 32. Defaults to `"MobileNetV2"`.
 * `backbone_kwargs` (`Kwargs | None`): Keyword arguments for the backbone node. The constructor reads them only when `context_backbone` is a string. It merges `kwargs` into this dictionary, and a key in `kwargs` replaces the same key here. A dictionary that is not empty changes in place, so the caller sees the merged keys. `None` and an empty dictionary start from a new empty dictionary.
 * `**kwargs`: Keyword arguments forwarded to [BaseNode](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/base_node.md). A backbone built from a name gets them too.

##### forward

```python
def forward(inputs: Tensor) -> list[Tensor]:
```

Fuse the spatial detail with the refined context features.

[SpatialPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md) and [ContextPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md) both give 128 channels at 1/8 of the input size. [FeatureFusionBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md) concatenates the two maps and fuses them into 256 channels. The node ignores the second output of [ContextPath](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md).

> **Example**
> ```pycon
>>> import torch
>>> from luxonis_train.nodes import ContextSpatial
>>> node = ContextSpatial()
>>> [tuple(t.shape) for t in node(torch.zeros(2, 3, 64, 64))]
[(2, 256, 8, 8)]
```

A size that is not a multiple of 32 can work too:

```pycon
>>> [tuple(t.shape) for t in node(torch.zeros(2, 3, 62, 94))]
[(2, 256, 8, 12)]
```

Parameters

 * `inputs` (`Tensor`): Image batch of shape `[B, 3, H, W]`. All multiples of `32` for `H` and `W` work. Other sizes can give
   feature maps of different sizes, and PyTorch then raises `RuntimeError`. In the training state, `B` must be larger than `1`.
   Otherwise, a batch norm raises `ValueError`.

Returns

 * `list[Tensor]`: A list with one fused feature map of shape `[B, 256, ceil(H / 8), ceil(W / 8)]`.

#### Attributes

##### context_path

##### ffm

##### spatial_path

### SpatialPath

Spatial path of
[ContextSpatial](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/contextspatial.md)
that keeps the image detail.

Three strided
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
layers reduce the input to 1/8 of its size. The first layer has a `7x7` kernel, and the other two have a `3x3` kernel. Each of
them has stride `2` and 64 output channels. A `1x1`
[ConvBlock](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/blocks/blocks.md)
then maps the result to `out_channels`. Every layer uses batch norm and ReLU.

> **Example**
> ```pycon
>>> import torch
>>> path = SpatialPath(3, 128)
>>> path(torch.zeros(1, 3, 64, 64)).shape
torch.Size([1, 128, 8, 8])
```

#### Methods

##### init

```python
def __init__(in_channels: int, out_channels: int):
```

Initialize the four convolutions.

Parameters

 * `in_channels` (`int`): Number of input channels.
 * `out_channels` (`int`): Number of output channels.

##### forward

```python
def forward(x: Tensor) -> Tensor:
```

Reduce `x` to 1/8 of its size and map its channels.

Parameters

 * `x` (`Tensor`): Input of shape `[B, in_channels, H, W]`.

Returns

 * `Tensor`: Output of shape `[B, out_channels, ceil(H / 8), ceil(W / 8)]`.

#### Attributes

##### conv_1x1

##### conv_3x3_1

##### conv_3x3_2

##### conv_7x7
