# rope_position_encoding

Python API: `luxonis_train.nodes.backbones.dinov3.rope_position_encoding`

The rotary position embedding of DINOv3, changed to export to ONNX.

## Classes

### RopePositionEmbedding

Axial rotary position embedding of DINOv3, without learnable weights.

The module computes the sine and the cosine tables that the DINOv3 attention uses to rotate the query and the key vectors. Each
patch of an `H x W` grid gets the coordinates of its center.
[forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/dinov3/rope_position_encoding.md)
divides them as `normalize_coords` selects and maps each coordinate c to 2c − 1. With `"separate"` or `"max"`, the result is in
`[-1, 1]`. With `"min"`, the coordinates of the longer axis can be above `1`. Each angle depends on one axis only, so the axes do
not mix. For the coordinate c and the period p, the angle is 2π**c ⁄ p.

The periods come from one of two settings. D is the head dimension, `embed_dim // num_heads`.

 * `base`: pi = base2i ⁄ (D ⁄ 2) for i = 0, …, D ⁄ 4 − 1.
 * `min_period` and `max_period`: D ⁄ 4 periods in a geometric sequence from `min_period` to `max_period`.

The code comes from DINOv3.
[forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/dinov3/rope_position_encoding.md)
uses `repeat` instead of `tile`, because `tile` does not export to ONNX.

> **Example**
> ```pycon
>>> rope = RopePositionEmbedding(64, num_heads=4)
>>> [round(p, 2) for p in rope.periods.tolist()]
[1.0, 3.16, 10.0, 31.62]
>>> sin, cos = rope(H=2, W=3)
>>> sin.shape, cos.shape
(torch.Size([6, 16]), torch.Size([6, 16]))
```

The second setting of the periods:

```pycon
>>> rope = RopePositionEmbedding(
...     64, num_heads=4, base=None, min_period=0.5, max_period=10.0
... )
>>> [round(p, 2) for p in rope.periods.tolist()]
[0.5, 1.36, 3.68, 10.0]
```

#### Methods

##### init

```python
def __init__(embed_dim: int, *, num_heads: int, base: float | None = 100.0, min_period: float | None = None, max_period: float |
None = None, normalize_coords: Literal['min', 'max', 'separate'] = 'separate', shift_coords: float | None = None, jitter_coords:
float | None = None, rescale_coords: float | None = None, dtype: torch.dtype | None = None, device: torch.device | None = None):
```

Store the settings and compute the periods.

Give either `base`, or `base=None` with both `min_period` and `max_period`. When `base` is set, the constructor ignores a single `min_period` or `max_period`.

Parameters

 * `embed_dim` (`int`): The embedding dimension of the transformer. It must be a multiple of `4 * num_heads`.
 * `num_heads` (`int`): The number of attention heads.
 * `base` (`float | None`): The base of the periods. `None` selects the `min_period` and `max_period` setting.
 * `min_period` (`float | None`): The smallest period. The module uses it only when `base` is `None`.
 * `max_period` (`float | None`): The largest period. The module uses it only when `base` is `None`.
 * `normalize_coords` (`Literal['min', 'max', 'separate']`): The divisor of the patch coordinates. `"separate"` divides the rows by `H` and the columns by `W`. `"max"` divides both by `max(H, W)`, and `"min"` divides both by `min(H, W)`.
 * `shift_coords` (`float | None`): In training mode, [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/dinov3/rope_position_encoding.md) adds a random shift to each axis. Each axis gets its own shift, uniform in `[-shift_coords, shift_coords]`. `None` adds no shift.
 * `jitter_coords` (`float | None`): In training mode, [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/dinov3/rope_position_encoding.md) multiplies each axis by its own random factor. The factor is log-uniform in `[1 / jitter_coords, jitter_coords]`. `None` applies no jitter.
 * `rescale_coords` (`float | None`): In training mode, [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/dinov3/rope_position_encoding.md) multiplies both axes by one random factor. The factor is log-uniform in `[1 / rescale_coords, rescale_coords]`. `None` applies no rescale.
 * `dtype` (`torch.dtype | None`): The data type of the periods and the coordinates. `None` selects the default data type.
 * `device` (`torch.device | None`): The device of the periods buffer. [forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/backbones/dinov3/rope_position_encoding.md) computes on the device of that buffer.

Raises

 * `AssertionError`: When `embed_dim` is not a multiple of `4 * num_heads`.
 * `ValueError`: When `base` is `None` and one of the periods is `None`, or when `base` and both periods are set.

##### forward

```python
def forward(*, H: int, W: int) -> tuple[Tensor, Tensor]:
```

Compute the sine and the cosine tables for an `H x W` grid.

The patches follow row-major order. For each patch, the method computes D ⁄ 4 angles from the row coordinate and then D ⁄ 4 angles from the column coordinate. It appends a copy of these D ⁄ 2 angles, which gives D angles. In training mode, the method first shifts, jitters, and rescales the coordinates, as the constructor arguments enable.

Parameters

 * `H` (`int`): The number of patch rows.
 * `W` (`int`): The number of patch columns.

Returns

 * `tuple[Tensor, Tensor]`: The sine and the cosine of the angles, each of shape `[H * W, D]`, where `D` is the head dimension.

Raises

 * `ValueError`: When `normalize_coords` is not `"min"`, `"max"`, or `"separate"`.

#### Attributes

##### periods

The D ⁄ 4 periods, in a persistent buffer.
