# triple_lr_sgd

Python API: `luxonis_train.strategies.triple_lr_sgd`

An SGD strategy with a warmup and a decay of the learning rate.

The strategy splits the parameters into `BatchNorm2d` weights, other weights, and biases.

## Classes

### TripleLRSGDStrategy

SGD over three parameter groups, with a warmup and a decay.

The rules of the strategy split the parameters into three groups:

 * `triple_lr/batch_norm_weights`: the `weight` of each `BatchNorm2d`, without weight decay.
 * `triple_lr/weights`: every other parameter named `weight`, with `weight_decay`.
 * `triple_lr/biases`: every parameter named `bias`, without weight decay.

Every group uses SGD with `lr`, `momentum`, and `nesterov`. A parameter with another name goes to the default rule, which uses the
configs of
[get_base_configs](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/triple_lr_sgd.md).

A `LambdaLR` scheduler sets the learning rate of each group to lr⋅f(e) in epoch e, counted from `0`. Let r = lre ⁄ lr, and let E
be `trainer.epochs`. With `cosine_annealing`, the factor is:

f(e) = 1 + (r − 1)⋅(1 − cos(π**e ⁄ E))/(2)

Without `cosine_annealing`, the factor falls linearly:

f(e) = r + (1 − r)⋅max(1 − e ⁄ E, 0)

Both factors go from `1` in the first epoch to r in epoch E.

During the warmup,
[update_parameters](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/triple_lr_sgd.md)
sets the learning rate of the groups of the three rules on each step. The rate moves linearly from a start value to lr⋅f(e). The
bias group starts from `warmup_bias_lr`, and the other groups start from `0`.

> **Example**
> The `trainer` section of a config:

```yaml
trainer:
  training_strategy:
    name: TripleLRSGDStrategy
    params:
      lr: 0.02
      lre: 0.0002
      warmup_epochs: 3
      cosine_annealing: true
```

#### Methods

##### init

```python
def __init__(pl_module: lxt.LuxonisLightningModule, lr: float = 0.02, momentum: float = 0.937, weight_decay: float = 0.0005, nesterov: bool = True, warmup_epochs: int = 3, warmup_bias_lr: float = 0.1, warmup_momentum: float = 0.8, lre: float = 0.0002, cosine_annealing: bool = True):
```

Store the settings and compute the length of the warmup.

The number of batches in an epoch is `ceil(len(pl_module.core.loaders["train"]) / trainer.batch_size)`. The warmup lasts
`warmup_epochs` times that number of steps, rounded, and at least `100` steps.

Parameters

 * `pl_module` (`lxt.LuxonisLightningModule`): The module to train. The strategy reads its `cfg`, its `core.loaders`, and its
   `current_epoch`.
 * `lr` (`float`): The base learning rate of every group.
 * `momentum` (`float`): The SGD momentum of every group.
 * `weight_decay` (`float`): The weight decay of the `triple_lr/weights` group.
 * `nesterov` (`bool`): Whether SGD uses Nesterov momentum.
 * `warmup_epochs` (`int`): The length of the warmup, in epochs.
 * `warmup_bias_lr` (`float`): The learning rate of the bias group at the start of the warmup.
 * `warmup_momentum` (`float`): The strategy stores the value, but does not use it.
 * `lre` (`float`): The learning rate at the end of the training.
 * `cosine_annealing` (`bool`): Whether the learning rate factor follows a cosine curve. With `False`, it falls linearly.

##### get_base_configs

```python
def get_base_configs(self) -> tuple[OptimizerConfig, SchedulerConfig]:
```

Return the SGD config and the `LambdaLR` config.

Returns

 * `tuple[OptimizerConfig, SchedulerConfig]`: The SGD config with `lr`, `momentum`, and `nesterov`, without weight decay. The
   `LambdaLR` config, whose `lr_lambda` is the learning rate factor f that the class describes.

##### rules

```python
def rules(self) -> list[StrategyRule]:
```

Return the three SGD rules of the strategy.

The `BatchNorm2d` rule comes before the weight rule, so a `BatchNorm2d` weight goes to the batch-norm group. No rule sets a
scheduler, so every group uses the `LambdaLR` of
[get_base_configs](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/triple_lr_sgd.md).

Returns

 * `list[StrategyRule]`: The rules with the tags `BATCH_NORM_TAG`, `WEIGHT_TAG`, and `BIAS_TAG`, in this order. Only the
   `WEIGHT_TAG` rule sets `weight_decay`.

##### update_parameters

```python
def update_parameters(self):
```

Set the warmup learning rates of the groups.

[TrainingManager](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/training_manager.md)
calls the method after each backward pass. A call counts as one step, also with gradient accumulation. The step in the epoch is
the call count modulo the number of batches in an epoch. The global step adds `current_epoch` times that number.

While the global step is at most the warmup length, the method sets `lr` of each group of the three tags. At step `0`, the value
is the start value of the group. At the last warmup step, it is lr⋅f(e), with e equal to `current_epoch`. Between the two steps,
the value changes linearly. After the warmup, the method changes no group.

#### Attributes

##### BATCH_NORM_TAG

##### BIAS_TAG

##### WEIGHT_TAG
