# training_plan

Python API: `luxonis_train.lightning.training_plan`

The partition of the model parameters into optimizer groups.

The module turns the node `finetuning` entries, the rules of a training strategy, and the node freezing into one optimizer
configuration. It works in two phases:

 1. [resolve_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
    builds a
    [TrainingPlan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
    from the config and the nodes. The plan holds every parameter of every node in exactly one parameter group, frozen parameters
    included. The parameters of a legacy training strategy are the only exception. Each group belongs to one inner optimizer and
    its scheduler. The function does not create the optimizers or the schedulers of the plan.
 2. [build_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
    creates the inner optimizers and their member schedulers from the plan. With one inner optimizer, Lightning receives that
    optimizer and its scheduler directly, so the checkpoint of a plain config holds a plain optimizer state. With several inner
    optimizers, one
    [CompositeOptimizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md)
    and composite schedulers wrap them, so the model stays in the automatic optimization of Lightning.

The partition is static. When a node unfreezes, its parameters stay in their groups. Only `requires_grad` changes, and a torch
optimizer skips a parameter whose gradient is `None`. A resume from a checkpoint is therefore a plain `state_dict` round trip.

## Classes

### GroupHandle

The stable address of one parameter group.

The handle holds indices, not the group dictionary. `Optimizer.load_state_dict` replaces the group dictionaries when a checkpoint
loads. The partition is static, so the indices stay valid.

#### Attributes

##### group_index

The index of the group in
[InnerSpec.groups](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
and in the `param_groups` of the inner optimizer.

##### inner_index

The index of the inner optimizer in
[TrainingPlan.inners](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
and in
[TrainingPlanRuntime.inner_optimizers](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md).

### GroupSpec

One parameter group of the plan.

#### Attributes

##### name

The name of the group, unique in the whole plan. It is the label of the rule, followed by `/<node>` when
[resolve_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
gives a node its own group of the rule. A repeated name gets the suffix `-2`, `-3`, and so on.

##### node_names

The names of the nodes that have parameters in the group, in the order of the first claim.

##### options

The parameter-group options, from the `params` of the optimizer of the rule.

##### parameter_names

The name of each parameter, in the order of `parameters`. A name is the node name, a dot, and the dotted name of the parameter in
the node, such as `"backbone.stem.conv.weight"`.

##### parameters

The parameters of the group, in claim order.

### InnerSpec

One inner optimizer of the plan and its scheduler.

#### Attributes

##### groups

The parameter groups of the optimizer, in creation order.

##### optimizer_name

The class name of the optimizer in the `OPTIMIZERS` registry.

##### scheduler

The scheduler of the optimizer.

### OptimizerSpec

The optimizer of one rule.

Rules with the same optimizer name and the same scheduler
[SchedulerSpec.key](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
share one inner optimizer. The `params` do not affect that choice, because each group carries them as its own options.

#### Methods

##### from_config

```python
def from_config(config: OptimizerConfig) -> OptimizerSpec:
```

Create the specification from an optimizer config.

Parameters

 * `config` (`OptimizerConfig`): The optimizer config.

Returns

 * `OptimizerSpec`: The name of `config` and a shallow copy of its `params`.

Raises

 * `KeyError`: When the `OPTIMIZERS` registry has no optimizer with the name of `config`.

#### Attributes

##### name

The class name of the optimizer in the `OPTIMIZERS` registry.

##### params

The optimizer parameters, such as `lr`. Each parameter group of the rule receives them as its options.

### Rule

A rule that claims parameters into the groups of the plan.

[resolve_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
creates one rule for each node `finetuning` entry, one rule for each
[StrategyRule](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md),
and one default rule that claims all parameters that are left.

#### Attributes

##### label

The base name of the groups of the rule: `<node>/<index>` for a `finetuning` entry, `strategy/<tag>` for a strategy rule, and
`default` for the default rule.

##### optimizer

The optimizer of the claimed parameters.

##### scheduler

The scheduler of the claimed parameters.

##### selector

The predicate that decides which parameters the rule claims.

##### tag

The tag of a strategy rule, or `None` for the other rules.

### SchedulerSpec

The scheduler of one rule.

#### Methods

##### from_config

```python
def from_config(config: SchedulerConfig, total_epochs: int) -> SchedulerSpec:
```

Create the specification from a scheduler config.

For `CosineAnnealingLR`, the method sets `T_max` to `total_epochs` when `params` has no `T_max`, and logs a warning. It also logs
a warning when `T_max` is not equal to `total_epochs`. The method adds the default before it computes `key`. A rule that omits
`T_max` and a rule that sets it to `total_epochs` therefore get the same `key`.

> **Example**
> ```pycon
>>> from luxonis_train.config.config import SchedulerConfig
>>> step = SchedulerConfig(
...     name="StepLR", params={"step_size": 5, "gamma": 0.5}
... )
>>> SchedulerSpec.from_config(step, total_epochs=10).key
'{"name": "StepLR", "params": {"gamma": 0.5, "step_size": 5}}'
```

Parameters

 * `config` (`SchedulerConfig`): The scheduler config.
 * `total_epochs` (`int`): The number of training epochs, from `trainer.epochs`.

Returns

 * `SchedulerSpec`: The name of `config`, a shallow copy of its `params` with the `T_max` default when the method adds one, and the `key`.

Raises

 * `KeyError`: When the `SCHEDULERS` registry has no scheduler with the name of `config`.

#### Attributes

##### key

A JSON string of `name` and `params` with sorted keys. A value that JSON cannot encode goes into the string as its `repr`. Rules with the same optimizer name and the same `key` share one inner optimizer. The equality check of the dataclass ignores this field.

##### name

The class name of the scheduler in the `SCHEDULERS` registry.

##### params

The scheduler parameters. [SchedulerSpec.from_config](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) adds `T_max` when a `CosineAnnealingLR` config has no `T_max`.

### Selector

A predicate that decides whether a rule claims a parameter.

[pattern_selector](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) builds a selector from the `parameters` patterns of a node `finetuning` entry. A training strategy gives its own selector in each [StrategyRule](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md). A selector sees only the parameters that no earlier rule claimed.

### StrategyRule

A parameter-group rule that a training strategy contributes.

[BaseTrainingStrategy.rules](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/base_strategy.md) returns these rules. [resolve_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) evaluates them in order, after the `finetuning` entries of every node and before the default rule. A rule visits every node before the next rule starts. It claims each parameter that its `selector` accepts and that no earlier rule claimed.

The groups of a rule are named `strategy/<tag>`. A node with `freezing.active` gets its own group `strategy/<tag>/<node>`.

#### Attributes

##### optimizer

The optimizer of the claimed parameters. The plan uses it as it is and does not merge it with the base optimizer.

##### scheduler

The scheduler of the claimed parameters. `None` uses the base scheduler of the strategy, or `trainer.scheduler` when [BaseTrainingStrategy.get_base_configs](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/base_strategy.md) raises `NotImplementedError`.

##### selector

The predicate that decides which parameters the rule claims.

##### tag

The name of the rule. [BaseTrainingStrategy.attach](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/base_strategy.md) receives the handles of the groups of the rule under this key. When the rule claims no parameter, the mapping has no entry for the tag.

### TrainingPlan

The static partition of the model parameters into groups.

[resolve_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) puts every parameter of every node into exactly one group, frozen parameters included. The only exception is a parameter that a legacy training strategy claims for its own optimizers.

#### Attributes

##### handles_by_node

For each node name, the handles of the groups that hold parameters of the node, in plan order. A node without a parameter in the plan has no entry.

##### handles_by_tag

For each strategy rule tag, the handles of the groups of the rule, in plan order. A tag whose rule claims no parameter has no entry.

##### inners

The specifications of the inner optimizers, in creation order.

### TrainingPlanRuntime

The optimizers and schedulers that [build_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) creates.

The runtime also gives access to a parameter group through its [GroupHandle](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md). The freeze schedule and the training strategy use this access to change the options of their groups, such as the learning rate.

#### Methods

##### group

```python
def group(handle: GroupHandle) -> dict[str, Any]:
```

Return the parameter group at a handle.

Parameters

 * `handle` (`GroupHandle`): The address of the group.

Returns

 * `dict[str, Any]`: The entry of `param_groups` in the inner optimizer. It is not a copy, so a change to it changes the optimizer.

##### handles_for_node

```python
def handles_for_node(node_name: str) -> tuple[GroupHandle, ...]:
```

Return the handles of the groups of a node.

Parameters

 * `node_name` (`str`): The name of the node.

Returns

 * `tuple[GroupHandle, ...]`: The handles of the groups that hold parameters of the node, in plan order. The tuple is empty when no group holds a parameter of `node_name`.

##### set_group_base_lr

```python
def set_group_base_lr(handle: GroupHandle, lr: float):
```

Set a new base learning rate for one parameter group.

The method sets `lr` and `initial_lr` of the group. It also sets the entry of the group in the `base_lrs` of the member scheduler. For a `SequentialLR` or a `ChainedScheduler`, it sets the entry in each child scheduler. The next scheduler steps therefore start from the new rate. The method skips a scheduler without `base_lrs`, such as a `ReduceLROnPlateau`. When the inner optimizer has no scheduler, only the group changes.

Parameters

 * `handle` (`GroupHandle`): The address of the group.
 * `lr` (`float`): The new base learning rate.

#### Attributes

##### inner_optimizers

One optimizer for each [InnerSpec](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) of the plan, in plan order. The optimizers of a legacy training strategy follow them.

##### members

The scheduler of each inner optimizer, in the same order. The entry is `None` for a legacy optimizer without a scheduler.

##### optimizer

The optimizer that Lightning receives. It is the only inner optimizer, or a [CompositeOptimizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md) over all inner optimizers.

##### plan

The plan of the runtime.

##### scheduler_configs

The schedulers and scheduler configs that Lightning receives. [build_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) describes the entries.

## Functions

### build_training_plan

```python
def build_training_plan(plan: TrainingPlan, cfg: Config, main_metric_monitor: str | None, strategy: BaseTrainingStrategy | None =
None) -> TrainingPlanRuntime:
```

Create the optimizers and the schedulers of a plan.

For each inner optimizer of the plan, the function creates the optimizer from the `OPTIMIZERS` registry. Each group of the plan becomes one parameter group with the group options. When the plan has more than one group, each parameter group also gets the `name` of its group. The `LearningRateMonitor` key of each group then ends with the group name instead of `pg1`, `pg2`, and so on.

For each of these optimizers, the function also creates the member scheduler from the `SCHEDULERS` registry. A `SequentialLR` is the exception: the function creates it directly from torch. Its `params` give the `milestones`, the `last_epoch`, and the child schedulers, which come from the registry. A `ReduceLROnPlateau` monitors `main_metric_monitor` in `max` mode and `val/loss` in any other mode.

The optimizers of [BaseTrainingStrategy.opaque_inners](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/base_strategy.md) follow the optimizers of the plan. Only a legacy strategy returns such optimizers.

With one optimizer in total, Lightning receives that optimizer and its scheduler:

 * A `ReduceLROnPlateau` goes into a dictionary with the keys `scheduler`, `monitor`, and `frequency`.
 * Another scheduler goes as it is.
 * The scheduler or config of a legacy strategy goes as it is. A legacy optimizer without a scheduler gives no scheduler entry.

With several optimizers, one [CompositeOptimizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md) wraps them. Lightning then receives these scheduler configs:

 * One [CompositeLRScheduler](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/schedulers/composite_scheduler.md), named `lr`, that steps all schedulers without a monitor. It is present only when at least one scheduler has no monitor.
 * One [CompositeReduceLROnPlateau](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/schedulers/composite_scheduler.md) for each monitor, with the `monitor`, `frequency`, and `reduce_on_plateau` keys. The configs are named `lr-plateau`, `lr-plateau-1`, and so on.

In this case, the function reads only the `scheduler` and the `monitor` keys of a legacy scheduler config. A legacy scheduler with a `monitor` must be a `ReduceLROnPlateau`.

[CompositeOptimizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md) raises `ValueError` when one of the several optimizers is an `LBFGS` optimizer. It also raises `ValueError` when the plan and the strategy give no optimizer.

Each `frequency` above is `trainer.validation_interval`.

Parameters

 * `plan` (`TrainingPlan`): The plan from [resolve_training_plan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md).
 * `cfg` (`Config`): The config. The function reads `trainer.validation_interval`.
 * `main_metric_monitor` (`str | None`): The logged name of the main metric, or `None` when the model has no main metric.
 * `strategy` (`BaseTrainingStrategy | None`): The training strategy, or `None`. The function reads its [BaseTrainingStrategy.opaque_inners](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/base_strategy.md).

Returns

 * `TrainingPlanRuntime`: The optimizers, the schedulers, and the scheduler configs for Lightning.

Raises

 * `TypeError`: When a group option is not a key of the `defaults` of its optimizer.
 * `ValueError`: When a `ReduceLROnPlateau` in `max` mode has no `main_metric_monitor`.

### merge_config_items

```python
def merge_config_items(base: OptimizerConfig | SchedulerConfig, override: FinetuningOptimizerConfig | FinetuningSchedulerConfig |
None) -> OptimizerConfig | SchedulerConfig:
```

Merge a finetuning override into a base optimizer or scheduler.

The merge follows these rules:

 * Without an override, the result is a copy of `base`.
 * An override without a `name`, or with the name of `base`, keeps the name of `base`. The result has the `params` of `base`, updated with the `params` of the override.
 * An override with a different `name` replaces `base`. The result has only the `params` of the override.

> **Example**
> ```pycon
>>> from luxonis_train.config.config import (
...     FinetuningOptimizerConfig,
...     OptimizerConfig,
... )
>>> base = OptimizerConfig(
...     name="SGD", params={"lr": 0.01, "momentum": 0.9}
... )
>>> lower_lr = FinetuningOptimizerConfig(params={"lr": 0.001})
>>> merge_config_items(base, lower_lr)
OptimizerConfig(name='SGD', params={'lr': 0.001, 'momentum': 0.9})
>>> adam = FinetuningOptimizerConfig(name="Adam", params={"lr": 0.001})
>>> merge_config_items(base, adam)
OptimizerConfig(name='Adam', params={'lr': 0.001})
```

Parameters

 * `base` (`OptimizerConfig | SchedulerConfig`): The base config: `trainer.optimizer`, `trainer.scheduler`, or a base config of
   the training strategy.
 * `override` (`FinetuningOptimizerConfig | FinetuningSchedulerConfig | None`): The override of a node `finetuning` entry, or
   `None`.

Returns

 * `OptimizerConfig | SchedulerConfig`: A new config of the type of `base`, so its `name` is never `None`. The arguments do not
   change.

### pattern_selector

```python
def pattern_selector(patterns: Sequence[ParameterPattern]) -> Selector:
```

Build a
[Selector](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md)
from the patterns of a `finetuning` entry.

The selector joins `module_name` and `parameter_name` into the dotted parameter name, such as `"stem.conv.weight"`. When
`module_name` is empty, the dotted name is `parameter_name` alone. The selector accepts the parameter when at least one pattern
matches. A pattern matches when its `name` matches the dotted name and its `module_type` matches the class name of `module`. A
field left as `None` matches everything. Both fields are regular expressions. `re.search` matches them without anchors and without
case, as
[ParameterPattern](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/config/config.md)
describes.

> **Example**
> ```pycon
>>> from torch import nn
>>> from luxonis_train.config.config import ParameterPattern
>>> linear = nn.Linear(2, 2)
>>> select = pattern_selector([ParameterPattern(name="head.weight")])
>>> select(linear, "head", linear.weight, "weight")
True
>>> select(linear, "head", linear.bias, "bias")
False
```

Parameters

 * `patterns` (`Sequence[ParameterPattern]`): The patterns. An empty sequence gives a selector that accepts no parameter.

Returns

 * `Selector`: A function with the arguments of [Selector.call](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md). It returns `True` when a pattern matches the parameter.

### resolve_training_plan

```python
def resolve_training_plan(cfg: Config, nodes: Nodes, strategy: BaseTrainingStrategy | None = None) -> TrainingPlan:
```

Resolve the parameter rules of a model into a [TrainingPlan](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md).

The function combines the node `finetuning` entries, the rules of the training strategy, the node freezing, and the base optimizer and scheduler. It does not create the optimizers or the schedulers of the plan.

The first rule that accepts a parameter claims it. The rules claim in this order:

 1. The `finetuning` entries of a node, in config order. An entry claims only parameters of its own node. An entry without `parameters` claims every free parameter of the node. Its optimizer and scheduler overrides merge into the base configs through [merge_config_items](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md).
 2. The rules of `strategy`, in order. Each rule visits every node before the next rule starts. A legacy strategy claims its own parameters before these rules. Those parameters stay outside the plan.
 3. The default rule. It claims every parameter that is left, with the base optimizer and scheduler.

Without a strategy, the default rule runs for each node directly after the `finetuning` entries of that node. With a strategy, the entries of all nodes run first, then the strategy rules, then the default rule. The function visits the nodes in the order of `nodes` and the parameters in module order, so the result is deterministic. The rules also claim frozen parameters, so a node that unfreezes already has an optimizer.

The groups get these names:

 * `<node>/<index>` for a `finetuning` entry. `<index>` is the position of the entry in the `finetuning` list of the node, from `0`.
 * `strategy/<tag>` for a strategy rule, shared by all nodes.
 * `default` for the default rule, shared by all nodes.
 * A node with `freezing.active` gets its own strategy and default groups, `strategy/<tag>/<node>` and `default/<node>`. The `lr_after_unfreeze` rate of the node therefore changes no group of another node.
 * Without a strategy, when any node has `finetuning` entries, every node gets its own default group `default/<node>`.

Rules with the same optimizer name and the same [SchedulerSpec.key](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) share one inner optimizer.

The base optimizer and scheduler are `trainer.optimizer` and `trainer.scheduler`. With a strategy, they come from [BaseTrainingStrategy.get_base_configs](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/strategies/base_strategy.md), unless that method raises `NotImplementedError`.

[SchedulerSpec.from_config](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/training_plan.md) logs a warning for each `CosineAnnealingLR` rule whose `T_max` is missing or differs from `trainer.epochs`.

Parameters

 * `cfg` (`Config`): The config. The function reads `trainer.epochs`, `trainer.optimizer`, and `trainer.scheduler`.
 * `nodes` (`Nodes`): The nodes of the model. The function reads the `name`, `module`, `finetuning`, and `unfreeze_after` of each [NodeWrapper](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/utils.md).
 * `strategy` (`BaseTrainingStrategy | None`): The training strategy, or `None`.

Returns

 * `TrainingPlan`: The plan.

Raises

 * `ValueError`: When a `finetuning` entry claims no parameter, for example because earlier entries claimed all its matches. Also when a legacy strategy claims a parameter that a `finetuning` entry claimed.
 * `KeyError`: When an optimizer or a scheduler name is not in its registry.
