# composite_optimizer

Python API: `luxonis_train.optimizers.composite_optimizer`

An optimizer that drives several inner optimizers.

Lightning sees one optimizer, so gradient accumulation and gradient clipping still work for any number of inner optimizers that
the finetuning rules and the training strategy produce.

## Classes

### CompositeOptimizer

One `torch.optim.Optimizer` that drives several inner optimizers.

[param_groups](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md)
joins the `param_groups` of the inner optimizers, and holds the same dictionary objects. Lightning can therefore drive several
optimizer configurations in its automatic optimization:

 * one `step` call,
 * one gradient clipping pass over all groups,
 * one gradient scaler slot.

The partition of the parameters is static. No group joins, leaves, or moves after the constructor. When a node freezes, only
`requires_grad` changes, and an inner optimizer skips a parameter whose gradient is `None`.

The constructor does not call `Optimizer.__init__`, because that method builds its own parameter groups. The class itself
implements the methods that Lightning uses: the `Optimizable` protocol,
[step](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md),
[zero_grad](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md),
[state_dict](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md),
and
[load_state_dict](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md).

> **Example**
> ```pycon
>>> from torch import nn
>>> from torch.optim import SGD, Adam
>>> sgd = SGD(nn.Linear(4, 8).parameters(), lr=0.1)
>>> adam = Adam(nn.Linear(8, 2).parameters(), lr=0.01)
>>> composite = CompositeOptimizer([sgd, adam])
>>> [group["lr"] for group in composite.param_groups]
[0.1, 0.01]
>>> "betas" in composite.defaults
False
>>> composite.state_dict()["optimizers"]
['SGD', 'Adam']
```

#### Methods

##### init

```python
def __init__(inners: Sequence[Optimizer]):
```

Wrap the inner optimizers.

The `defaults` of the composite hold only the keys that all inner optimizers share, with the values of the first one.

Parameters

 * `inners` (`Sequence[Optimizer]`): The inner optimizers, in the order of their groups in [param_groups](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md).

Raises

 * `ValueError`: If `inners` is empty. Also if `inners` has more than one optimizer and one of them is an `LBFGS` optimizer. `LBFGS` needs the step closure, and [step](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md) does not pass the closure to the inner optimizers.

##### add_param_group

```python
def add_param_group(param_group: dict[str, Any]):
```

Reject a new parameter group, because the partition is fixed.

Parameters

 * `param_group` (`dict[str, Any]`): The group. The method does not use it.

Raises

 * `RuntimeError`: Always.

##### load_state_dict

```python
def load_state_dict(state_dict: dict[str, Any]):
```

Load a composite state into the inner optimizers.

The method loads each entry of `"inners"` into the inner optimizer at the same position.

Parameters

 * `state_dict` (`dict[str, Any]`): A state that the [CompositeOptimizer.state_dict](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md) method returned.

Raises

 * `ValueError`: If `"format"` is not `"luxonis_composite"`, for example in the checkpoint of a single optimizer. Also if `"version"` is not `1`, or if the class names in `"optimizers"` differ from the current inner optimizers.

##### param_groups.setter

```python
def param_groups.setter(value: Any):
```

Reject a new value, because the partition is fixed.

##### state.setter

```python
def state.setter(value: Any):
```

Reject a new value, because the state is a view.

##### state_dict

```python
def state_dict(self) -> dict[str, Any]:
```

Return the state of every inner optimizer.

Returns

 * `dict[str, Any]`: A dictionary with these keys. * `"format"` is `"luxonis_composite"`.
    * `"version"` is `1`.
    * `"optimizers"` holds the class name of each inner optimizer.
    * `"inners"` holds the `state_dict()` of each inner optimizer.

##### step

```python
def step(closure: Callable[[], Any] | None = None) -> Any:
```

Run the closure once, then step every inner optimizer.

The method calls `closure` with gradients on. It then calls `step()` of each inner optimizer in order, without a closure. The inner steps fire the step hooks of torch. The composite fires no hooks of its own.

Parameters

 * `closure` (`Callable[[], Any] | None`): The function that computes the loss and the gradients, or `None`.

Returns

 * `Any`: The return value of `closure`, or `None` without a closure.

##### zero_grad

```python
def zero_grad(set_to_none: bool = True):
```

Reset the gradients of every inner optimizer.

Parameters

 * `set_to_none` (`bool`): Whether to set the gradients to `None` instead of to zero. The method passes it to each inner optimizer.

#### Attributes

##### defaults

##### inner_optimizers

The inner optimizers, in the order of the constructor.

##### param_groups

The parameter groups of all inner optimizers, in order.

Each access builds a new list, but the list holds the group dictionaries of the inner optimizers. A change to a group therefore changes its inner optimizer. An assignment to the property raises `TypeError`.

##### state

A live view over the states of the inner optimizers.

A write to `state[parameter]` goes to the inner optimizer that owns the parameter. A parameter that no inner optimizer owns raises `KeyError`. An assignment to the property raises `TypeError`.

##### STATE_DICT_FORMAT

## Functions

### unwrap_optimizers

```python
def unwrap_optimizers(optimizers: Sequence[Optimizer]) -> list[Optimizer]:
```

Replace each [CompositeOptimizer](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/optimizers/composite_optimizer.md) with its inner optimizers.

A plain optimizer stays as it is. A caller can therefore treat a run with one optimizer and a run with a composite in the same way.

> **Example**
> ```pycon
>>> from torch import nn
>>> from torch.optim import SGD, Adam
>>> sgd = SGD(nn.Linear(2, 2).parameters(), lr=0.1)
>>> adam = Adam(nn.Linear(2, 2).parameters(), lr=0.01)
>>> composite = CompositeOptimizer([sgd, adam])
>>> unwrap_optimizers([composite]) == [sgd, adam]
True
>>> unwrap_optimizers([sgd]) == [sgd]
True
```

Parameters

 * `optimizers` (`Sequence[Optimizer]`): The optimizers, such as `trainer.optimizers` of Lightning.

Returns

 * `list[Optimizer]`: A new list with the plain optimizers, in order.
