# luxonis_progress_bar

Python API: `luxonis_train.callbacks.luxonis_progress_bar`

The progress bars of the trainer, and the optimizer summary.

[LuxonisRichProgressBar](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md)
and
[LuxonisTQDMProgressBar](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md)
show the batch progress and print the results of each evaluation epoch. Both bars mirror the printed results to the log file.

[build_optimizer_summary](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md)
and
[log_optimizer_summary](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md)
report how the parameters of the model are split across the optimizers and their parameter groups. The summary shows the effect of
the finetuning rules and of the training strategy of the config.

## Classes

### BaseLuxonisProgressBar

Base class for the progress bars of the trainer.

The class drops the `v_num` item from the bar and adds the running mean of the train loss as `Loss`. A subclass prints the results
of an evaluation epoch with
[BaseLuxonisProgressBar.print_results](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md)
and
[BaseLuxonisProgressBar.print_table](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).
The subclasses also write one summary line per train epoch to the log file.

#### Methods

##### format_matrix_for_printing

```python
def format_matrix_for_printing(node: Any, name: str, value: Tensor) -> dict[str, Any]:
```

Convert a matrix metric into a printable dictionary.

The row and column labels are the class names of `node` when the matrix has one row (column) per class. When the matrix has one
extra row (column), the labels are the class names plus `"no match"`. Otherwise the labels are the indices as strings. When the
class names of `node` raise a `RuntimeError`, for example because the node has no dataset metadata, the class names count as
empty. Any other exception of that lookup propagates.

The result has the keys `"values"` (the matrix as a nested list), `"row_labels"`, `"col_labels"`, `"row_axis"` (`"GT"`), and
`"col_axis"` (`"Pred"`).

> **Example**
> ```pycon
>>> import torch
>>> from types import SimpleNamespace
>>> bar = LuxonisTQDMProgressBar()
>>> node = SimpleNamespace(class_names=["cat", "dog"])
>>> matrix = torch.tensor([[3, 1, 0], [0, 2, 1]])
>>> info = bar.format_matrix_for_printing(node, "cm", matrix)
>>> info["row_labels"], info["col_labels"]
(['cat', 'dog'], ['cat', 'dog', 'no match'])
>>> info["values"]
[[3, 1, 0], [0, 2, 1]]
>>> info["row_axis"], info["col_axis"]
('GT', 'Pred')
```

Parameters

 * `node` (`Any`): The node the metric is attached to. When the object has a `module` attribute, as a [NodeWrapper](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/utils.md) has, the method reads the class names from that attribute.
 * `name` (`str`): Name of the metric. Unused.
 * `value` (`Tensor`): The matrix, of shape `[R, C]`.

Returns

 * `dict[str, Any]`: The matrix values and their labels.

##### get_metrics

```python
def get_metrics(trainer: pl.Trainer, pl_module: lxt.LuxonisLightningModule) -> dict[str, int | str | float | dict[str, float]]:
```

Return the items shown at the end of the progress bar.

The items are the metrics that the model logs with `prog_bar=True`, as `ProgressBar.get_metrics` collects them, without the `v_num` entry. When the train loss accumulator of `pl_module` holds a `"loss"` entry, the method adds it as `Loss`. That value is the running mean of the total loss over the batches of the current epoch so far.

Parameters

 * `trainer` (`pl.Trainer`): The trainer.
 * `pl_module` (`lxt.LuxonisLightningModule`): The model. Its train loss accumulator provides the `Loss` value.

Returns

 * `dict[str, int | str | float | dict[str, float]]`: The items to show, keyed by name.

##### print_results

```python
def print_results(stage: str, loss: float, metrics: Mapping[str, Mapping[str, int | str | float]], matrices: Mapping[str,
Mapping[str, Mapping[str, Any]]]):
```

Print the results of an evaluation epoch.

An implementation must print the stage name, the loss, one table per node in `metrics`, and one table per matrix in `matrices`.

Parameters

 * `stage` (`str`): Name of the stage, for example `"Validation"`.
 * `loss` (`float`): Mean loss of the epoch.
 * `metrics` (`Mapping[str, Mapping[str, int | str | float]]`): Scalar metrics as `{node_name: {metric_name: value}}`.
 * `matrices` (`Mapping[str, Mapping[str, Mapping[str, Any]]]`): Matrix metrics as `{node_name: {metric_name: matrix}}`. Each matrix is a dictionary in the format of [BaseLuxonisProgressBar.format_matrix_for_printing](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).

##### print_table

```python
def print_table(title: str, table: Iterable[tuple[str | int | float, ...]], column_names: list[str]):
```

Print one table.

An implementation must print `title`, then a table with one header per entry of `column_names` and one row per tuple of `table`.

Parameters

 * `title` (`str`): Title of the table.
 * `table` (`Iterable[tuple[str | int | float, ...]]`): The rows. Each row is a tuple with one value per column.
 * `column_names` (`list[str]`): Names of the columns.

### LuxonisRichProgressBar

Progress bar that prints styled text with `rich`.

[LuxonisModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/core.md) uses this bar when `rich_logging` is `True` in the config. The bar prints the results of an evaluation epoch as `rich` tables on the console. It renders the same tables without terminal styling into a buffer and writes the buffer to the log file only.

#### Methods

##### init

```python
def __init__(self):
```

Initialize the bar with `leave=True` and a log console.

The finished train bar stays in the terminal at the end of each epoch. The log console writes into an in-memory buffer without terminal styling. [LuxonisRichProgressBar.print_results](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md) writes the buffer to the log file and clears it.

##### on_train_epoch_end

```python
def on_train_epoch_end(trainer: pl.Trainer, pl_module: lxt.LuxonisLightningModule):
```

Refresh the train task and log the epoch summary.

Lightning calls this hook at the end of every train epoch. The Lightning `RichProgressBar` base class updates the metrics column of an enabled display and refreshes the display. The items come from [BaseLuxonisProgressBar.get_metrics](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md), with every tensor replaced by its `item()` value. Then one summary line goes to the log file only, in the form `[Epoch <n>/<max>] Duration: <s>s | Train Loss: <loss>`. The duration is the time since [LuxonisRichProgressBar.on_train_epoch_start](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md), the validation of the epoch included. The loss is the `train/loss` value of `trainer.callback_metrics`.

Lightning runs this hook before [LuxonisLightningModule.on_train_epoch_end](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/luxonis_lightning.md), which logs `train/loss`. So the line shows the value of the previous epoch. It shows `N/A` when the value is missing or zero, as in the first epoch.

Parameters

 * `trainer` (`pl.Trainer`): The trainer.
 * `pl_module` (`lxt.LuxonisLightningModule`): The model. The base class reads its metrics for the metrics column.

##### on_train_epoch_start

```python
def on_train_epoch_start(trainer: pl.Trainer, pl_module: lxt.LuxonisLightningModule):
```

Start the train task and record the epoch start time.

Lightning calls this hook at the start of every train epoch. The Lightning `RichProgressBar` base class adds a train task named after the epoch. From the second epoch on, it first stops the current display and starts a new one, because `leave` is `True`. A disabled bar skips all of that. The start time feeds the duration in [LuxonisRichProgressBar.on_train_epoch_end](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).

Parameters

 * `trainer` (`pl.Trainer`): The trainer.
 * `pl_module` (`lxt.LuxonisLightningModule`): The model. Unused.

##### print_results

```python
def print_results(stage: str, loss: float, metrics: Mapping[str, Mapping[str, int | str | float]], matrices: Mapping[str,
Mapping[str, Mapping[str, Any]]]):
```

Print the results of an evaluation epoch.

On the console, the output starts with a magenta rule that holds the stage name, then the loss and a `Metrics:` heading. Each node in `metrics` gets a `rich` table with its name as the title and the columns `Name` and `Value`. The matrices of that node follow its table. The matrices of the nodes that have no scalar metrics come last, under the title `<node>/<matrix title>`. The title of a matrix is its name in title case, with spaces for underscores. A matrix prints as a table. Its header holds the `row_axis` and `col_axis` names of the matrix and the column labels. Each row starts with its row label. A closing rule ends the output. The same output, without terminal styling, goes to the log file only. The method then clears the log buffer.

Parameters

 * `stage` (`str`): Name of the stage, for example `"Validation"`.
 * `loss` (`float`): Mean loss of the epoch.
 * `metrics` (`Mapping[str, Mapping[str, int | str | float]]`): Scalar metrics as `{node_name: {metric_name: value}}`.
 * `matrices` (`Mapping[str, Mapping[str, Mapping[str, Any]]]`): Matrix metrics as `{node_name: {metric_name: matrix}}`, in the format of [BaseLuxonisProgressBar.format_matrix_for_printing](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).

Raises

 * `RuntimeError`: When the console of the bar does not exist yet, see [LuxonisRichProgressBar.console](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).

##### print_table

```python
def print_table(title: str, table: Iterable[tuple[str | int | float, ...]], column_names: list[str], console: Console | None =
None):
```

Print one table as a `rich` table.

The title is bold and the headers are bold magenta. The first column is magenta and the other columns are white. `str` converts the first element of each row. A `float` in the other elements prints with five decimals. Any other element prints with `str`.

Parameters

 * `title` (`str`): Title of the table.
 * `table` (`Iterable[tuple[str | int | float, ...]]`): The rows. Each row is a tuple with one value per column.
 * `column_names` (`list[str]`): Names of the columns.
 * `console` (`Console | None`): The console to print to. `None` means the console of the bar, the terminal.

Raises

 * `RuntimeError`: When `console` is `None` and the console of the bar does not exist yet, see [LuxonisRichProgressBar.console](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).

#### Attributes

##### console

The `rich` console that the bar prints to.

The Lightning `RichProgressBar` base class creates the console when a stage starts and the bar is enabled. A bar that is disabled when the stage starts gets no console.

Raises

 * `RuntimeError`: When the console does not exist yet. The message asks the user to set `rich_logging` to `False` in the config.

### LuxonisTQDMProgressBar

Progress bar that prints plain text with `tqdm`.

[LuxonisModel](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/core.md) uses this bar when `rich_logging` is `False` in the config. The bar prints the results of an evaluation epoch as `tabulate` grids through the logger, so that the console and the log file receive the same text.

#### Methods

##### init

```python
def __init__(self):
```

Initialize the bar with `leave=True`.

The finished train bar stays in the terminal at the end of each epoch, and the next epoch gets a new bar.

##### on_train_epoch_end

```python
def on_train_epoch_end(trainer: pl.Trainer, pl_module: lxt.LuxonisLightningModule):
```

Close the train bar and log the epoch summary.

Lightning calls this hook at the end of every train epoch. The Lightning `TQDMProgressBar` base class sets the postfix of an enabled bar from [BaseLuxonisProgressBar.get_metrics](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md) and closes the bar. Then one summary line goes to the log file only, in the form `[Epoch <n>/<max>] Duration: <s>s | Train Loss: <loss>`. The duration is the time since [LuxonisTQDMProgressBar.on_train_epoch_start](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md), the validation of the epoch included. The loss is the `train/loss` value of `trainer.callback_metrics`.

Lightning runs this hook before [LuxonisLightningModule.on_train_epoch_end](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/luxonis_lightning.md), which logs `train/loss`. So the line shows the value of the previous epoch. It shows `N/A` when the value is missing or zero, as in the first epoch.

Parameters

 * `trainer` (`pl.Trainer`): The trainer.
 * `pl_module` (`lxt.LuxonisLightningModule`): The model. The base class reads its metrics for the bar postfix.

##### on_train_epoch_start

```python
def on_train_epoch_start(trainer: pl.Trainer, pl_module: lxt.LuxonisLightningModule):
```

Start a new train bar and record the epoch start time.

Lightning calls this hook at the start of every train epoch. The Lightning `TQDMProgressBar` base class creates a new bar, because `leave` is `True`, and sets its description to `Epoch N`. The start time feeds the duration in [LuxonisTQDMProgressBar.on_train_epoch_end](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).

Parameters

 * `trainer` (`pl.Trainer`): The trainer.
 * `pl_module` (`lxt.LuxonisLightningModule`): The model. Unused.

##### print_results

```python
def print_results(stage: str, loss: float, metrics: Mapping[str, Mapping[str, int | str | float]], matrices: Mapping[str,
Mapping[str, Mapping[str, Any]]]):
```

Print the results of an evaluation epoch through the logger.

The output starts with a rule that holds the stage name, then the loss and a `Metrics:` heading. Each node in `metrics` gets a rule with its name and a `tabulate` grid with the columns `Name` and `Value`. The matrices of that node follow its grid. The matrices of the nodes that have no scalar metrics come last, under the title `<node>/<matrix title>`. The title of a matrix is its name in title case, with spaces for underscores. A matrix prints as a rule with its title and a grid. The header of the grid holds the `row_axis` and `col_axis` names of the matrix and the column labels. Each row starts with its row label. A closing rule ends the output. Every line goes through `logger.info`, so it reaches the console and the log file.

Parameters

 * `stage` (`str`): Name of the stage, for example `"Validation"`.
 * `loss` (`float`): Mean loss of the epoch.
 * `metrics` (`Mapping[str, Mapping[str, int | str | float]]`): Scalar metrics as `{node_name: {metric_name: value}}`.
 * `matrices` (`Mapping[str, Mapping[str, Mapping[str, Any]]]`): Matrix metrics as `{node_name: {metric_name: matrix}}`, in the format of [BaseLuxonisProgressBar.format_matrix_for_printing](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).

##### print_table

```python
def print_table(title: str, table: Iterable[tuple[str | int | float, ...]], column_names: list[str]):
```

Print one table as a `tabulate` grid through the logger.

The output is a rule with `title`, then the table in the `fancy_grid` format with right-aligned numbers.

Parameters

 * `title` (`str`): Title of the table.
 * `table` (`Iterable[tuple[str | int | float, ...]]`): The rows. Each row is a tuple with one value per column.
 * `column_names` (`list[str]`): Names of the columns.

## Functions

### build_optimizer_summary

```python
def build_optimizer_summary(optimizers: Sequence[Optimizer], schedulers: Sequence[LRSchedulerTypeUnion | LRSchedulerConfig],
modules: Mapping[str, nn.Module]) -> dict[str, Any]:
```

Build the summary of the optimizers and their parameter groups.

The summary is a nested dictionary that [log_optimizer_summary](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md) renders and writes as JSON to the log file. Each parameter group lists its hyperparameters and its owners, the modules of `modules` whose parameters it holds.

Two denominators are in use:

 * The group-level percentages `tensors_pct_of_model` and `params_pct_of_model` are relative to all parameters of the model. The sum over all groups of all optimizers is 100% minus the share of the parameters that no group holds. A parameter that two groups hold counts twice in that sum.
 * The owner-level percentages `tensors_pct_of_owner` and `params_pct_of_owner` are relative to all parameters of that owner. Over all appearances of one owner they add up to 100% when every parameter of the owner is in exactly one group. This shows how the parameters of a node are split across groups.

A frozen parameter counts in its group like any other parameter. The percentages include frozen parameters, and each group and owner reports its trainable and frozen counts separately.

The owner `"<external>"` collects the parameters that a group holds but that no module in `modules` owns. A parameter that several modules share belongs to the first module of `modules` that lists it. In every count, `*_tensors` is a number of parameter tensors and `*_params` is a number of elements. A percentage is `0.0` when its denominator is zero.

The result has these keys:

 * `n_optimizers`: Number of optimizers.
 * `model_tensors`, `model_params`: Totals over all owners, external ones included.
 * `trainable_tensors`, `trainable_params`, `frozen_tensors`, `frozen_params`: The totals split by `requires_grad`.
 * `optimizers`: One entry per optimizer with `index`, `optimizer` and `scheduler` (class names), `n_groups`, and `groups`.

Each group holds `index`, `n_tensors`, `n_params`, the trainable and frozen counts, `tensors_pct_of_model`, `params_pct_of_model`, `hyperparams`, and `owners`. `hyperparams` holds the entries of the parameter group except `params`, callables, lists, tuples, and dictionaries. `owners` lists the owners in descending order of their `n_params` in the group. Each owner holds `name`, `n_tensors`, `n_tensors_of_owner`, `tensors_pct_of_owner`, `n_params`, `n_params_of_owner`, `params_pct_of_owner`, and the trainable and frozen counts.

> **Example**
> ```pycon
>>> from torch import nn
>>> from torch.optim import SGD
>>> from torch.optim.lr_scheduler import ConstantLR
>>> backbone = nn.Linear(4, 8)
>>> head = nn.Linear(8, 2)
>>> optimizer = SGD(
...     [
...         {"params": backbone.parameters()},
...         {"params": head.parameters(), "lr": 0.1},
...     ],
...     lr=0.01,
... )
>>> summary = build_optimizer_summary(
...     [optimizer],
...     [ConstantLR(optimizer, factor=1.0)],
...     {"backbone": backbone, "head": head},
... )
>>> summary["n_optimizers"], summary["model_params"]
(1, 58)
>>> group = summary["optimizers"][0]["groups"][0]
>>> group["hyperparams"]["lr"]
0.01
>>> round(group["params_pct_of_model"], 1)
69.0
>>> [owner["name"] for owner in group["owners"]]
['backbone']
```

Parameters

 * `optimizers` (`Sequence[Optimizer]`): The optimizers, in order.
 * `schedulers` (`Sequence[LRSchedulerTypeUnion | LRSchedulerConfig]`): One entry per optimizer. A dictionary contributes the
   class name of its `"scheduler"` entry, as a Lightning scheduler config dictionary holds one. Any other object contributes its
   own class name, so `None` for an optimizer without a scheduler shows as `"NoneType"`. A Lightning `LRSchedulerConfig` dataclass
   is not a dictionary, so it shows as `"LRSchedulerConfig"`.
 * `modules` (`Mapping[str, nn.Module]`): The owner modules keyed by name, usually the nodes of the model keyed by node name.

Returns

 * `dict[str, Any]`: The summary described above.

Raises

 * `ValueError`: When `optimizers` and `schedulers` differ in length.

### log_optimizer_summary

```python
def log_optimizer_summary(summary: dict[str, Any], use_rich: bool = True):
```

Log the summary from
[build_optimizer_summary](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).

With `use_rich`, the summary goes to the global `rich` console as nested panels. A header panel holds the totals. Then one panel
per optimizer holds one panel per group. A group panel holds one panel with the hyperparameters and one with the owners. Without
`use_rich`, the summary goes through `logger.info` as an indented plain-text list, which reaches the console and the log file. In
both cases the function also writes the summary as JSON to the log file only, with `str` for a value that JSON cannot encode.

Parameters

 * `summary` (`dict[str, Any]`): The summary from
   [build_optimizer_summary](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/luxonis_progress_bar.md).
 * `use_rich` (`bool`): Whether to render the summary with `rich` panels.
