# gpu_stats_monitor

Python API: `luxonis_train.callbacks.gpu_stats_monitor`

Logs the GPU statistics of `nvidia-smi` while a model trains.

Copyright The PyTorch Lightning team.

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License.
You may obtain a copy of the License at [http://www.apache.org/licenses/LICENSE-2.0](http://www.apache.org/licenses/LICENSE-2.0)

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS"
BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language
governing permissions and limitations under the License.

## Classes

### GPUStatsMonitor

Callback that logs the GPU statistics of `nvidia-smi`.

The callback queries `nvidia-smi --query-gpu` for each GPU of the trainer at the start and at the end of a train batch. It logs
the values with the logger of the trainer, at `trainer.global_step`. It logs only on the steps where the trainer updates its logs.
Lightning selects these steps with `log_every_n_steps` of the trainer. The batch hooks run on rank 0 only.

Each flag of
[GPUStatsMonitor.init](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/gpu_stats_monitor.md)
turns on a set of queries:

 * `gpu_utilization`: `utilization.gpu`, in percent. It is the part of the last sample period in which one or more kernels ran on
   the GPU. The callback logs it at the start and at the end of a batch.
 * `memory_utilization`: `memory.used` and `memory.free`, in MiB, and `utilization.memory`, in percent. The last value is the part
   of the last sample period in which the GPU read or wrote its memory. The callback logs them at the start and at the end of a
   batch.
 * `fan_speed`: `fan.speed`, in percent of the maximum speed. It is the intended speed, not a measured speed. The callback logs it
   at the end of a batch only.
 * `temperature`: `temperature.gpu` and `temperature.memory`, in degrees Celsius. The callback logs them at the end of a batch
   only.

A metric name has the form `GPU_<device>/<query> - <unit>`, for example `GPU_0/utilization.gpu - percent` or `GPU_0/memory.used -
MB`. `<device>` is the logical device index of the trainer. `<unit>` is `percent`, `MB`, or `°C`. The `MB` label marks a value in
MiB. The callback logs `0.0` for a value that is not a number, such as `[N/A]`.

Two more flags log the time of the batches, in milliseconds:

 * `intra_step_time`: `batch_time/intra_step (ms)`, the time from the start to the end of a batch. The callback logs it at the end
   of a batch.
 * `inter_step_time`: `batch_time/inter_step (ms)`, the time from the end of the previous batch to the start of this batch. The
   callback logs it at the start of a batch. The first batch of an epoch has no value.

Each time includes the `nvidia-smi` queries that run in its interval.

The callback is in the `CALLBACKS` registry, so a config can add it:

```yaml
trainer:
  callbacks:
    - name: GPUStatsMonitor
      params:
        temperature: true
        intra_step_time: true
```

#### Methods

##### init

```python
def __init__(memory_utilization: bool = True, gpu_utilization: bool = True, intra_step_time: bool = False, inter_step_time: bool = False, fan_speed: bool = False, temperature: bool = False):
```

Initialize the callback with the statistics to log.

The class docstring describes each statistic. The constructor checks only for `nvidia-smi`.
[GPUStatsMonitor.setup](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/gpu_stats_monitor.md)
checks for a logger and for CUDA.

Parameters

 * `memory_utilization` (`bool`): Log `memory.used`, `memory.free`, and `utilization.memory` at the start and at the end of a
   batch.
 * `gpu_utilization` (`bool`): Log `utilization.gpu` at the start and at the end of a batch.
 * `intra_step_time` (`bool`): Log the time from the start to the end of a batch as `batch_time/intra_step (ms)`.
 * `inter_step_time` (`bool`): Log the time from the end of a batch to the start of the next batch as `batch_time/inter_step
   (ms)`.
 * `fan_speed` (`bool`): Log `fan.speed` at the end of a batch.
 * `temperature` (`bool`): Log `temperature.gpu` and `temperature.memory` at the end of a batch.

Raises

 * `MisconfigurationException`: When the `nvidia-smi` executable is not on the `PATH`. The message says that the NVIDIA driver is
   not installed.

##### is_available

```python
def is_available() -> bool:
```

Return whether this machine can run the callback.

The constructor and
[GPUStatsMonitor.setup](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/callbacks/gpu_stats_monitor.md)
do not call this method. They do their own checks and raise an error instead.

Returns

 * `bool`: `True` when the `nvidia-smi` executable is on the `PATH` and the Lightning `CUDAAccelerator` reports CUDA as available.

##### on_train_batch_end

```python
def on_train_batch_end(trainer: pl.Trainer, pl_module: pl.LightningModule, outputs: STEP_OUTPUT, batch: Any, batch_idx: int):
```

Log all enabled GPU statistics after a batch.

Lightning calls this hook after every train batch. The hook runs on rank 0 only. When `inter_step_time` is on, the hook first
records the end time of the batch. It stops there when the trainer does not update its logs on this step.

Otherwise the hook queries `nvidia-smi` for the statistics of the enabled flags among `gpu_utilization`, `memory_utilization`,
`fan_speed`, and `temperature`. When `intra_step_time` is on, the hook adds `batch_time/intra_step (ms)`. It logs the values with
`trainer.logger.log_metrics` at `trainer.global_step`.

Parameters

 * `trainer` (`pl.Trainer`): The trainer. Its logger receives the values.
 * `pl_module` (`pl.LightningModule`): The model. Unused.
 * `outputs` (`STEP_OUTPUT`): The output of the training step. Unused.
 * `batch` (`Any`): The batch. Unused.
 * `batch_idx` (`int`): The index of the batch. Unused.

##### on_train_batch_start

```python
def on_train_batch_start(trainer: pl.Trainer, pl_module: pl.LightningModule, batch: Any, batch_idx: int):
```

Log the GPU utilization and memory before a batch.

Lightning calls this hook before every train batch. The hook runs on rank 0 only. When `intra_step_time` is on, the hook first
records the start time of the batch. It stops there when the trainer does not update its logs on this step.

Otherwise the hook queries `nvidia-smi` for the statistics of the enabled flags among `gpu_utilization` and `memory_utilization`.
When `inter_step_time` is on and the previous batch recorded its end time, the hook adds `batch_time/inter_step (ms)`. It logs the
values with `trainer.logger.log_metrics` at `trainer.global_step`.

Parameters

 * `trainer` (`pl.Trainer`): The trainer. Its logger receives the values.
 * `pl_module` (`pl.LightningModule`): The model. Unused.
 * `batch` (`Any`): The batch. Unused.
 * `batch_idx` (`int`): The index of the batch. Unused.

##### on_train_epoch_start

```python
def on_train_epoch_start(trainer: pl.Trainer, pl_module: pl.LightningModule):
```

Clear the recorded batch times.

Lightning calls this hook at the start of every train epoch, on every rank. Because of the reset, the first batch of an epoch logs
no `batch_time/inter_step (ms)` value.

Parameters

 * `trainer` (`pl.Trainer`): The trainer. Unused.
 * `pl_module` (`pl.LightningModule`): The model. Unused.

##### setup

```python
def setup(trainer: pl.Trainer, pl_module: pl.LightningModule, stage: str | None = None):
```

Check the trainer and find the GPUs to query.

Lightning calls this hook at the start of every stage. The hook stores the sorted, unique device indices of `trainer`. It maps
each index to a physical GPU ID through the `CUDA_VISIBLE_DEVICES` environment variable. When the variable is not set, each GPU ID
is equal to its index.

Parameters

 * `trainer` (`pl.Trainer`): The trainer. It must have a logger.
 * `pl_module` (`pl.LightningModule`): The model. Unused.
 * `stage` (`str | None`): The stage that starts. Unused.

Raises

 * `MisconfigurationException`: When `trainer` has no logger, or when CUDA is not available.
