# annotate_utils

Python API: `luxonis_train.core.utils.annotate_utils`

Pre-annotation of a directory of images with a trained model, into a new dataset.

## Functions

### annotate_from_directory

```python
def annotate_from_directory(model: lxt.LuxonisModel, img_paths: Iterable[PathType], dataset_name: str, bucket_storage: Literal['local', 'gcs'] = 'local', delete_local: bool = True, delete_remote: bool = True, team_id: str | None = None) -> LuxonisDataset:
```

Annotate image files with a model into a new dataset.

The function runs these steps:

 * It builds a loader over `img_paths` with
   [create_loader_from_directory](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md),
   with a batch size of `1`. That function replaces a local dataset named `infer_from_directory`.
 * It creates the dataset `dataset_name` and adds the records of
   [annotated_dataset_generator](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/annotate_utils.md)
   to it.
 * It splits a non-empty dataset into `train`, `val`, and `test` in the ratio 0.8, 0.1, and 0.1. For an empty dataset, it logs a
   warning.
 * It deletes the local copy of the `infer_from_directory` dataset.

The function does not load weights.
[LuxonisModel.annotate](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/core.md)
loads them before the call.

Parameters

 * `model` (`lxt.LuxonisModel`): The model that predicts the annotations. The loader applies its `trainer.preprocessing`.
 * `img_paths` (`Iterable[PathType]`): The image files to annotate.
 * `dataset_name` (`str`): The name of the new dataset.
 * `bucket_storage` (`Literal['local', 'gcs']`): The storage backend of the new dataset.
 * `delete_local` (`bool`): Delete the local files of an existing dataset named `dataset_name` before the function creates the new
   dataset. With `False` and `"local"` storage, the records go into the existing dataset.
 * `delete_remote` (`bool`): Delete the remote files of an existing dataset named `dataset_name`. The value has an effect only
   with `"gcs"` storage.
 * `team_id` (`str | None`): The team that owns the dataset. `None` reads `LUXONISML_TEAM_ID` from the environment.

Returns

 * `LuxonisDataset`: The new dataset with the annotations.

### annotated_dataset_generator

```python
def annotated_dataset_generator(model: lxt.LuxonisModel, loader: torch_data.DataLoader) -> DatasetIterator:
```

Yield the dataset records that the heads of a model predict.

The generator puts the Lightning module of `model` in eval mode. For each batch of `loader`, it runs
[LuxonisLightningModule.full_forward](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/lightning/luxonis_lightning.md)
without gradients. For each output node that is a
[BaseHead](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/base_head.md),
it calls
[BaseHead.annotate](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/nodes/heads/base_head.md).
The call gets the outputs of the head, the image paths from the `"path"` sample metadata, and the `trainer.preprocessing` of the
config. The generator skips the other output nodes.

A record from `annotate` that is a dictionary becomes a `DatasetRecord`. When the only validation error of a record is a bounding
box outside the clipping range, the generator skips the record and logs a debug message.

Parameters

 * `model` (`lxt.LuxonisModel`): The model that predicts the annotations.
 * `loader` (`torch_data.DataLoader`): A loader whose batches hold the inputs, the labels, and a list with the metadata of each
   sample, as
   [create_loader_from_directory](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/luxonis-train-api-reference/core/utils/infer_utils.md)
   builds with `return_sample_metadata=True`.

Returns

 * `DatasetIterator`

Yields

 * The records that the heads give for the images.

Raises

 * `ValidationError`: When a record fails the validation of `DatasetRecord` for another reason.
