# Data Preparation

The default `LuxonisLoaderTorch` reads a
[LuxonisDataset](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-ml/luxonis-dataset.md). It can
open an existing dataset by name or parse a source directory with
[LuxonisParser](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-ml/luxonis-parser.md).

## Data Directory

Point `dataset_dir` at a directory in a supported format, a ZIP file, or a remote dataset source. `dataset_name` names the parsed
LuxonisDataset:

```yaml
loader:
  name: LuxonisLoaderTorch
  params:
    dataset_dir: /path/to/dataset
    dataset_name: my_dataset
    delete_existing: false
```

The parser detects the dataset format. You can set `dataset_type` explicitly when needed. See
[LuxonisParser](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-ml/luxonis-parser.md) for
supported layouts and formats.

When `dataset_dir` is set, `delete_existing` defaults to `true`: an existing local dataset with the same name is deleted and the
source is parsed again. Set it to `false` to reuse an existing dataset without parsing the directory again. After preparing a
dataset, you can omit `dataset_dir` and select it by `dataset_name` alone.

Remote sources are downloaded to `data/` in the working directory before parsing. For example, to load a COCO-format export from
Roboflow:

```yaml
loader:
  params:
    dataset_name: coco_test
    dataset_dir: "roboflow://team-roboflow/coco-128/2/coco"
    delete_existing: false
```

A Roboflow source requires `ROBOFLOW_API_KEY` in the environment. Its URL has the form
`roboflow://workspace/project/version/format`. Cloud paths such as `s3://bucket/path` and `gs://bucket/path` require credentials
and filesystem support for the selected backend.

## Existing LuxonisDataset

For a dataset that is already prepared, configure its name and storage backend:

```yaml
loader:
  params:
    dataset_name: my_dataset
    bucket_storage: local
```

`bucket_storage` accepts `local`, `s3`, `gcs`, or `azure`. For a remote dataset, `update_mode: missing` downloads only media files
that are absent locally; `update_mode: all` downloads all media files again. Annotations and metadata are synchronized in both
modes. Use `team_id` to select the dataset's team, or set `LUXONISML_TEAM_ID` in the environment.

### Dataset Views

The training, validation, and test views select dataset splits. By default, each view uses its matching split. Map them to your
dataset's split names as needed:

```yaml
loader:
  train_view: [train]
  val_view: [validation]
  test_view: [test]
  params:
    dataset_name: my_dataset
```

Each view accepts one or more split names. A CLI option such as `--view val` selects the configured validation view, even when its
underlying split is named `validation`. Keep validation and test data separate from training data when measuring model
performance.

### Tasks and Labels

The loader supplies dataset classes and task metadata to the model. For datasets with multiple tasks, set `task_name` on the
corresponding heads. Use `loader.params.filter_task_names` to load a subset of tasks, `class_order_per_task` to set class order,
or `kpts_mapping_per_task` to remap keypoints.

## Inspect the Dataset

Inspect samples and their labels before training:

```bash
luxonis_train inspect --config config.yaml --view train --list-augmentations
```

You can also inspect data using a packaged model's preprocessing:

```bash
luxonis_train inspect \
  --model detection --variant light \
  --view val \
  loader.params.dataset_name "my_dataset"
```

The command opens a visualization window. Press any key for the next sample, or `q` or `Esc` to close it. Normalization is
disabled for display. `--list-augmentations` lists the tracked augmentations applied to each sample; use `-s 0.5` to display at
half size. This command needs a graphical environment.

Configure image size, normalization, and augmentations under `trainer.preprocessing`, as described in
[Training](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/training.md).

## Custom Loaders

Subclass `BaseLoaderTorch` and implement `input_shapes`, `__len__`, `__getitem__`, and `get_classes`. A loader with keypoint
labels also implements `get_n_keypoints`. Override `collate_fn` if your labels need custom batching.

Set `loader.name` to the subclass name and pass its arguments in `loader.params`. Import the class before creating the model, or
load its file with `--source`, as described in
[Customization](https://docs.luxonis.com/software-v3/ai-inference/model-source/training/luxonis-train/concepts.md).
