# yolo

Python API: `depthai_nodes.node.parsers.utils.yolo`

## Classes

### YOLOSubtype

#### Attributes

##### DEFAULT

##### GOLD

##### P

##### V10

##### V11

##### V12

##### V26

##### V3

##### V3T

##### V3U

##### V3UT

##### V4

##### V4T

##### V5

##### V5U

##### V6

##### V6R1

##### V6R2

##### V7

##### V8

##### V9

## Functions

### compute_yolo_detections

```python
def compute_yolo_detections(*, subtype: YOLOSubtype, layer_names: list[str], outputs_values: list[np.ndarray], strides: list[int] | None = None, conf_threshold: float, n_classes: int, iou_threshold: float, max_det: int, anchors: np.ndarray | None, n_keypoints: int, label_names: list[str] | None, keypoint_label_names: list[str] | None, keypoint_edges: list[tuple[int, int]] | None, input_shape: tuple[int, int] | None = None, kpts_outputs: list[np.ndarray] | None = None, masks_outputs_values: list[np.ndarray] | None = None, protos_output: np.ndarray | None = None, protos_len: int | None = None, mask_conf: float = 0.5, v26_mask_coeffs: np.ndarray | None = None, v26_protos: np.ndarray | None = None, v26_pose_kpts: np.ndarray | None = None) -> dict[str, np.ndarray | list[str] | int | None]:
```

Decode YOLO detection, pose, or segmentation tensors.

Parameters

 * `subtype` (`YOLOSubtype`): YOLO variant controlling tensor decoding.
 * `layer_names` (`list[str]`): Output layer names used to distinguish detection, pose, and segmentation modes.
 * `outputs_values` (`list[np.ndarray]`): Detection tensors ordered by output head.
 * `strides` (`list[int] | None`): Output head strides. If omitted, use the subtype-specific defaults.
 * `conf_threshold` (`float`): Minimum detection confidence used to filter candidates.
 * `n_classes` (`int`): Number of object classes encoded in the detection tensors.
 * `iou_threshold` (`float`): Intersection-over-union threshold for non-maximum suppression.
 * `max_det` (`int`): Maximum retained YOLO26 detections. Other subtypes use the suppression defaults in `decode_yolo_output`.
 * `anchors` (`np.ndarray | None`): Precomputed anchor coordinates used to decode model predictions.
 * `n_keypoints` (`int`): Number of keypoints encoded per prediction.
 * `label_names` (`list[str] | None`): Optional class-name lookup indexed by predicted class ID.
 * `keypoint_label_names` (`list[str] | None`): Optional names for the keypoints in each detection.
 * `keypoint_edges` (`list[tuple[int, int]] | None`): Optional pairs of keypoint indexes defining skeleton edges.
 * `input_shape` (`tuple[int, int] | None`): Model input image shape as `(height, width)`.
 * `kpts_outputs` (`list[np.ndarray] | None`): Per-head pose tensors for non-YOLO26 models.
 * `masks_outputs_values` (`list[np.ndarray] | None`): Mask coefficient tensors ordered to match the detection heads.
 * `protos_output` (`np.ndarray | None`): Batched prototype tensor with shape `(1, channels, height, width)`.
 * `protos_len` (`int | None`): Number of prototype channels used by each mask coefficient vector.
 * `mask_conf` (`float`): Probability threshold used to binarize mask logits.
 * `v26_mask_coeffs` (`np.ndarray | None`): YOLO26 mask coefficients aligned with detection queries.
 * `v26_protos` (`np.ndarray | None`): YOLO26 prototype masks.
 * `v26_pose_kpts` (`np.ndarray | None`): YOLO26 pose coordinates and confidences aligned with detection queries.

Returns

 * `dict[str, np.ndarray | list[str] | int | None]`: A dictionary containing `mode` (0 detection, 1 pose, 2 segmentation),
   normalized `bboxes`, `scores`, `labels`, `label_names`, `keypoints`, `keypoints_scores`, `keypoint_label_names`,
   `keypoint_edges`, and `masks`. Boxes use center-XY/width/height. A segmentation mask contains int16 detection indexes and -1
   background; other modes return `None` for masks.

Raises

 * `ValueError`: If required YOLO26 input geometry or detection outputs are missing, class/keypoint counts disagree with tensor
   shapes, or the mask instance count exceeds int16 capacity.

### decode_yolo26

```python
def decode_yolo26(raw: np.ndarray, conf_threshold: float, max_det: int, extra_raw: np.ndarray | None = None) -> tuple[np.ndarray, np.ndarray | None]:
```

Decode YOLO26 output for detection, segmentation, or pose.

YOLO26 end2end output is already decoded to xyxy pixel boxes and includes a pre-computed confidence score (ReduceMax over class
scores) in column 4. This path only applies confidence thresholding and top-k filtering. It can also filter an auxiliary tensor
such as mask coefficients or keypoints using the kept rows.

Parameters

 * `raw` (`np.ndarray`): Raw detection tensor (N, A, 5+nc) where columns are [x1, y1, x2, y2, conf, cls_0, ..., cls_nc-1].
 * `conf_threshold` (`float`): Confidence threshold.
 * `max_det` (`int`): Maximum number of detections.
 * `extra_raw` (`np.ndarray | None`): Optional auxiliary tensor (N, A, M) such as mask coefficients or keypoints. When provided
   the kept rows are returned as the second element.

Returns

 * `tuple[np.ndarray, np.ndarray | None]`: Tuple of (detection results (K, 6), kept auxiliary data (K, M) or None).

### decode_yolo_output

```python
def decode_yolo_output(yolo_outputs: list[np.ndarray], strides: list[int], anchors: np.ndarray | None = None, kpts: list[np.ndarray] | None = None, conf_thres: float = 0.5, iou_thres: float = 0.45, num_classes: int = 1, det_mode: bool = False, subtype: YOLOSubtype = YOLOSubtype.DEFAULT, max_nms: int = 3000) -> np.ndarray:
```

Decode the output of an YOLO instance segmentation or pose estimation model.

Parameters

 * `yolo_outputs` (`list[np.ndarray]`): List of YOLO outputs.
 * `strides` (`list[int]`): List of strides.
 * `anchors` (`np.ndarray | None`): An optional array of anchors.
 * `kpts` (`list[np.ndarray] | None`): An optional list of keypoints.
 * `conf_thres` (`float`): Confidence threshold.
 * `iou_thres` (`float`): Intersection over union threshold.
 * `num_classes` (`int`): Number of classes.
 * `det_mode` (`bool`): Detection only mode. If True, the output will only contain bbox detections.
 * `subtype` (`YOLOSubtype`): YOLO version.
 * `max_nms` (`int`): Maximum number of boxes to keep after NMS.

Returns

 * `np.ndarray`: NMS output.

### make_grid_numpy

```python
def make_grid_numpy(ny: int, nx: int, na: int) -> np.ndarray:
```

Create a grid of shape (1, na, ny, nx, 2)

Parameters

 * `ny` (`int`): Number of y coordinates.
 * `nx` (`int`): Number of x coordinates.
 * `na` (`int`): Number of anchors.

Returns

 * `np.ndarray`: Grid.

### non_max_suppression

```python
def non_max_suppression(prediction: np.ndarray, conf_thres: float = 0.5, iou_thres: float = 0.45, classes: list | None = None, num_classes: int = 1, agnostic: bool = False, multi_label: bool = False, max_det: int = 300, max_time_img: float = 0.05, max_nms: int = 30000, max_wh: int = 7680, kpts_mode: bool = False, det_mode: bool = False) -> list[np.ndarray]:
```

Performs Non-Maximum Suppression (NMS) on inference results.

Parameters

 * `prediction` (`np.ndarray`): Prediction from the model, shape = (batch_size, boxes, xy+wh+...)
 * `conf_thres` (`float`): Confidence threshold.
 * `iou_thres` (`float`): Intersection over union threshold.
 * `classes` (`list | None`): For filtering by classes.
 * `num_classes` (`int`): Number of classes.
 * `agnostic` (`bool`): Runs NMS on all boxes together rather than per class if True.
 * `multi_label` (`bool`): Multilabel classification.
 * `max_det` (`int`): Limiting detections.
 * `max_time_img` (`float`): Maximum time for processing an image.
 * `max_nms` (`int`): Maximum number of boxes.
 * `max_wh` (`int`): Maximum width and height.
 * `kpts_mode` (`bool`): Keypoints mode.
 * `det_mode` (`bool`): Detection only mode. If True, the output will only contain bbox detections.

Returns

 * `list[np.ndarray]`: An array of detections. If det_mode is False, the detections may include kpts or segmentation outputs.

### parse_kpts

```python
def parse_kpts(kpts: np.ndarray, n_keypoints: int, img_shape: tuple[int, int]) -> list[tuple[float, float, float]]:
```

Parse keypoints.

Parameters

 * `kpts` (`np.ndarray`): Result keypoints.
 * `n_keypoints` (`int`): Number of keypoints.
 * `img_shape` (`tuple[int, int]`): Image shape of the model input in (height, width) format.

Returns

 * `list[tuple[float, float, float]]`: Parsed keypoints.

### parse_yolo_output

```python
def parse_yolo_output(out: np.ndarray, stride: int, num_outputs: int, anchors: np.ndarray | None = None, head_id: int = -1, kpts: np.ndarray | None = None, det_mode: bool = False, subtype: YOLOSubtype = YOLOSubtype.DEFAULT) -> np.ndarray:
```

Parse a single channel output of an YOLO model.

Parameters

 * `out` (`np.ndarray`): A single output of an YOLO model for the given channel.
 * `stride` (`int`): Stride.
 * `num_outputs` (`int`): Number of outputs of the model.
 * `anchors` (`np.ndarray | None`): Anchors for the given head.
 * `head_id` (`int`): Head ID.
 * `kpts` (`np.ndarray | None`): A single output of keypoints for the given channel.
 * `det_mode` (`bool`): Detection only mode.
 * `subtype` (`YOLOSubtype`): YOLO version.

Returns

 * `np.ndarray`: Parsed output.

### resolve_yolo_strides

```python
def resolve_yolo_strides(strides: list[int] | tuple[int, ...] | None, subtype: YOLOSubtype, num_outputs: int) -> list[int]:
```

Resolve YOLO strides from metadata or default fallback.

Parameters

 * `strides` (`list[int] | tuple[int, ...] | None`): Optional strides from NNArchive head metadata.
 * `subtype` (`YOLOSubtype`): YOLO subtype.
 * `num_outputs` (`int`): Number of YOLO output heads.

Returns

 * `list[int]`: Resolved YOLO strides.

## Attributes

### logger
