Data Module¶
The data module provides dataset loading, image tiling, data augmentation, and evaluation utilities.
DOTA Dataset¶
Label file format (official DOTA: comma-separated)¶
We produce and use the official DOTA format: comma-separated lines:
- Writing:
format_dota_line(),DOTAAnnotation.to_line(),tools/tile_dota.py, andtools/playground_to_dota.pyall output this format. - Reading:
DOTAAnnotation.from_line()accepts both comma-separated (official) and space-separated (legacy). The official comma grammar is used only when the line looks like 8 numeric coords + category + difficult; a space-separated line whose category contains commas (e.g. Airbus) stays on the space path. PreferDOTAAnnotation.from_corners()when building from structured fields.
Loading DOTA Dataset¶
The DOTA loader supports two modes for discovering annotation files:
Mode 1: Pattern Matching (Default)¶
from oriented_det.data import DOTADataset
dataset = DOTADataset(
root_dir="/path/to/dota",
split="train",
allowed_classes=["plane", "ship", "vehicle"],
difficult_strategy="drop"
)
Mode 2: Split File (Official DOTA Convention)¶
dataset = DOTADataset(
root_dir="/path/to/dota",
split="train",
split_file="train.txt", # Lists image names, one per line
difficult_strategy="drop"
)
Mode 3: Separate Folders for Splits¶
If your dataset has train, val, and test in separate folders:
data_root/
├── train/
│ ├── labelTxt/
│ └── images/
├── val/
│ ├── labelTxt/
│ └── images/
└── test/
├── labelTxt/
└── images/
You can specify custom label_dir and image_dir paths:
# Training set
train_dataset = DOTADataset(
root_dir="/path/to/data_root", # Base directory (not used when label_dir/image_dir specified)
split="train",
label_dir="/path/to/data_root/train/labelTxt",
image_dir="/path/to/data_root/train/images",
difficult_strategy="drop"
)
# Validation set
val_dataset = DOTADataset(
root_dir="/path/to/data_root",
split="val",
label_dir="/path/to/data_root/val/labelTxt",
image_dir="/path/to/data_root/val/images",
difficult_strategy="drop"
)
# Or using build_dota_loader
from oriented_det.data import build_dota_loader
train_loader = build_dota_loader(
root_dir="/path/to/data_root",
split="train",
label_dir="/path/to/data_root/train/labelTxt",
image_dir="/path/to/data_root/train/images",
batch_size=4,
shuffle=True
)
Mode 4: Images and annotations in the same folder¶
You can keep images (.jpg/.png) and DOTA annotation files (.txt) in the same directory. Pass that directory as both label_dir and image_dir:
# Single folder per split: images and .txt labels together
train_dataset = DOTADataset(
root_dir="/path/to/train",
split="train",
label_dir="/path/to/train",
image_dir="/path/to/train",
difficult_strategy="drop"
)
For config-based training, set dataset.same_folder: true in your config so that train_tiles_dir and val_tiles_dir are used as the single folder for each split (no images/ or labels/ subdirectories):
{
"dataset": {
"data_root": "/path/to/dota",
"train_tiles_dir": "/path/to/dota/train",
"val_tiles_dir": "/path/to/dota/val",
"same_folder": true,
"difficult_strategy": "drop"
}
}
Annotation files must match image base names (e.g. P0001.png with P0001.txt, or P0001_train.txt with P0001_train.png when using split-suffix pattern matching).
Detailed Loading Examples:
Example 1: Standard DOTA structure with pattern matching
from oriented_det.data import DOTADataset
# Standard DOTA structure:
# /path/to/dota/
# ├── train/
# │ ├── images/
# │ └── labelTxt/
# ├── val/
# │ ├── images/
# │ └── labelTxt/
# └── test/
# ├── images/
# └── labelTxt/
dataset = DOTADataset(
root_dir="/path/to/dota",
split="train",
difficult_strategy="drop",
allowed_classes=["plane", "ship", "small-vehicle"] # Optional: filter classes
)
Example 2: Using split file (official DOTA convention)
# DOTA structure with split files:
# /path/to/dota/
# ├── train.txt # Lists image names: P0001.png, P0002.png, ...
# ├── val.txt
# ├── images/
# └── labelTxt/
dataset = DOTADataset(
root_dir="/path/to/dota",
split="train",
split_file="train.txt", # Explicit split file
difficult_strategy="drop"
)
Example 3: Custom directory structure
# Custom structure with separate folders:
# /mnt/data/
# ├── dota_train/
# │ ├── images/
# │ └── labels/
# └── dota_val/
# ├── images/
# └── labels/
train_dataset = DOTADataset(
root_dir="/mnt/data", # Not used when label_dir/image_dir specified
split="train",
label_dir="/mnt/data/dota_train/labels",
image_dir="/mnt/data/dota_train/images",
difficult_strategy="drop"
)
Edge Case Handling:
- Missing annotation files: Dataset skips images without corresponding annotation files
- Malformed annotations: Lines that can't be parsed are skipped with a warning
- Empty tiles: By default, tiles with no ground-truth objects are included (empty target list). Set
dataset.filter_empty_gt: truein the training config to drop them at dataset init (afterdifficult_strategy,allowed_classes, andignore_labels), matching MMRotateDOTADataset(filter_empty_gt=True). The DOTA pretrain recipesdota_le90_1x.jsonanddota_le90_3x.jsonenable this. To keep empty tiles that the model still false-positives on, leavefilter_empty_gtfalse and setdataset.drop_easy_empty_tiles: truewithdataset.tile_metrics_csvfrommake train-preds: training then drops only vacuous tiles (tp=fp=fn=0) and can oversample hard empties. - Difficult objects: Use
dataset.difficult_strategy: drop: remove difficult objects at read-time (never reach training/eval targets)ignore: keep difficult objects but treat them as “don’t care” (MMRotate/MMDet style)keep: treat difficult objects as normal GT
Performance Considerations:
- Mode 1 (Pattern matching): Fastest, good for standard DOTA structure
- Mode 2 (Split file): Slightly slower due to file reading, but more explicit
- Mode 3 (Separate folders): Most flexible, similar performance to Mode 1
Best Practices:
- For MMRotate parity, use
difficult_strategy="ignore"(difficult objects should not contribute to loss) - Filter classes early with
allowed_classesif you only need specific classes - Use
split_filefor explicit control over train/val/test splits - Set
label_dirandimage_dirfor non-standard directory structures
DOTA Polygon Format¶
DOTA uses 8 coordinates to define quadrilaterals:
Important Notes:
- Corners are ordered sequentially around the polygon perimeter
- The loader converts polygons to QBox (which normalizes point order) and then to RBox
- QBox ensures counter-clockwise orientation and orders points starting from top-most
The loader automatically: 1. Parses polygons from annotation files 2. Converts to QBox (normalizes point order, ensures counter-clockwise) 3. Converts to RBox (computes center, dimensions, angle)
Using with PyTorch DataLoader¶
from oriented_det.data import build_dota_loader
loader = build_dota_loader(
root_dir="/path/to/dota",
split="train",
batch_size=4,
shuffle=True,
num_workers=4
)
Filtering¶
# Filter by class
sample = sample.filter_by_class(allowed_classes=["plane", "ship"])
# Drop difficult annotations on an already-loaded sample
sample = sample.filter_by_class(drop_difficult=True)
Airbus Playground¶
CSV + split-file datasets (dataset.format: airbus_playground in Configuration):
annotations_file— object CSV from Playground exportsplit_file— fold or train/val column (val_split_idfor integer folds)train_includes_val— whentrue, train on all folds;val_split_idfold is still used for validation/monitoring only (DOTAtrain_tiles_dirstrainval parity)ignore_labels,map_labels— filter and rename classes (map_labelsis exact match on the full concatenatedclass_namestring)difficult_tags— exact Playground tags (e.g.["Partially Hidden"]) that set DOTAdifficult=1and are stripped from the semantic class name at load and CSV generation timedifficult_strategy— same as DOTA (drop/ignore/keep). For Partially Hidden don't-care, use"ignore"- Lookalike confusers (hard negatives, not a semantic class) — see Lookalike confusers below
Playground JSON stores multiple tags per object as ", ".join(sorted(tags)). Class names may therefore contain commas (e.g. car, van and pickup). The Airbus loader builds annotations from structured CSV fields (dota_coords + class_name + difficult) and does not drop boxes when the category contains a comma.
Label routing (do not confuse):
| Mechanism | Effect |
|---|---|
ignore_labels |
Drop the box at read-time |
difficult=1 + difficult_strategy: "ignore" |
Don't-care: stays in the sample, routed to rboxes_ignore; no FG/BG/TP/FN/FP |
lookalike / lookalike_labels |
Hard negative: overlapping non-positives forced to background |
Example vehicles config fragment:
"dataset": {
"format": "airbus_playground",
"difficult_strategy": "ignore",
"difficult_tags": ["Partially Hidden"],
"ignore_labels": [],
"map_labels": {
"truck and bus": "truck",
"car, van and pickup": "car"
}
}
Partially Hidden, car → semantic car, difficult=1 (not a lookalike, not dropped). Compounds that remain after stripping must be mapped (or they become extra classes); with allowed_classes set, an unmapped leftover raises instead of silently dropping.
filter_empty_gt keeps tiles that still have don't-care or lookalike boxes (same as lookalike-only tiles). With difficult_strategy: "drop", Partially Hidden boxes are removed and a don't-care-only tile becomes empty and is dropped.
Keep dataset JSON in your own config tree and inherit @odet:configs/_base_/... fragments. Prep tools: odet playground-csv, odet playground-to-dota (see tools/README.md). Pass --difficult-tag 'Partially Hidden' when regenerating CSVs so the CSV difficult column matches the loader.
Lookalike confusers¶
lookalike is a reserved class name. It never enters class_map / num_classes. Boxes mapped onto it are trained as hard negatives: overlapping non-positive RPN/ROI/RetinaNet/FCOS samples are forced to background (not ignore), and two-stage heads preferentially sample those negatives.
This is not the same as DOTA difficult / difficult_tags (e.g. Partially Hidden): lookalike = hard negative (BG); difficult ignore = don't-care (neither FG nor BG).
Recipe (Airbus Playground CSV or any loader that applies map_labels):
"dataset": {
"map_labels": { "Confuser": "lookalike" },
"ignore_labels": [],
"lookalike_labels": null
}
Notes:
- CSV / DOTA labels must still contain the boxes. If they were dropped with
--ignore-label Confuser/ignore_labels, regenerate or clear that filter. - Optional
dataset.lookalike_labelsadds extra aliases treated the same way;"lookalike"is always included. Example without renaming:"lookalike_labels": ["Confuser"]. - Lookalike wins over
ignore_labelsand is kept even when not inallowed_classes. Lookalike-only tiles are not dropped byfilter_empty_gt. - At eval, lookalikes are not positives; a real-class detection on that region counts as an FP (desired). Do not put them on
rboxes_ignoreat eval. - Do not put Partially Hidden on
ignore_labelsor map it tolookalike; usedifficult_tags+difficult_strategy: "ignore".
HRSC2016¶
Native XML loader (dataset.format: hrsc2016).
Download¶
The original paper site (escience.cn) is often offline. Use one of these copies of the official 2016 release (1,061 images: 436 train / 181 val / 444 test):
| Source | Size | Notes |
|---|---|---|
| IEEE DataPort | ~3.5 GB (HRSC2016_dataset.zip) |
Free IEEE account; closest hosted archive |
| Baidu AI Studio | — | Link used by MMRotate; needs a Baidu account |
Kaggle guofeng/hrsc2016 |
— | Same layout; kaggle datasets download -d guofeng/hrsc2016 |
Do not use HRSC2016-MS (a later multi-scale variant). After unzip, point dataset.data_root at the folder that contains FullDataSet/ and ImageSets/ (a wrapping HRSC2016/ directory is also accepted).
Paper: Liu, Yuan, Weng, Yang, A High Resolution Optical Satellite Image Dataset for Ship Recognition and Some New Baselines, ICPRAM 2017. DOI.
Official layout:
HRSC2016/
FullDataSet/AllImages/*.bmp
FullDataSet/Annotations/*.xml
ImageSets/{train,val,test,trainval}.txt
- Single class:
ship(fine-grainedClass_IDvalues are ignored). - XML
mbox_cx/cy/w/h/anguses radians; boxes are converted through the same polygon → RBox path as DOTA (le90). - Default ImageSets mapping (MMRotate): train →
trainval, val →test. Override withdataset.train_split/dataset.val_split. - Oriented R-CNN, Faster R-CNN, and FCOS 1×/3× use
keep_ratio(long edge 800) +pad_size_divisor32. FCOS HRSC 1×/3× and Oriented R-CNN / Faster R-CNN 3× enable random rotate at p=0.5 ±20°. Two-stage 1× recipes leave rotate off. Oriented R-CNN uses Smooth L1 + ProbIoU aux; Faster R-CNN keeps ProbIoU main + Smooth L1 aux.make eval-val/odet predsuse the same whole-image path forpad/keep_ratio(no native sliding windows). DOTA eval-val stays onfixedpre-tiled rasters. - HRSC / DOTA NMS split: train
modeland eval-valevaluation.final_nms_iou_threshold0.1 (MMRotate test parity); deployproduction.final_nms_iou_threshold0.3. Two-stage HRSC keeps max 2000 dets/image and score 0.05. - Optional DOTA export (for
odet tile-dota):odet hrsc-to-dota --data-root /path/to/HRSC2016 --output-dir /path/to/HRSC2016-dota.
{
"dataset": {
"format": "hrsc2016",
"data_root": "/path/to/data/HRSC2016",
"train_split": "trainval",
"val_split": "test"
}
}
Recipes: configs/oriented_rcnn/hrsc2016_le90_1x.json, configs/oriented_rcnn/hrsc2016_le90_3x.json (36 epochs, milestones 24/33, ±20° rotate), configs/rotated_faster_rcnn/hrsc2016_le90_1x.json, configs/rotated_faster_rcnn/hrsc2016_le90_3x.json, configs/rotated_fcos/hrsc2016_le90_1x.json, configs/rotated_fcos/hrsc2016_le90_3x.json.
Image Tiling¶
Split large images into overlapping patches:
from oriented_det.data import ImageTiler
# Create tiler
tiler = ImageTiler(
tile_size=1024,
overlap=0.2, # 20% overlap
min_box_area=64, # Filter small boxes
min_overlap_ratio=0.3, # Keep box if >= 30% overlaps tile
edge_handling="clip" # "clip", "ignore", or "keep"
)
# Generate tiles
tiles = tiler.generate_tiles(image_width=4000, image_height=4000)
# Process each tile
for tiled_sample in tiler.tile_image(
image_path=Path("large_image.png"),
image_width=4000,
image_height=4000,
rboxes=annotations,
class_names=classes
):
# Process tiled_sample
process_tile(tiled_sample)
Visualizing Tiles¶
from oriented_det.data import visualize_tiles
visualize_tiles(
image_path=Path("large_image.png"),
image_width=4000,
image_height=4000,
tiles=tiles,
rboxes=annotations,
class_names=classes,
output_path=Path("tiles_vis.png")
)
Data Augmentation¶
Training collate applies geometric augs after spatial resize: random flips (preprocessing.enable_flip_*, MMRotate RRandomFlip) then optional random rotate (enable_random_rotate, random_rotate_prob, random_rotate_angle_range in degrees — MMRotate PolyRandomRotate, auto_bound=False). Val and inference do not flip or rotate. FCOS HRSC 1×/3× and Oriented R-CNN / Faster R-CNN HRSC 3× use p=0.5 ±20°; two-stage 1× and DOTA leave rotate off.
1. Geometric Transforms (Oriented Bounding Box Aware)¶
Image+box helpers in oriented_det.data (flips.py / rotates.py). Boxes are re-normalized to le90. geometry.transforms.rotate is rbox-only math (y-up); do not use it on PIL images without negating the angle.
import math
from oriented_det.data import (
apply_flip_to_image,
apply_flip_to_rboxes,
apply_random_train_flips,
apply_random_train_rotate,
apply_rotate_to_image,
apply_rotate_to_rboxes,
)
# Same path as training collate
image, rboxes = apply_random_train_flips(
image, rboxes, image_width=512, image_height=512,
enable_horizontal=True, enable_vertical=True, enable_diagonal=True,
)
image, rboxes = apply_random_train_rotate(
image, rboxes, image_width=512, image_height=512,
prob=0.5, angle_range_deg=180.0,
)
# Fixed flip / rotate (PIL visual CCW)
image = apply_flip_to_image(image, "horizontal")
rboxes = apply_flip_to_rboxes(rboxes, "horizontal", image_width=512, image_height=512)
image = apply_rotate_to_image(image, 90.0)
rboxes = apply_rotate_to_rboxes(rboxes, math.radians(90.0), image_width=512, image_height=512)
2. Albumentations (Non-Geometric Only)¶
For additional augmentation options, you can use albumentations with non-geometric transforms. Note: Only non-geometric augmentations (color, contrast, blur, noise, etc.) are supported because albumentations does not support oriented bounding boxes.
from oriented_det.data import create_albumentations_augmentation
# Create default augmentation pipeline
aug = create_albumentations_augmentation(
brightness_limit=0.2,
contrast_limit=0.2,
gamma_limit=(80, 120),
gauss_noise_var_limit=(10.0, 50.0),
blur_limit=3,
clahe_clip_limit=4.0,
p_brightness_contrast=0.5,
p_gamma=0.3,
p_noise=0.2,
p_blur=0.2,
p_clahe=0.3,
)
# Apply to PIL Image
augmented_image = aug(image) # Returns PIL Image
Supported Non-Geometric Augmentations: - Color/Contrast: RandomBrightnessContrast, RandomGamma, CLAHE - Noise: GaussNoise - Blur: GaussianBlur
Not Supported (Geometric):
- Rotation, Scaling, Translation, Affine transforms (use apply_rotate_to_* / apply_random_train_rotate)
- Any transform that would require updating oriented bounding box coordinates
You can also create custom albumentations pipelines:
import albumentations as A
from oriented_det.data import AlbumentationsTransform
# Create custom non-geometric augmentation
custom_aug = A.Compose([
A.RandomBrightnessContrast(p=0.5),
A.RandomGamma(p=0.3),
A.GaussNoise(p=0.2),
A.GaussianBlur(p=0.2),
])
transform = AlbumentationsTransform(custom_aug)
augmented_image = transform(image)
Evaluation¶
Oriented mAP¶
Compute mean Average Precision (mAP) for oriented detection using compute_oriented_map(). This function implements the DOTA evaluation protocol for oriented bounding boxes.
Basic usage:
from oriented_det.data import Detection, GroundTruth, compute_oriented_map
from oriented_det.geometry import RBox
# Prepare detections (from model predictions)
detections = {
"img1": [
Detection(
rbox=RBox(100, 200, 50, 30, 0.5),
score=0.9,
class_id=0,
class_name="plane"
),
Detection(
rbox=RBox(300, 400, 80, 40, 0.2),
score=0.85,
class_id=1,
class_name="ship"
),
],
"img2": [
Detection(
rbox=RBox(150, 250, 60, 35, 0.1),
score=0.75,
class_id=0,
class_name="plane"
),
],
}
# Prepare ground truths (from dataset)
ground_truths = {
"img1": [
GroundTruth(
rbox=RBox(100, 200, 50, 30, 0.5),
class_id=0,
class_name="plane",
difficult=0
),
GroundTruth(
rbox=RBox(310, 410, 75, 45, 0.25),
class_id=1,
class_name="ship",
difficult=0
),
],
"img2": [
GroundTruth(
rbox=RBox(150, 250, 60, 35, 0.1),
class_id=0,
class_name="plane",
difficult=0
),
],
}
# Compute mAP
mean_ap, class_aps = compute_oriented_map(
detections,
ground_truths,
iou_threshold=0.5
)
print(f"mAP: {mean_ap:.4f}")
print("Per-class AP:")
for class_name, ap in class_aps.items():
print(f" {class_name}: {ap:.4f}")
Advanced usage:
import torch
from oriented_det.data import compute_oriented_map
# Use GPU acceleration for large evaluations
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
mean_ap, class_aps = compute_oriented_map(
detections,
ground_truths,
iou_threshold=0.5, # IoU threshold for positive matches
class_names=["plane", "ship", "small-vehicle"], # Evaluate specific classes
show_progress=True, # Show progress bars
max_iou_calculations_per_class=10_000_000, # Optimization threshold
device=device, # GPU acceleration
)
Parameters:
detections: Dictionary mappingimage_id -> List[Detection]- Each
Detectionhas:rbox(RBox),score(float),class_id(int),class_name(str) ground_truths: Dictionary mappingimage_id -> List[GroundTruth]- Each
GroundTruthhas:rbox(RBox),class_id(int),class_name(str),difficult(int) iou_threshold: IoU threshold for positive matches (default: 0.5)class_names: Optional list of class names to evaluate (default: all classes)show_progress: Whether to show progress bars (default: True)max_iou_calculations_per_class: Maximum IoU calculations before using optimization (default: 10M)device: Optional torch device for GPU acceleration
Returns:
mean_ap: Mean Average Precision across all classes (float)class_aps: Dictionary mappingclass_name -> APfor each class
Integration with model evaluation:
from oriented_det.data import Detection, GroundTruth, compute_oriented_map
from oriented_det.geometry import RBox
# After running inference
model.eval()
detections = {}
ground_truths = {}
for image_id, (image, target) in enumerate(val_loader):
with torch.no_grad():
outputs = model([image])
# Convert model outputs to Detection objects
output = outputs[0]
detections[image_id] = [
Detection(
rbox=rbox,
score=float(score),
class_id=int(label),
class_name=class_names[int(label) - 1] # 1-indexed labels
)
for rbox, score, label in zip(
output["rboxes"],
output["scores"],
output["labels"]
)
]
# Convert ground truth to GroundTruth objects
ground_truths[image_id] = [
GroundTruth(
rbox=rbox,
class_id=int(label),
class_name=class_names[int(label) - 1],
difficult=0
)
for rbox, label in zip(target["rboxes"], target["labels"])
]
# Compute mAP
mean_ap, class_aps = compute_oriented_map(
detections,
ground_truths,
iou_threshold=0.5
)
Performance Notes:
- For large evaluations (>10M IoU calculations per class), the function automatically uses batch IoU computation
- GPU acceleration is available when
deviceis set to a CUDA device - Progress bars show computation status for each class
- Difficult objects are included in evaluation (set
difficult=1in GroundTruth to mark them)
DOTA Protocol Compatibility:
- Uses oriented IoU for matching (not axis-aligned IoU)
- Follows DOTA evaluation protocol
- Supports multiple IoU thresholds (call multiple times with different thresholds)
- Handles class-agnostic evaluation (set
class_names=None)
See Also¶
- API Reference - Complete API documentation
- Models Guide - Using data with models