|
Description
|
## Motivation and scope Motivation: controlled study of domain shift, domain robustness, classification and segmentation; this is not a substitute for natural data.
## Composition Six domains (clean, rotation, geometry, noise, clipping, composite); five classes (circle, octagon, rectangle, square, triangle); 51,000 64*64 grayscale PNG images and 51,000 binary L-mode PNG masks. Per domain: 5,000 train, 1,000 validation and 2,500 test images.
## Generation and splits Images are reproducibly generated controlled 2D shape rasters. Metadata schema version 2.1 records rendering parameters and file linkage. Train/validation/test splits are independent within each domain.
## Intended uses and limitations Suitable for controlled classification, segmentation, invariance and educational experiments. It has no natural textures, camera model, complex scenes or external validity for natural imagery. Strong clipping can reduce semantic identifiability. (2026-08-11)
|
|
Notes
| Exact SHA-256 screening found 143 cross-domain pixel-identical groups: 41 expected parameter collisions and 102 expected symmetry cases; no group has different class labels. Fifty groups span train/test across domains. Strict cross-domain benchmarks must use CROSS_DOMAIN_DUPLICATES.csv as an exclusion/filter list and report the applied rule. |