482
J. Eur. Opt. Society-Rapid Publ. 22, 49( 2026)
time-consuming and prone to human error [ 9 ]. These factors make it difficult to obtain sufficiently large and diverse labeled datasets for training and evaluating modern deep learning-based detectors. To alleviate these limitations, we construct a synthetic UAV swarm detection dataset using a controllable simulation pipeline. The dataset provides rich variations in swarm configurations and environmental conditions while ensuring precise, automatically generated annotations. It can serve as a primary training source or as an auxiliary domain for methods combining real and synthetic data. The dataset and generation pipeline are publicly available at https:// github. com / marisinpiper / Synthetic-UAV-Swarm-Dataset. The main contributions of this work can be summarized as follows:
We construct a large-scale synthetic UAV swarm dataset comprising 7000 high-resolution images( 1920 1080) with pixel-accurate annotations, featuring systematic variations in swarm density, formation patterns, target scales, and environmental conditions.
We present a controllable simulation pipeline that enables precise 6-DOF pose specification for UAV swarms, effectively bypassing the difficulty of physical swarm coordination and the prohibitive cost of manual annotation.
We benchmark representative deep learning detectors on the proposed dataset and conduct cross-dataset experiments, demonstrating the effectiveness of synthetic data for UAV swarm detection.
2 Related work
2.1 Vision-based UAV detection
With the rapid proliferation of commercial and recreational drones, anti-UAV technologies have attracted increasing attention from both academia and industry. Existing systems typically comprise three key components – detection, tracking, and identification or classification – implemented using heterogeneous sensing modalities such as radio frequency( RF) signals, radar, acoustics, infrared( IR) and visible-light cameras, or their combinations in multi-sensor fusion frameworks. Among these modalities, vision-based approaches have become particularly prominent due to their relatively low deployment cost, rich semantic information, and compatibility with modern deep learning techniques.
Recent surveys provide comprehensive overviews of vision-based and multi-modal anti-UAV methods, and consistently highlight the challenges posed by small, fast-moving aerial targets in complex backgrounds. Wang et al. [ 10 ] systematically review vision-based anti-UAV techniques, summarizing deep-learning-based detection and tracking frameworks and emphasizing the difficulty of accurately localizing tiny UAV targets in cluttered scenes. Dong et al. [ 11 ] present a broader survey of anti-UAV systems, covering optical, radar, RF, infrared, acoustic, and multimodal fusion methods, and benchmarking representative approaches on public datasets.
In the vision domain, many works adapt generic object detectors such as Faster R-CNN, SSD and YOLO [ 12 – 14 ] to UAV detection tasks by introducing multi-scale feature fusion, feature pyramid networks, or attention mechanisms to better handle small targets. However, most existing benchmarks focus on single or sparsely distributed UAVs, and only a few recent studies consider more challenging scenarios such as dense swarms or multiple coordinated targets. For example, multi-sensor and multi-view datasets like MMFW-UAV provide valuable resources for air-to-air vision tasks involving fixed-wing UAVs, but they mainly target single-platform perception and do not explicitly model swarm formations. Overall, while the literature has made substantial progress on anti-UAV detection, no publicly available dataset explicitly addresses swarm scenarios with controllable formation patterns and systematic scale variations, leaving a critical gap for training robust multitarget detectors.
To provide a clearer picture of the current landscape, Table 1 presents a systematic comparison of representative existing UAV detection datasets alongside the proposed SynthSwarm dataset.
As shown in Table 1, existing UAV detection datasets predominantly focus on single-UAV or sparsely distributed scenarios and lack explicit swarm formation modeling. Notably, all three real-world datasets contain no swarm configurations, and their instance counts equal their image counts, indicating that each image contains at most one UAV target. In contrast, SynthSwarm contains an average of 4.5 UAV instances per image, explicitly modeling multi- UAV swarm formations. Furthermore, SynthSwarm adopts an automatic annotation pipeline, eliminating the time-consuming and error-prone manual labeling process required by all real-world counterparts. These characteristics make SynthSwarm uniquely suited for training and evaluating detectors in dense multi-target aerial surveillance scenarios.
Overall, while the literature has made substantial progress on anti-UAV detection, no publicly available dataset explicitly addresses swarm scenarios with controllable formation patterns and systematic scale variations, leaving a critical gap for training robust multi-target detectors.
2.2 Synthetic data and virtual dataset generation
Data-driven perception systems, especially deep learningbased detectors, typically require large-scale, diverse, and well-annotated datasets to achieve robust performance and generalization. However, as discussed in the Introduction, constructing such datasets for UAV swarm detection is particularly difficult due to safety constraints, regulatory restrictions, and the high cost of collecting and labeling realworld swarm data. These limitations have motivated a growing interest in using synthetic data and virtual environments to complement or partially replace real data for training and evaluation in aerial vision tasks.
Thanks to advances in computer graphics and game engines, highly realistic virtual scenes with controllable environmental conditions, sensor configurations, and objects can now be created at relatively low marginal cost. By leveraging simulation platforms and physically based rendering pipelines, researchers can generate large numbers of annotated images or video sequences while maintaining fine-grained control over object appearance, pose, motion