J. Eur. Opt. Society-Rapid Publ. 22, 49( 2026) 483
Table 1. Comparison of representative UAV detection datasets.
Dataset |
Images |
Instances |
Swarm |
Source |
Det-Fly [ 15 ] |
13,271 |
13,271 |
No |
Real |
DUT Anti-UAV [ 16 ] |
10,000 |
10,109 |
No |
Real |
MMFW-UAV |
147,417 |
147,417 |
No |
Real |
SynthSwarm( Ours) |
7000 |
31,542 |
Yes |
Synthetic |
patterns, background complexity, illumination, and weather. Annotations such as bounding boxes, segmentation masks, depth maps, and optical flow can be obtained automatically and accurately, eliminating the need for time-consuming and error-prone manual labeling.
Several simulation platforms have been widely adopted in autonomous driving and aerial vision research. AirSim [ 17 ], developed by Microsoft, is a high-fidelity simulator specifically designed for aerial and ground vehicle perception, supporting hardware-in-the-loop simulation with customizable environments. UnrealCV [ 18 ] provides a plugin for Unreal Engine that enables the generation of synthetic images with ground-truth depth, surface normal, and semantic segmentation labels. These platforms have demonstrated that synthetic data generated from modern game engines can effectively supplement or even replace real data for training deep learning models, particularly when real data collection is expensive or impractical.
In aerial and remote-sensing applications, several works have demonstrated the effectiveness of synthetic imagery for target recognition and small-object detection. Nasir and Khurshid [ 19 ] propose a multiscale attention generator – discriminator framework to produce synthetic remotesensing aircraft images and show that such synthetic data can significantly improve aircraft recognition performance. Patel et al. [ 20 ] develop a CGI-based synthetic data generation and detection pipeline for small objects in aerial imagery, combining synthetic training data with modern object detectors to enhance drone-based image recognition and small-object detection performance. These studies indicate that carefully designed synthetic aerial datasets can serve both as primary training sources and as auxiliary domains for pre-training or data augmentation, especially when real data are scarce or cover only a limited range of operating conditions.
These studies indicate that carefully designed synthetic aerial datasets can serve both as primary training sources and as auxiliary domains for pre-training or data augmentation, especially when real data are scarce or cover only a limited range of operating conditions.
A key challenge in leveraging synthetic data is the domain gap between virtual and real imagery. Differences in texture fidelity, lighting models, noise characteristics, and motion blur patterns can degrade detector performance when models trained on synthetic data are directly applied to real-world images. Common strategies to mitigate this gap include domain randomization [ 21 ], style transfer [ 22 ], and unsupervised domain adaptation [ 23 ]. In the context of UAV detection, understanding and addressing this domain gap is particularly important, as real-world UAV imagery often exhibits complex atmospheric effects and sensor-specific artifacts that are difficult to fully replicate in simulation.
Despite these advantages, applying synthetic data generation to UAV swarm detection poses several unique challenges. First, realistic modeling of swarm requires coherent control over the relative positions, formations, and motion patterns of multiple UAVs, rather than treating each target as an independent object. Second, to faithfully reflect realworld operating conditions, the virtual environment must capture diverse backgrounds, complex clutter, and various camera viewpoints, including both ground-based and aerial sensors. Third, the synthetic swarm data should include a wide range of target scales and densities, from sparse formations to dense clusters, so that models can learn to detect numerous small UAVs that may occupy only a few pixels in high-resolution images.
In this work, we follow the general paradigm of synthetic data generation for aerial vision, but tailor the simulation and rendering pipeline specifically to the requirements of UAV swarm detection. By explicitly modeling swarm size and inter-UAV spacing, we generate a virtual dataset that emphasizes small-object and multi-target detection in cluttered environments. The automatically generated annotations provide dense, accurate labels for UAV instances in each image, offering a scalable and flexible resource for training, benchmarking, and transferring deep learning-based anti-UAV detectors to real-world swarm scenarios. 3 Dataset generation method
Our training data are generated synthetically in a controllable 3D simulation environment. As illustrated in Figure 1, UAV instances are sampled within a bounded 3D observation volume and projected onto the image plane using a pinhole camera model. For each rendered frame, 2D bounding boxes are obtained automatically from the projected 3D bounds, providing pixel-accurate annotations without manual labeling. This framework enables systematic variation of swarm size, spatial density and camera configuration for dataset generation.
3.1 6-DOF state parameterization
We model each UAV instance as a rigid body with six degrees of freedom( 6-DOF) in three-dimensional Euclidean space. A global right-handed world coordinate frame is defined, and all UAV poses are expressed with respect to