490
J. Eur. Opt. Society-Rapid Publ. 22, 49( 2026)
Table 4. Cross-dataset evaluation results.
Training set Test set mAP 50 mAP 50 – 95
MMFW-UAV |
MMFW-UAV |
0.856 |
0.521 |
SynthSwarm |
MMFW-UAV |
0.743 |
0.412 |
Another typical failure case arises in cluttered environments, such as Woodland and Urban Street, where background structures may cause false positives. These failure cases indicate that small object detection under complex and low-contrast backgrounds remains a challenging problem. Future work will focus on incorporating more diverse training samples and enhancing feature representation to further improve robustness in these scenarios.
4.4 Cross-dataset experiments
To further validate the generalizability of the proposed SynthSwarm dataset, we conduct a cross-dataset evaluation using the MMFW-UAV dataset as an additional test domain. Specifically, the YOLOv13 detector trained on SynthSwarm is directly evaluated on the MMFW-UAV test set without any fine-tuning, and the results are compared against a model trained on MMFW-UAV itself. The results are summarized in Table 4.
As shown in Table 4, the model trained on SynthSwarm achieves a mAP 50 of 0.743 and mAP 50 – 95 of 0.412 on the MMFW-UAV test set without any fine-tuning, compared to 0.856 and 0.521 for the in-domain baseline. While a performance gap is expected due to the inherent domain shift between synthetic rendering and real-world imaging conditions, the SynthSwarm-trained model retains approximately 87 % of the in-domain mAP 50 performance, demonstrating reasonable cross-domain transferability. These results suggest that SynthSwarm captures sufficient visual diversity and structural characteristics of UAV swarm targets to serve as a viable pre-training or standalone training source for real-world UAV detection tasks.
5 Conclusion
In this paper, we proposed a synthetic UAV swarm dataset specifically designed to study small-object, multi-target detection in long-range aerial surveillance scenarios. By leveraging a controllable virtual environment built in Unity, we modeled a fixed sky-facing camera and instantiated multiple 3D UAV models with randomized positions and orientations in a three-dimensional observation volume. This setup enabled the automatic generation of thousands of high-resolution images together with pixel-accurate 2D bounding box annotations, without the need for manual labeling.
We provided a detailed description of the dataset design and generation pipeline, and analyzed the resulting data in terms of target scale, swarm density, and scene diversity. Extensive experiments with six representative detectors spanning three paradigms – one-stage detectors( YOLOX,
YOLOv13, YOLOv12, YOLOv6), a two-stage detector( Faster R-CNN), and a Transformer-based detector( RT- DETR) – demonstrated that the proposed dataset poses significant challenges due to the predominance of small UAVs and the presence of dense swarms. Among all evaluated models, YOLOv13 achieved the best overall performance, while the four one-stage YOLO-based detectors consistently outperformed Faster R-CNN and RT-DETR in both detection accuracy and inference efficiency, suggesting that lightweight one-stage architectures are better suited for this task. The consistent performance trends across detector paradigms indicate that SynthSwarm supports fair and meaningful cross-architecture comparisons. Furthermore, cross-dataset experiments on the MMFW- UAV dataset confirmed that SynthSwarm possesses reasonable transferability to real-world UAV detection scenarios, validating its potential as a pre-training data source.
In future work, we plan to further enrich the dataset by incorporating additional sensor modalities, more diverse weather and illumination conditions, and multiple object categories such as birds or manned aircraft to better capture real-world confusion scenarios. We also intend to investigate specialized detection architectures tailored to extremely small and crowded UAV targets, as well as domain adaptation techniques that more effectively bridge the gap between synthetic and real imagery. We hope that the release of this dataset will foster further research on robust and scalable UAV swarm detection in the computer vision and remote sensing communities.
Acknowledgments
The authors would like to thank the anonymous reviewers for their valuable comments and suggestions that helped improve the quality of manuscript.
Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Conflicts of interest The authors declare that they have no competing interests.
Data availability statement
The dataset and generation pipeline are publicly available at https:// github. com / marisinpiper / Synthetic-UAV-Swarm-Dataset. Author contribution statement
All authors take part in the discussion of the work described in this paper. These authors contributed equally to this work.