488
J. Eur. Opt. Society-Rapid Publ. 22, 49( 2026)
3.9, and PyTorch framework The models were optimized using the SGD optimizer with a batch size of 16, an initial learning rate of 0.01, a learning rate decay of 0.0005, and an input image size set to 640 640 pixels.
We select six representative object detection architectures as baselines. The proposed synthetic UAV swarm dataset is divided into non-overlapping training, validation, and test subsets. The specific partition is as follows:
Training set: 4940 images( 70 %), used for learning detector parameters.
Validation set: 706 images( 10 %), used for model selection and hyperparameter tuning.
Test set: 1412 images( 20 %), reserved for final performance evaluation.
4.2 Evaluation metrics
In order to validate the performance of the model, the following performance metrics were selected for measurement: P( precision), R( recall), mean average precision( mAP 50), mAP 50 – 95, parameters, and Model Size. True Positive( TP) denotes the number of correctly detected targets, False Positive( FP) indicates the number of backgrounds detected as targets, and False Negative( FN) denotes the number of targets detected as backgrounds.
P denotes the proportion of positive samples correctly classified by the model, reflecting the ability of the model to correctly classify. The calculation is as show in equation( 11).
TP P ¼ ð11Þ
TP þ FP
R denotes the ratio of the number of correctly predicted samples to the total number of samples, reflecting the ability of the model to detect the target comprehensively, and measuring the number of samples that the model misses in identifying the target. The calculation is as show in equation( 12).
TP R ¼
TP þ FN: ð12Þ
The mAP is the average accuracy of all targets detected by the model, reflecting the ability to generate predictive frames and labels that overlap. The higher the value of this metric, the better the model’ s detection effect on different categories. The calculation is as show in equation( 13).
X n mAP ¼ 1 n AP ð13Þ i¼1
where n denotes the average accuracy for each category. In our UAV detection task, n = 1. mAP 50 denotes the average accuracy when the IOU( Intersection over Union) is 50 %. mAP 50 – 95 denotes that the threshold of mAP ranges from 50 % to 95 %, and the average of 10 mAP values is obtained at 5 % intervals.
Model Size is used to evaluate the complexity of the model. In general, the smaller the Model Size is, the less computing power the model requires, the lower the performance requirements for hardware, and the easier it is to build in low-end devices.
4.3 Experimental results
Among all evaluated models, YOLOv13 achieves the best overall performance, attaining the highest Precision( 0.903), Recall( 0.897), mAP 50( 0.934), and mAP 50 – 95( 0.612). YOLOX ranks second with a mAP 50 of 0.912, followed by YOLOv6( mAP 50 = 0.901) and YOLOv12( mAP 50 = 0.887). Among the non-YOLO detectors, RT- DETR achieves a mAP 50 of 0.889, comparable to the mid-range YOLO variants, while Faster R-CNN yields the lowest accuracy( mAP 50 = 0.856), likely due to the limited effectiveness of its region proposal mechanism on predominantly small UAV targets. Overall, the four one-stage YOLO-based detectors consistently outperform both Faster R-CNN and RT-DETR in detection accuracy on this dataset.
A noteworthy observation is the significant drop in mAP 50 – 95 across all detectors, with values ranging from 0.487( Faster R-CNN) to 0.612( YOLOv13), representing a decrease of approximately 30 – 35 percentage points relative to their mAP 50 counterparts. This indicates that precise bounding box regression for small-scale UAV targets remains challenging across all evaluated architectures.
In terms of computational efficiency, the four YOLO variants achieve inference speeds of 75 – 105 FPS with compact model sizes( 17.3 – 38.0 MB), making them suitable for real-time deployment. RT-DETR operates at 48 FPS with a 64.0 MB model, while Faster R-CNN is the most resourceintensive, with only 22 FPS and a 160.0 MB model size.
Figure 7 illustrates qualitative detection results of four YOLO-based detectors across six representative scene environments. These scenes cover a wide range of background complexities, including dense vegetation, complex cloud distributions, low-texture sky backgrounds, urban street environments, rural landscapes, and snow-covered mountainous areas. Such diversity enables a comprehensive evaluation of detector robustness under realistic and challenging conditions. Note that the qualitative visualization focuses on the four YOLO-based detectors, as all quantitative comparisons including Faster R-CNN and RT-DETR are comprehensively presented in Table 3.
As shown in the figure, all detectors are capable of detecting UAV targets in relatively simple scenes. However, noticeable performance differences emerge in more challenging scenarios. In the Clear Sky and Snowy Mountain scenes, UAVs appear as extremely small objects against low-texture backgrounds, making detection particularly difficult. In these cases, missed detections and low-confidence predictions are frequently observed, especially for earlier-generation models. In contrast, more recent YOLO variants demonstrate improved robustness, benefiting from enhanced feature representation and multi-scale feature aggregation.
Overall, the quantitative results in Table 3 are consistent with the qualitative findings in Figure 7, collectively confirming that YOLOv13 offers the most robust detection performance across diverse and challenging scene conditions, while also highlighting that small object localization accuracy( as reflected by mAP 50 – 95) remains an open challenge for all current detectors.