JEOS RP ISSN03 | Page 433

426 J. Eur. Opt. Society-Rapid Publ. 22, 43( 2026)
TP P ¼ ð7Þ
TP þ FP
Recall R represents the proportion of correctly predicted positive samples to the total number of actual positive samples, reflecting the model’ s comprehensive target detection capability, and can be used to measure the missing detection rate of the model in target recognition tasks. Its calculation formula is as follows equation( 8).
TP R ¼ ð8Þ
TP þ FN
The mean Average Precision( mAP) is the average precision of the model for all detected targets [ 22 ], which reflects the model’ s ability to generate prediction boxes with overlapping regions matching the labels. A higher value of this metric indicates a better detection performance of the model for targets of different categories, and its calculation formula is as follows equation( 9).
mAP ¼ 1 n
X n
i¼1
AP ð9Þ
where n denotes the number of categories for average precision calculation. In the UAV detection task of this study, n = 1. mAP 50 represents the mean average precision when the Intersection over Union threshold is set to 50 %; mAP 50�95 refers to the metric obtained by gradually adjusting the IoU threshold from 50 % to 95 % with a step size of 5 % and averaging the 10 average precision values obtained within this interval.
The model size( Model Size) is used to evaluate the complexity of the model. Generally speaking, the smaller the model size, the less computing power it requires and the lower the hardware performance requirements, making it easier to deploy on low-end devices.
3.3 Ablation experiments
On the self-built infrared thermal imaging dataset, performance verification was conducted for the enhancement modules of the YOLO-KMM model [ 23 ]. Table 1 compares the baseline BASE model with its improved versions integrated with different modules( C2KD, C3K2-MU). The baseline BASE model achieves a mAP 50 of 80.4 % and a mere mAP 50�95 of 42.5 % on this dataset, with Precision( P) and Recall( R) reaching 86.9 % and 71.4 % respectively. After introducing the C2KD module, the model’ smAP 50 increases to 82.8 %, mAP 50�95 rises to 45.3 %, P synchronously climbs to 91.9 %, and R slightly improves to 72.3 %. By further stacking the C3K2-MU module, the mAP 50 exceeds 84.1 %, mAP 50�95 grows to 47.2 %, while P and R reach 92.6 % and 72.8 % respectively. The final YOLO-KMM model integrating all modules achieves the optimal performance: mAP 50 hits 88.2 %, mAP 50�95 increases to 48.4 %, with P and R reaching 94.0 % and 74.3 % respectively. In addition, the prediction heatmaps of different model versions reveal that: the BASE model has weak heatmap focusing ability and low confidence for some targets; after adding the C2KD module, the heatmap focuses more on target regions and the confidence scores are significantly improved, indicating that this module enhances the model’ s target recognition and focusing capability. With the C3K2-MU module stacked, the target localization accuracy of the heatmap is further improved and the confidence distribution is more stable, verifying the module’ s optimization effect on localization accuracy. The heatmap of YOLO-KMM presents the clearest target boundaries and the highest confidence, confirming the effectiveness of all enhancement modules in improving the model’ s detection performance, Figure 8 shows the improvement effects of each module.
3.4 Performance comparison
On the self-built dataset, comparative experiments were conducted between the proposed YOLO-KMM model and mainstream YOLO series models [ 24 – 26 ] to verify its performance advantages, as shown in Figure 9, theresultsofeach comparison model are presented, and the core indicators of each model are presented in Table 2. Among them, the YOLO-KMM model achieved outstanding detection performance: The Figures 10 and 11 provide a detailed comparison of FPS and mAP 50 across different models, its mAP 50 reached 88.2 %, which was 7.8 % points higher than that of the same-level YOLOv11( 80.4 %). Meanwhile, the inference frame rate( FPS) of YOLO-KMM reached 246.18, with a single inference time of only about 4.06 ms, achieving a good balance between detection accuracy and real-time performance. In terms of the balance between precision and recall, the precision of YOLO-KMM reached 94.0 % and the recall was 74.3 %. Compared with YOLOv11( precision 86.9 %, recall 71.4 %) [ 27 ], both detection accuracy and target coverage were significantly optimized. It can also be seen from the indicator distribution that the number of parameters( 2.3 M) and computation( 5.4 GFLOPs) of YOLO-KMM were at a low level, indicating smaller computation and parameter scale. These results show that the optimization strategies introduced in YOLO- KMM effectively enhance the feature extraction and target recognition capabilities on the premise of controlling the model scale and computing cost. Its characteristics of high precision, lightweight and fast speed make it more suitable for deployment scenarios with limited computing resources such as UAV inspection and embedded devices [ 28 ], and also prove the adaptability of the constructed dataset.
The lightweight architecture of YOLO-KMM and efficient inference speed strongly indicate its feasibility for deployment on resource- constrained platforms such as NVIDIA Jetson series devices. The optimization strategies in the C3K2-MU module, including channel pruning and depthwise separable convolutions, are specifically designed to facilitate efficient execution on embedded systems. The computational complexity ratio of C = N demonstrates that TSSA achieves approximately 80 – 95 % complexity reduction compared to standard self-attention, which directly translates to reduced memory footprint on edge devices. While comprehensive performance quantification on specific edge platforms remains a valuable direction for future work, the model’ s design principles align with successful