JEOS RP ISSN03 | Page 427

420
J. Eur. Opt. Society-Rapid Publ. 22, 43( 2026)
target size is defined as the number of pixels in the annotated bounding box. The results show that the average width of targets in the dataset is only 11.2 pixels and the average height is only 6.6 pixels [ 13 ]; moreover, ultra-small targets with a width of less than 40 pixels account for as high as 99.7 %. This indicates that the dataset is dominated by ultra-small targets, which is highly consistent with the characteristics of long-distance and low pixel occupancy of UAV targets in actual air-to-air detection scenarios, and also highlights the arduousness of object detection tasks in such scenarios.
Weather Characteristic Analysis: The statistical results of sample quantities under different weather conditions are presented in Figure 2: Sunny days account for 51.3 % of the dataset with 2043 images, cloudy days for 19.7 % with 787 images, snowy days for 15.5 % with 620 images, and hazy days for 13.5 % with 543 images. This distribution covers both normal and severe weather conditions [ 14 ], which can effectively verify the robustness of the model in complex atmospheric environments.
Target Signal-to-Noise Ratio Distribution quantify the difficulty of distinguishing infrared targets from the background under different weather conditions, the grayscale mean values and signal-to-noise ratio( SNR) of target and background regions were calculated. The SNR calculation formula is defined as follows: SNR ¼ jl t�l b j r b
, where l t denotes the grayscale mean of the target region, l t represents the grayscale mean of the background region, and r b stands for the grayscale standard deviation of the background region.
Under cloudy conditions, the grayscale difference between targets and the background is the largest, with an average target grayscale of 146.7 and an average background grayscale of 122.8, corresponding to the highest SNR of 0.42, which makes target features the easiest to identify. Under sunny conditions, the target grayscale of 107.0 is slightly higher than the background grayscale of 96.5, with an SNR of 0.21, indicating that targets have a certain degree of distinguishability. Under hazy conditions, the target grayscale of 90.7 is close to the background grayscale of 96.6, and the SNR is as low as �0.09, meaning that targets are prone to being submerged by background noise. Snowy conditions exhibit a special reverse contrast characteristic: the average target grayscale of 40.0 is significantly lower than the average background grayscale of 59.6, with an SNR of �0.30. In such scenarios, targets become dark targets because their thermal radiation is weaker than that of the low-temperature background, which further increases the detection difficulty. The comparison between the target and the environment is shown in Figure 3.
The above results directly reflect the impact of different weather conditions on infrared target detection. Among them, snowy and hazy days are the core challenging scenarios of this dataset, which also point out the direction for subsequent model improvement.
2.2 The purpose method
In view of the detection pain points of the self-built SIM- AIR dataset, such as the high proportion of ultra-small targets, sparse features, complex background environment and prominent dynamic interference, and combined with the rigid requirements of lightweight model and real-time inference in resource-constrained scenarios such as UAV inspection and embedded equipment, this study proposes an improved YOLO-KMM object detection model, with the core goal of“ accurately adapting the characteristics of the dataset, improving the performance of small object detection, and ensuring the feasibility of deployment”. Figure 4 shows the network architecture of YOLO-KMM in detail, and the co-design of the three modules of“ feature enhancement, efficient detection, and direction awareness” realizes the collaborative optimization of detection performance and computing efficiency. In terms of specific design, the C2KD feature enhancement module integrates high-level semantic features and low-level detail features by constructing a cross-scale feature fusion channel, and introduces a spatial attention mechanism to achieve precise focus on small target areas, effectively strengthening the characterization of weak features and suppressing background noise, and specially adapting to the characteristics of low signalto-noise ratio and weak target-background contrast ratio of datasets. Combined with the dynamic channel number optimization strategy, the model parameters and computational amount are greatly reduced under the premise of retaining the effective feature expression ability of the detection head, so as to meet the deployment requirements of resource-limited scenarios. The Conv-M direction sensing module captures the directional characteristics and radiation distribution of small infrared targets through the efficient combination of asymmetric padding, multi-branch parallel convolution and feature splicing, makes up for the lack of directional features caused by atmospheric scattering, and further improves the positioning accuracy of ultra-small targets. The above three modules are integrated into the original YOLOv11 architecture in the form of embedded replacement, and key parameters are optimized according to the target size distribution and category characteristics of the dataset. This improvement idea of“ targeted module design and native architecture compatibility” not only avoids the compatibility problems caused by large-scale refactoring, but also accurately matches the scenario requirements of air-to-air infrared detection, and finally ensures that the model achieves the synergistic improvement of small target detection accuracy and inference speed without increasing deployment costs, meeting the dual requirements of practical application scenarios.
2.2.1 C2KD
As the core feature extraction component of YOLO11, the C2PSA module adopts a hybrid architecture combining CSPNet and multi-head self-attention, enabling simultaneous capture of local feature correlations and global feature dependencies. However, the standard self-attention in the original module requires calculating the pairwise similarity between all input tokens, leading to a quadratic increase in computational complexity O( n 2) and excessive memory consumption. This issue is particularly prominent when processing high-resolution images or large-scale token sequences, severely limiting the scalability of the model in