422
J. Eur. Opt. Society-Rapid Publ. 22, 43( 2026)
accuracy and computational efficiency, and exhibits outstanding performance especially in scenarios with small targets and complex backgrounds. This improved C3K2 module achieves outstanding performance especially in scenarios with small targets and complex backgrounds.
Figure 3. Target and background average gray level comparison.
this feature map is sent to two parallel 3 3convolution branches to generate two complementary feature maps( F1 andF2) focusing on different feature dimensions; then the MU operation is performed on these two feature maps, capturing the nonlinear correlations between corresponding feature elements via element-wise multiplication; finally, the fused feature map recovers the channel dimension through a 1 1 convolutional layer to form the final output feature map( Ffinal). The feature map generation process of the MU module can be expressed as the following equation( 4):
F final ¼ Conv restore ðF 1 F 2 þ Conv reduce ðXÞÞ ð4Þ
Here, X denotes the input feature map of the MU bottleneck block; Conv – reduce( X) represents the 1 1convolution operation applied to X for channel dimensionality reduction; Conv – branch1 and Conv – branch2 denote two parallel 3 3 convolution operations used for feature transformation of the dimensionality-reduced feature map; the symbol represents element-wise multiplication between the feature maps F 1 and F 2 of the two branches; Conv – restore denotes the 1 1 convolution operation for restoring the channel dimension of the fused feature map; the introduction of residual connection( summing Conv – reduce( X) and the fused feature map) can alleviate the gradient vanishing problem and retain the original feature information [ 18 ].
To further optimize the adaptability of feature extraction, this paper retains the residual connection mechanism of the original C3K module and introduces the ReLU6 activation function after the dual-branch convolution of the MU bottleneck block. ReLU6 constrains the output values within the interval [ 0, 6 ], which enhances the model’ s robustness to numerical instability and improves the inference efficiency on edge devices simultaneously [ 19 ]. This combination enables the MU bottleneck block to adaptively capture complex feature correlations according to input content, which significantly enhances the model’ s flexibility and representational capability compared with the traditional linear summation method. The improved C3K2 module achieves a better balance between detection
2.2.3
Conv-M
The Conv-M module generates enhanced feature maps through an efficient process combining asymmetric padding, multi-branch parallel convolution and feature concatenation [ 20 ]. The input feature map X is first processed by four parallel asymmetric padding convolution operations, where each convolution branch is designed with a dedicated convolution kernel and padding mode for different spatial directions: the horizontal direction adopts a 1 3 convolution kernel with left-right asymmetric padding, and the vertical direction adopts a 3 1convolution kernel with top-bottom asymmetric padding, yielding four complementary feature branches( X1, X2, X3, X4). Then, channel concatenation is performed on the feature maps of the four branches to integrate multi-directional feature information. Finally, a 2 2 convolutional layer is used for feature fusion and dimension adjustment to generate the final output feature map, the structural process is shown in Figure 7. Batch Normalization and SiLU activation function are appended after each convolution layer throughout the process [ 21 ], which ensures training stability and the nonlinear expression capability of features. The feature map generation process of the Conv-M module can be expressed as the following equation( 5):
F f ¼ SiLUðBNðCatðX 1; X 2; X 3; X 4 Þ W 22 ÞÞ ð5Þ
Here, X 1 X 4 denote the output feature maps of the four |
parallel convolution branches, and their generation methods |
are as follows equation( 6): |
|
�
�
X 1 ¼ SiLU BN X
Pð1; 0; 0; 3Þ
|
W 13
|
|
|
�
�
X 2 ¼ SiLU BN X
Pð0; 3; 0; 1Þ
�
�
X 3 ¼ SiLU BN X
Pð0; 1; 3; 0Þ
|
W 31
W 13
|
ð6Þ |
|
�
�
X 4 ¼ SiLU BN X
Pð3; 0; 1; 0Þ
|
W 31
|
|
Here, X P( left, right, top, bottom) denotes the asymmetric padding of the input feature map X with specified pixels in each direction( the numbers in parentheses are the padding pixel counts for left, right, top and bottom directions respectively); W 13 and W 31 represent the directional convolution kernels of 1 3and3 1 respectively; W 22 denotes the 2 2 convolution kernel for final feature fusion; is the convolution operator; Cat() denotes the channel concatenation operation of feature maps.
Therefore, replacing the standard convolution in the C3K2 module with Conv�M can match the Gaussian distribution characteristics of infrared small targets, strengthen the weak target feature extraction capability through multi-directional feature capture and efficient receptive field expansion, while avoiding redundant