J. Eur. Opt. Society-Rapid Publ. 22, 43( 2026) 421
k-th attention head( p is the head dimension), and k is the k-th column of the group assignment matrix P, which records the membership probability of each token belonging to the k-th group.
The theoretical complexity ratio between KD and standard self-attention is shown in the equation( 2):
Ra ¼
FLOPs KD
¼ K N C p ¼ C FLOPs Self�Attention K N 2 p N ð2Þ
Figure 2. Proportion of the dataset under different weather conditions.
real-time detection scenarios. To address this limitation, we replace the standard self-attention in C2PSA with our proposed KD mechanism( based on Token Statistical Self-Attention, TSSA) to form the C2KD module. This integration combines the KD mechanism’ scomputational efficiency with CSPNet’ s local feature extraction capability, ultimately achieving a balance between computational efficiency and representational capability. Different from the traditional self-attention based on pairwise similarity, the proposed method adopts an efficient attention computation paradigm based on token statistical features. As shown in the feature map generation process detailed in Figure 5.
The KD module generates discriminative feature maps via a data-driven statistical learning process, which consists of four core steps: token projection, group membership probability estimation, second-order moment calculation, and weighted feature update; the tokenized input feature map Z is first projected into K low-dimensional subspaces through learnable projection matrices { U k }( k = 1,2,...,
K), yielding the projected token features( U T k Z), then softmax-based group membership probability estimation assigns to each token the probability of belonging to each subspace, forming the group assignment matrix P, next the empirical second-order moment statistics of token features within each subspace are calculated to measure the“ feature strength” inside the group, and finally adaptive weighting coefficients are generated based on these statistics to update the original token features, suppressing irrelevant feature directions and enhancing discriminative feature representation, with the feature optimization process of the KD module formally definable by equation( 1):
Z ¼ Z � s n
X K
k¼1
U k D k U k
ZDiagðp k Þ ð1Þ
Here, Z denotes the input token sequence with the shape B N C, where B is the batch size, N is the number of tokens and C is the feature dimension; s is the gradient step size parameter, and n is the total number of tokens. U k 2 R ^( C p) represents the projection matrix of the
Since the number of tokens N( e. g., N = 4096 for a 640 640 image divided by 16 16 patches) is much larger than the feature dimension C( e. g., C = 256 or 512 in YOLO11), the computational cost of TSSA is significantly lower than that of standard self-attention. To further enhance the feature extraction flexibility of the C2PSA module, the KD mechanism integrates an adaptive group assignment strategy based on token features. Unlike fixed partitioning strategies( e. g., sliding windows or block partitioning), KD dynamically estimates the group assignment matrix P via equation( 3):
P j; k ¼ softmax 1
2g kU k z j y j k k2 2 þ bk j ð3Þ
Here z j, denotes the j – th token in Z, g is a learnable temperature parameter, y k j, isthe‘ 2 normalization vector of the projected token( ensuring feature scale consistency), and b kj is a learnable additional bias( used to compensate for cumulative calculation errors in causal scenarios). This adaptive assignment method can allocate tokens with similar semantic features to the same subspace, enhancing the model’ s ability to capture semantic-level feature dependencies.
2.2.2 C3K2-MU
The C3K2 module serves as the core feature extraction component in the latest YOLO11 model, which leverages the CSPNet structure to split, process, and fuse input feature maps [ 15 ]. However, the traditional bottleneck blocks inside the C3K module adopt linear summation for feature fusion, which struggle to capture complex nonlinear feature correlations [ 16 ]; meanwhile, stacked convolutional layers introduce redundant computational overhead. To address these issues, this paper proposes the MU operation, which replaces the traditional bottleneck blocks in the C3K module with MU bottleneck blocks, achieving improved feature representation capability while reducing computational complexity. Unlike conventional bottleneck blocks relying on linear summation, the MU bottleneck block adopts a more efficient and powerful feature fusion strategy centered on the MU operation, whose structure is illustrated in Figure 6.
This operation enables nonlinear interaction between features without explicitly increasing the network width [ 17 ], thus enhancing the model’ s capability to capture fine-grainedfeaturepatterns. TheMUmodulegenerates enhanced feature maps through an efficient process combining dual-branch parallel convolution and element-wise multiplication. The input feature map X is first fed into a 1 1 convolutional layer for channel dimensionality reduction to obtain a dimensionality-reduced feature map; subsequently,