428
J. Eur. Opt. Society-Rapid Publ. 22, 43( 2026)
5 Jocher G. et al., YOLOv5: A state-of-the-art real-time object detection system, GitHub Repository( 2020). https:// github. com / ultralytics / yolov5. 6 Wang C.-Y., Bochkovskiy A., Liao H.-Y. M., YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for realtime object detectors, Proc. IEEE / CVF Conf. Comput. Vis. Pattern Recognit. 7464 – 7475( 2023). https:// doi. org / 10.1109 / CVPR52733.2023.00721.
7 Li C. et al., YOLOv6: A single-stage object detection framework for industrial applications, arXiv: 2209.02976( 2022). https:// doi. org / 10.48550 / arXiv. 2209.02976.
8 Terven J., Cordova-Esparza D., A comprehensive review of YOLO: From YOLOv1 to YOLOv8 and beyond, arXiv: 2304.00501( 2023). https:// doi. org / 10.48550 / arXiv. 2304.00501.
9 Vaswani A. et al., Attention is all you need, Adv. Neural Inf. Process. Syst. 30, 5998 – 6008( 2017). https:// doi. org / 10.48550 / arXiv. 1706.03762.
10 Woo S., Park J., Lee J.-Y., Kweon I. S., CBAM: Convolutional block attention module, Proc. Eur. Conf. Comput. Vis., 3 – 19( 2018). https:// doi. org / 10.1007 / 978-3-030-01234-2 _ 1.
11 Hu J., Shen L., Sun G., Squeeze-and-excitation networks, Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 7132 – 7141( 2018). https:// doi. org / 10.1109 / CVPR. 2018.00745.
12 Howard A. G. et al., MobileNets: Efficient convolutional neural networks for mobile vision applications, arXiv: 1704.04861( 2017). https:// doi. org / 10.48550 / arXiv. 1704.04861.
13 Zhang X., Zhou X., Lin M., Sun J., ShuffleNet: An extremely efficient convolutional neural network for mobile devices, Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 6848 – 6856( 2018). https:// doi. org / 10.1109 / CVPR. 2018.00716.
14 Sandler M., Howard A., Zhu M., Zhmoginov A., Chen L.-C., MobileNetV2: Inverted residuals and linear bottlenecks, Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 4510 – 4520( 2018). https:// doi. org / 10.1109 / CVPR. 2018.00474.
15 Tan M., Le Q., EfficientNet: Rethinking model scaling for convolutional neural networks, Proc. Int. Conf. Mach. Learn. 97, 6105 – 6114( 2019). https:// doi. org / 10.48550 / arXiv. 1905.11946.
16 Liu S., Qi L., Qin H., Shi J., Jia J., Path aggregation network for instance segmentation, Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 8759 – 8768( 2018). https:// doi. org / 10.1109 / CVPR. 2018.00913.
17 Lin T.-Y., Dollár P., Girshick R., He K., Hariharan B., Belongie S., Feature pyramid networks for object detection, Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2117 – 2125( 2017). https:// doi. org / 10.1109 / CVPR. 2017.106.
18 Du D. et al., VisDrone-DET2019: The vision meets drone object detection in image challenge results, Proc. IEEE / CVF Int. Conf. Comput. Vis. Workshops, 0 – 0( 2019). https:// doi. org / 10.1109 / ICCVW. 2019.00400.
19 Zhu P. et al., Detection and tracking meet drones challenge, IEEE Trans. Pattern Anal. Mach. Intell. 44, 7380 – 7399( 2021). https:// doi. org / 10.1109 / TPAMI. 2021.3119563.
20 Cao Y., Chen S., Zhang Y., Zhang Q., LWIR vs. MWIR vs. SWIR: A comparative study of infrared imaging for UAVbased object detection, Proc. SPIE 11740, 117400K( 2021). https:// doi. org / 10.1117 / 12.2588235.
21 Everingham M., Van Gool L., Williams C. K. I., Winn J., Zisserman A., The Pascal visual object classes( VOC) challenge, Int. J. Comput. Vis. 88, 303 – 338( 2010). https:// doi. org / 10.1007 / s11263-009-0275-4.
22 Lin T.-Y. et al., Microsoft COCO: Common objects in context, Proc. Eur. Conf. Comput. Vis., 740 – 755( 2014). https:// doi. org / 10.1007 / 978-3-319-10602-1 _ 48.
23 Rezatofighi H. et al., Generalized intersection over union: A metric and a loss for bounding box regression, Proc. IEEE / CVF Conf. Comput. Vis. Pattern Recognit., 658 – 666( 2019). https:// doi. org / 10.1109 / CVPR. 2019.00075.
24 Redmon J., Farhadi A., YOLOv3: An incremental improvement arXiv: 1804.02767( 2018). https:// doi. org / 10.48550 / arXiv. 1804.02767.
25 Girshick R., Donahue J., Darrell T., Malik J., Rich feature hierarchies for accurate object detection and semantic segmentation, Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 580 – 587( 2014). https:// doi. org / 10.1109 / CVPR. 2014.81.
26 Ren S., He K., Girshick R., Sun J., Faster R-CNN: Towards real-time object detection with region proposal networks, Adv. Neural Inf. Process. Syst. 28( 2015). https:// doi. org / 10.48550 / arXiv. 1506.01497.
27 Liu W. et al., SSD: Single shot multibox detector, Proc. Eur. Conf. Comput. Vis., 21 – 37( 2016). https:// doi. org / 10.1007 / 978-3-319-46448-0 _ 2.
28 Deng J. et al., ImageNet: A large-scale hierarchical image database, Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 248 – 255( 2009). https:// doi. org / 10.1109 / CVPR. 2009.5206848.
29 Jiang C., Ren H., Ye X., Zhu J., Zeng H., Yang N., Sun M., Ren X., Huo H., Object detection from UAV thermal infrared images and videos using YOLO models, Int. J. Appl. Earth Obs. Geoinf. 112, 102912( 2022). https:// doi. org / 10.1016 / J. JAG. 2022.102912.
30 Andraši P., Radišić T., Muštra M., Ivošević J., Night-time detection of UAVs using thermal infrared camera, Transp. Res. Procedia 28, 183( 2017). https:// doi. org / 10.1016 / j. trpro. 2017.12.184.
31 Mittal P., A comprehensive survey of deep learning-based lightweight object detection models for edge devices, Artif. Intell. Rev. 57, 242( 2024). https:// doi. org / 10.1007 / S10462- 024-10877-1.
32 Fan Q., Li Y., Deveci M., Zhong K., Kadry S., LUD-YOLO: A novel lightweight object detection network for unmanned aerial vehicle, Inf. Sci. 686, 121366( 2025). https:// doi. org / 10.1016 / J. INS. 2024.121366.
33 Han B. G., Lee J. G., Lim K. T., Choi D. H., Design of a scalable and fast YOLO for edge-computing devices, Sensors 20, 6779( 2020). https:// doi. org / 10.3390 / S20236779.
34 Li J., Ye J., Edge-YOLO: Lightweight infrared object detection method deployed on edge devices, Appl. Sci. 13, 4402( 2023). https:// doi. org / 10.3390 / APP13074402.
35 Zhang R., Li H., Duan K., You S., Liu K., Wang F., Hu Y., Automatic detection of earthquake-damaged buildings by integrating UAV oblique photography and infrared thermal imaging, Remote Sens. 12, 2621( 2020). https:// doi. org / 10.3390 / rs12162621.
36 He K., Zhang X., Ren S., Sun J., Deep residual learning for image recognition, Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 770 – 778( 2016). https:// doi. org / 10.1109 / CVPR. 2016.90.
37 Wu D., Cao L., Zhou P., Li N., Li Y., Wang D., Infrared small-target detection based on radiation characteristics with a multimodal feature fusion network, Remote Sens. 14, 3570( 2022). https:// doi. org / 10.3390 / RS14153570.
38 Liu Z., Zou Y., Hu Z., Xue H., Li M., Rao B., Research on multi-modal fusion detection method for low-slow-small uavs based on deep learning, Drones 9, 852( 2025). https:// doi. org / 10.3390 / DRONES9120852.
39 Iandola F. N. et al., SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and < 0.5MB model size, arXiv: 1602.07360( 2016). https:// doi. org / 10.48550 / arXiv. 1602.07360.