1. 安徽理工大学 安全科学与工程学院,安徽,淮南,232001
2. 安徽理工大学 计算机科学与工程学院,安徽,淮南,232001
3. 安徽理工大学 公共安全与应急管理学院,安徽,合肥,231131
收稿:2025-06-10,
网络首发:2026-08-17,
移动端阅览
廖欢,顾成杰,朱东郡,等. 面向多光谱无人机遥感图像的双流融合跨模态目标检测方法[J/OL]. 兵工学报, 2026(2026-08-17). https://doi.org/10.12382/bgxb.2025.0480.
LIAO H, GU C, ZHU D J, et al. An object detection method for multispectral uav remote sensing images[J/OL]. Acta Armamentarii, 2026(2026-08-17). https://doi.org/10.12382/bgxb.2025.0480. (in Chinese)
廖欢,顾成杰,朱东郡,等. 面向多光谱无人机遥感图像的双流融合跨模态目标检测方法[J/OL]. 兵工学报, 2026(2026-08-17). https://doi.org/10.12382/bgxb.2025.0480. DOI:
LIAO H, GU C, ZHU D J, et al. An object detection method for multispectral uav remote sensing images[J/OL]. Acta Armamentarii, 2026(2026-08-17). https://doi.org/10.12382/bgxb.2025.0480. (in Chinese) DOI:
针对当前无人机多光谱遥感目标检测中存在的模态内特征弱化与跨模态交互低效问题,提出一种融合双流单模态增强与跨模态特征交互的目标检测方法。设计一种基于深度可分离卷积的小波变换模块,通过提取并增强特征图的高频细节信息,强化可见光分支对纹理细节的提取能力。提出一种跨尺度全局信息融合模块,利用多尺度特征整合和全局信息提取,增强红外分支对热物体的检测能力。构建一种多模态混合自注意力融合模块,通过自适应引导可见光与红外模态的特征交互,充分挖掘光谱间互补性。实验结果表明,该方法在可见光-红外多光谱数据集DVTOD和DroneVehicle上的mAP@0.50分别达到86.7%和86.0%,较基线模型分别提升1.3%和1.9%,有效提升了无人机多光谱目标检测性能。
To address the issues ofweakenedintra-modal featuresand inefficient cross-modal interaction in current UAV multispectral remote sensing object detection
an object detectionmethodthatintegratesdual-stream single-modal enhancement and cross-modal feature interaction is proposed. A wavelet transform module based on depthwise separable convolution is designed to extract and enhancethehigh-frequency detail information in feature maps
thusstrengthening the texture detail extraction capability of the visible light branch. A cross-scale global information fusion module is proposed
which utilizes multi-scale feature integration and global information extraction to enhance the detection capability of the infrared branch for thermal objects. A multimodal hybrid self-attention fusion module is constructed to adaptively guidethefeature interaction between visible light and infrared modalities
fully exploiting spectral complementarity. Experimental results show that the proposed method achieves mAP@0.5 of 86.7% and 86.0% on the visible-infrared multispectral datasets DVTOD and DroneVehicle
respectively
andimproves themby 1.3% and 1.9%
respectively
compared to the baseline model
effectively enhancingthe performance of UAV multispectral object detection.
苏向阳,汪洋,谭森起,等.PCA-YOLO11:全域复杂环境的轻量目标检测[J].兵工学报,2025,46(增刊2):250584.
SU X Y, WANG Y, TAN S Q, et al. PCA-YOLO11: lightweight object detection in multi-altitude complex environments[J]. Acta Armamentarii, 2025,46(S2): 250584. (in Chinese)
赵紫臣,萧浩东,凌焕章.基于YOLOv10n的轻量化红外小目标检测模型[J].兵工学报,2025,46(11):250146.
ZHAO Z C, XIAO H D, LING H Z. Lightweight infrared small target detection model based on YOLOv10n[J]. Acta Armamentarii, 2025,46(11):250146. (in Chinese)
王泉,叶广飞,陈祺东.YOLO-SWR:无人机视角下轻量级交通车辆检测方法[J].计算机工程与应用, 2025,61(14):
-122.
WANG Q, YE G F, CHEN Q D. Lightweight traffic vehicle detection algorithm from UAV perspective[J]. Computer Engineering and Applications, 2025, 61(14): 112-122. (in Chinese)
徐相华,周德靖,余超然,等.基于DCA-YOLO的受病虫害侵染树木农业无人机低空检测模型[J].农业机械学报,2025,56(10):479-491.
XU X H, ZHOU D J, YU C R, et al. Improved model for low altitude detection of trees infected by pests and doseases using agricultural dronesbased on DCA-YOLO[J]. Transactions of the Chinese Society for Agricultural Machinery, 2025, 56(10): 479-491. (in Chinese)
丁子天,喜文飞,钱堂慧,等.结合多特征的无人机雾天影像识别[J].红外技术,2025,47(7):833-841.
DING Z T, XI W F, QIAN T H, et al.Multiple feature fusion for unmanned aerial vehicle image recognition in foggy weather[J]. Infrared Technology, 2025, 47(7):833-841. (in Chinese)
孙备, 党昭洋,吴鹏,等. 多尺度互交叉注意力改进的单无人机对地伪装目标检测定位方法[J].仪器仪表学报, 2023, 44(6): 54-65.
SUN B, DANG Z Y, WU P, et al. Multi scale cross attention improved method of single unmanned aerial vehicle for ground camouflage target detection and localization[J].Chinese Journal of Scientific Instrument, 2023,44(6):54-65.(in Chinese)
付琨, 王佩瑾, 冯瑛超, 等.遥感跨模态智能解译:模型、数据与应用[J].中国科学:信息科学, 2023,53(8): 1529-1559.
FU K, WANG P J, FENG Y C, et al. Cross-modal remote sensing intelligent interpretation:method, data and application[J].Scientia Sinica(Informationis ),2023,53(8):1529-1559. (in Chinese)
XIE Z X, SHAO F, CHEN G, et al. Cross-modality double bidirectional interaction and fusion network for RGB-T Salient object detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2023, 33(8): 4149-4163.
WANG K P, TU Z Z, LI C L, et al. Learning adaptive fusion bank for multi-modal salient object detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(8): 7344-7358.
LIU X Y, QI J H, CHEN C, et al. Relation-aware weight sharing in decoupling feature learning network for UAV RGB-infrared vehicle re-identification[J]. IEEE Transactions on Multimedia, 2024, 26: 9839-9853.
ZHANG J Q, LEI J, XIE W Y, et al.SuperYOLO:super-resolution assisted object detection in multimodal remote sensing imagery:arXiv:2209.13351[R]. Ithaca,NY,US:Cornell University, 2022:2209.13351.
CHEN Y S, WANG B R, GUO X Y, et al. Deyolo: dual-feature-enhancement yolo for cross-modality object detection[J].Lecture Notes in Computer Science, 2024,15317: 236-252.
FANG Q Y, HAN D P, WANG Z K. Cross-modality fusion transformer for multispectral object detection:ArXiv:2111.00273[R].Ithaca,NY,US:Cornell University, 2021: 2111.00273.
SHEN J, CHEN Y, LIU Y, et al. ICAFusion: iterative cross-attention guided feature fusion for multispectral object detection[J]. Pattern Recognition, 2024, 145:109913.
代小波,任思羽,何小海,等.基于特征增强和对齐融合的可见光–红外图像融合目标检测[J].工程科学学报,2025,47(12):2554-2565.
DAI X B, REN S Y, HE X H, et al.Visible–infrared fusion object detection based on feature enhancement and alignment fusion [J]. Chinese Journal of Engineering, 2025, 47(12):2554-2565. (in Chinese)
XIAO Y M, MENG F M, WU Q Bet al.GM-DETR:generalized multispectral detection transformer with efficient fusion encoder for visible-infrared detection[C]//2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024:17157-17166.
SONG K C, XUE X T, WEN H W, et al. Misaligned visible-thermal object detection: a drone-based benchmark and baseline[J] IEEE Transactions on Intelligent Vehicles, 2024,9(11): 7449-7460.
FINDER S E, AMOYAL R, TREISTER E, et al. Wavelet convolutions for large receptive fields[C]//European Conference on Computer Vision. 2024: 363-380.
CAO Y, XU J R, LIN S, et al. GCNet: non-local networks meet squeeze-excitation networks and beyond[C]//2019 IEEE/CVF International Conference on Computer Vision Workshop. 2019:1971-1980.
SUN Y M, CAO B, ZHU P F, et al. Drone-based RGB-infrared cross-modality vehicle detection via uncertainty-aware learning[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(10): 6700-6713.
0
浏览量
0
下载量
0
CNKI被引量
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024360号