电子学报2026,Vol.54Issue(1):86-101,16.DOI:10.12263/DZXB.20250906
基于文本语义引导的红外与可见光图像融合方法
Textual Semantic Guidance for Infrared and Visible Image Fusion
摘要
Abstract
Infrared and visible image fusion(IVF)aims to integrate the complementary information contained in both image modalities by effectively combining the salient targets in infrared images with the rich texture details present in visi-ble images.Through this integration,IVF produces more informative and comprehensive fused images that surpass single-modality inputs.Existing research has demonstrated that deep learning-based fusion methods have achieved remarkable progress in improving fused image quality.However,most of these approaches focus mainly on low-level visual features,and the deep semantic associations between high-level semantic information and visual features have not yet been sufficient-ly explored.In recent years,with the rapid development of large vision-language models(VLMs),text-guided image fusion methods have exhibited great potential due to their flexibility and versatility.However,the effective integration and utiliza-tion of textual semantic information in the image fusion process remain insufficiently studied.To tackle these challenges,this paper proposes a textual semantic guidance method for infrared and visible image fusion,termed textual semantic guid-anc(TeSG),which guides the image synthesis process in a way that is optimized for downstream tasks such as object detec-tion and semantic segmentation.By explicitly introducing high-level semantic information generated by VLMs into the fu-sion pipeline,TeSG achieves precise regulation of the fusion process and enhances the semantic consistency of the fused re-sults.TeSG introduces textual semantics at two levels:the mask semantic level and the text semantic level.First,automati-cally generated textual descriptions from VLMs are employed as global text-level semantic guidance,providing high-level semantic constraints for the fusion process.Second,based on these textual descriptions,mask semantics corresponding to key target regions are constructed,enabling accurate localization and differentiated modeling of foreground and background regions.Building on this,three core modules are designed to implement the proposed framework.The semantic information generator(SIG)module generates both mask semantics and text semantics from automatically produced textual descrip-tions.The mask-guided cross-attention(MGCA)module performs preliminary attention-based fusion of visual features from both infrared and visible images under the guidance of mask semantics,thereby realizing mask-level cross-modal fea-ture interaction.Finally,the text-driven attentional fusion(TDAF)module achieves text-level fusion and dynamic weighting through text-guided attention and a gating mechanism,allowing semantic cues to modulate the contribution of different mo-dalities in an adaptive manner.Experimental results demonstrate that the proposed TeSG method,through its dual-level tex-tual semantic guidance paradigm,performs favorably against existing state of the art(SOTA)methods in preserving multi-modal texture information and enhancing contrast in the fused images.In addition,TeSG yields superior performance in downstream tasks such as object detection and semantic segmentation,highlighting its task-oriented fusion capability.Com-pared with current SOTA image fusion approaches,the proposed TeSG achieves an average improvement of 1.4%on down-stream tasks,validating its competitiveness and effectiveness while also exhibiting strong generalization ability across dif-ferent datasets and scene conditions.The proposed method effectively addresses the insufficient exploration of deep correla-tions between textual and visual features in existing image fusion algorithms,achieving simultaneous improvements in fu-sion quality and downstream task performance.关键词
图像融合/红外与可见光图像/文本语义引导/深度学习/视觉-语言模型/注意力Key words
image fusion/infrared and visible images/textual semantic guidance/deep learning/vision-language models/attention分类
信息技术与安全科学引用本文复制引用
朱明瑞,陈希茹,卫鑫,王楠楠,高新波..基于文本语义引导的红外与可见光图像融合方法[J].电子学报,2026,54(1):86-101,16.基金项目
国家自然科学基金(No.62576261,No.U22A2096) National Natural Science Foundation of China(No.62576261,No.U22A2096) (No.62576261,No.U22A2096)