跳至主導覽 跳至搜尋 跳過主要內容

CCSFusion: A Hierarchical Semantic Chain-of-Thought Reasoning Architecture for Infrared-Visible Image Fusion and Captioning

  • Miaoshan Lin
  • , Guoheng Huang
  • , Jietao Yang
  • , Jiehao Zheng
  • , Xiaochen Yuan
  • , Yan Li
  • , Xiaofeng Zhang
  • , Kim Fung Tsang
  • , Chi Man Pun
  • Guangdong University of Technology
  • Shenzhen Polytechnic
  • Shanghai Jiao Tong University
  • Shenzhen Institute of Advanced Technology
  • University of Macau

研究成果: Article同行評審

摘要

Infrared-Visible Image Fusion (IVIF) aims to generate a single, information-rich image for downstream tasks. However, prevailing methods exhibit two key limitations. First, many approaches lack explicit hierarchical semantic decoupling, failing to effectively integrate semantic features across different levels, which restricts their ability to capture complex scene structures. Second, task-driven fusion frameworks typically adopt a cascaded design, with unidirectional supervision provided by geometry-centric downstream tasks like detection. This architecture not only limits mutual reinforcement between the fusion and task networks, but also creates a ”supervision bottleneck” by lacking interaction with the linguistic modality that captures richer scene relationships. To tackle these challenges, we propose CCSFusion, the first framework that leverages Chain-of-Thought captioning as supervision, redirecting IVIF optimization from narrow geometric accuracy to multimodal scene comprehension. It establishes a mutually reinforcing coupling between the fusion network and the captioning task. Specifically, we introduce a Segmentation Mask Calibration Unit (SMCU) to refine coarse semantic priors, providing precise pixel-level guidance. Subsequently, the calibrated features are fed into Chained Semantic Fusion Module (CSFM) which explicitly decomposes the semantic priors into three hierarchical levels, and then feeds them into the Hierarchical Semantic Attention module. Finally, a bidirectional knowledge distillation mechanism transfers the reasoning ability of the teacher network to the student. Experiments show that CCSFusion achieves superior fusion performance and generates more semantically coherent images for high-level cognitive tasks.

原文English
期刊IEEE Internet of Things Journal
DOIs
出版狀態Accepted/In press - 2026

指紋

深入研究「CCSFusion: A Hierarchical Semantic Chain-of-Thought Reasoning Architecture for Infrared-Visible Image Fusion and Captioning」主題。共同形成了獨特的指紋。

引用此