TY - JOUR
T1 - CCSFusion
T2 - A Hierarchical Semantic Chain-of-Thought Reasoning Architecture for Infrared-Visible Image Fusion and Captioning
AU - Lin, Miaoshan
AU - Huang, Guoheng
AU - Yang, Jietao
AU - Zheng, Jiehao
AU - Yuan, Xiaochen
AU - Li, Yan
AU - Zhang, Xiaofeng
AU - Tsang, Kim Fung
AU - Pun, Chi Man
N1 - Publisher Copyright:
© 2014 IEEE.
PY - 2026
Y1 - 2026
N2 - Infrared-Visible Image Fusion (IVIF) aims to generate a single, information-rich image for downstream tasks. However, prevailing methods exhibit two key limitations. First, many approaches lack explicit hierarchical semantic decoupling, failing to effectively integrate semantic features across different levels, which restricts their ability to capture complex scene structures. Second, task-driven fusion frameworks typically adopt a cascaded design, with unidirectional supervision provided by geometry-centric downstream tasks like detection. This architecture not only limits mutual reinforcement between the fusion and task networks, but also creates a ”supervision bottleneck” by lacking interaction with the linguistic modality that captures richer scene relationships. To tackle these challenges, we propose CCSFusion, the first framework that leverages Chain-of-Thought captioning as supervision, redirecting IVIF optimization from narrow geometric accuracy to multimodal scene comprehension. It establishes a mutually reinforcing coupling between the fusion network and the captioning task. Specifically, we introduce a Segmentation Mask Calibration Unit (SMCU) to refine coarse semantic priors, providing precise pixel-level guidance. Subsequently, the calibrated features are fed into Chained Semantic Fusion Module (CSFM) which explicitly decomposes the semantic priors into three hierarchical levels, and then feeds them into the Hierarchical Semantic Attention module. Finally, a bidirectional knowledge distillation mechanism transfers the reasoning ability of the teacher network to the student. Experiments show that CCSFusion achieves superior fusion performance and generates more semantically coherent images for high-level cognitive tasks.
AB - Infrared-Visible Image Fusion (IVIF) aims to generate a single, information-rich image for downstream tasks. However, prevailing methods exhibit two key limitations. First, many approaches lack explicit hierarchical semantic decoupling, failing to effectively integrate semantic features across different levels, which restricts their ability to capture complex scene structures. Second, task-driven fusion frameworks typically adopt a cascaded design, with unidirectional supervision provided by geometry-centric downstream tasks like detection. This architecture not only limits mutual reinforcement between the fusion and task networks, but also creates a ”supervision bottleneck” by lacking interaction with the linguistic modality that captures richer scene relationships. To tackle these challenges, we propose CCSFusion, the first framework that leverages Chain-of-Thought captioning as supervision, redirecting IVIF optimization from narrow geometric accuracy to multimodal scene comprehension. It establishes a mutually reinforcing coupling between the fusion network and the captioning task. Specifically, we introduce a Segmentation Mask Calibration Unit (SMCU) to refine coarse semantic priors, providing precise pixel-level guidance. Subsequently, the calibrated features are fed into Chained Semantic Fusion Module (CSFM) which explicitly decomposes the semantic priors into three hierarchical levels, and then feeds them into the Hierarchical Semantic Attention module. Finally, a bidirectional knowledge distillation mechanism transfers the reasoning ability of the teacher network to the student. Experiments show that CCSFusion achieves superior fusion performance and generates more semantically coherent images for high-level cognitive tasks.
KW - Chain-of-Thought
KW - Hierarchical Semantics
KW - Image Captioning
KW - Infrared-Visible Image Fusion
KW - Task-Driven Fusion
UR - https://www.scopus.com/pages/publications/105040176157
U2 - 10.1109/JIOT.2026.3696435
DO - 10.1109/JIOT.2026.3696435
M3 - Article
AN - SCOPUS:105040176157
SN - 2327-4662
JO - IEEE Internet of Things Journal
JF - IEEE Internet of Things Journal
ER -