TY - JOUR
T1 - CFCM-Gen:Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation
AU - Shi, Jiahui
AU - Huang, Guoheng
AU - Yu, Yumian
AU - Yuan, Xiaochen
AU - Pun, Chi Man
AU - Cheng, Lianglun
AU - Li, Yan
AU - Yang, Qi
AU - Guo, Ling
AU - Lei, Baiying
AU - Li, Haojiang
N1 - Publisher Copyright:
© 2017 IEEE.
PY - 2026
Y1 - 2026
N2 - 3D medical imaging provides rich volumetric information and plays a critical role in disease diagnosis, however, due to the high-dimensional complexity of 3D radiology images, automatically generating radiology reports remains a challenging task. Specifically, this complexity makes cross-modal feature alignment between images and text challenging, weakening their semantic association mapping. To this end, we propose the Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation (CFCM-Gen), which aims to enhance cross-modal feature co-optimization to improve semantic alignment between images and text, thereby strengthening the reporting consistency and cross-modal semantic alignment of automatically generated reports. In addition, we construct the first 3D MRI-text aligned dataset specifically designed for nasopharyngeal carcinoma diagnosis, named NPC-RG, which provides high-quality data support for the 3D radiology report generation task. In the design of our method, firstly, we design Keywords-Prompt for Local-Global Feature Fusion Module (KPLG), which uses medical keywords as hints to guide the local features and global features of 3D radiology images to be fused. This enhances the semantic representation of key regions and provides more precise feature information for cross-modal alignment. Secondly, the Dynamic Bidirectional Modality Interaction Module (DBMI) is proposed to dynamically optimize bidirectional information flow. It refines the correspondence between image regions and textual words, mitigates semantic conflicts, and enhances the alignment accuracy of cross-modal features. Finally, we propose Block-specific Reinforcement Feedback Mechanism (BSRF) to optimize semantic modeling by text block partitioning and reinforcement learning reward signals. Experimental results show that CFCM-Gen performs well in terms of reporting consistency and semantic correspondence.
AB - 3D medical imaging provides rich volumetric information and plays a critical role in disease diagnosis, however, due to the high-dimensional complexity of 3D radiology images, automatically generating radiology reports remains a challenging task. Specifically, this complexity makes cross-modal feature alignment between images and text challenging, weakening their semantic association mapping. To this end, we propose the Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation (CFCM-Gen), which aims to enhance cross-modal feature co-optimization to improve semantic alignment between images and text, thereby strengthening the reporting consistency and cross-modal semantic alignment of automatically generated reports. In addition, we construct the first 3D MRI-text aligned dataset specifically designed for nasopharyngeal carcinoma diagnosis, named NPC-RG, which provides high-quality data support for the 3D radiology report generation task. In the design of our method, firstly, we design Keywords-Prompt for Local-Global Feature Fusion Module (KPLG), which uses medical keywords as hints to guide the local features and global features of 3D radiology images to be fused. This enhances the semantic representation of key regions and provides more precise feature information for cross-modal alignment. Secondly, the Dynamic Bidirectional Modality Interaction Module (DBMI) is proposed to dynamically optimize bidirectional information flow. It refines the correspondence between image regions and textual words, mitigates semantic conflicts, and enhances the alignment accuracy of cross-modal features. Finally, we propose Block-specific Reinforcement Feedback Mechanism (BSRF) to optimize semantic modeling by text block partitioning and reinforcement learning reward signals. Experimental results show that CFCM-Gen performs well in terms of reporting consistency and semantic correspondence.
KW - Radiology report generation
KW - modality alignment
KW - prompt learning
KW - reinforcement learning
UR - https://www.scopus.com/pages/publications/105038611287
U2 - 10.1109/TRPMS.2026.3688246
DO - 10.1109/TRPMS.2026.3688246
M3 - Article
AN - SCOPUS:105038611287
SN - 2469-7311
JO - IEEE Transactions on Radiation and Plasma Medical Sciences
JF - IEEE Transactions on Radiation and Plasma Medical Sciences
ER -