Skip to main navigation Skip to search Skip to main content

CFCM-Gen:Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation

  • Jiahui Shi
  • , Guoheng Huang
  • , Yumian Yu
  • , Xiaochen Yuan
  • , Chi Man Pun
  • , Lianglun Cheng
  • , Yan Li
  • , Qi Yang
  • , Ling Guo
  • , Baiying Lei
  • , Haojiang Li
  • Guangdong University of Technology
  • University of Macau
  • Shenzhen Polytechnic
  • Sun Yat-Sen University Cancer Center
  • Shenzhen University

Research output: Contribution to journalArticlepeer-review

Abstract

3D medical imaging provides rich volumetric information and plays a critical role in disease diagnosis, however, due to the high-dimensional complexity of 3D radiology images, automatically generating radiology reports remains a challenging task. Specifically, this complexity makes cross-modal feature alignment between images and text challenging, weakening their semantic association mapping. To this end, we propose the Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation (CFCM-Gen), which aims to enhance cross-modal feature co-optimization to improve semantic alignment between images and text, thereby strengthening the reporting consistency and cross-modal semantic alignment of automatically generated reports. In addition, we construct the first 3D MRI-text aligned dataset specifically designed for nasopharyngeal carcinoma diagnosis, named NPC-RG, which provides high-quality data support for the 3D radiology report generation task. In the design of our method, firstly, we design Keywords-Prompt for Local-Global Feature Fusion Module (KPLG), which uses medical keywords as hints to guide the local features and global features of 3D radiology images to be fused. This enhances the semantic representation of key regions and provides more precise feature information for cross-modal alignment. Secondly, the Dynamic Bidirectional Modality Interaction Module (DBMI) is proposed to dynamically optimize bidirectional information flow. It refines the correspondence between image regions and textual words, mitigates semantic conflicts, and enhances the alignment accuracy of cross-modal features. Finally, we propose Block-specific Reinforcement Feedback Mechanism (BSRF) to optimize semantic modeling by text block partitioning and reinforcement learning reward signals. Experimental results show that CFCM-Gen performs well in terms of reporting consistency and semantic correspondence.

Original languageEnglish
JournalIEEE Transactions on Radiation and Plasma Medical Sciences
DOIs
Publication statusAccepted/In press - 2026

Keywords

  • Radiology report generation
  • modality alignment
  • prompt learning
  • reinforcement learning

Fingerprint

Dive into the research topics of 'CFCM-Gen:Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation'. Together they form a unique fingerprint.

Cite this