跳至主導覽 跳至搜尋 跳過主要內容

CFCM-Gen:Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation

  • Jiahui Shi
  • , Guoheng Huang
  • , Yumian Yu
  • , Xiaochen Yuan
  • , Chi Man Pun
  • , Lianglun Cheng
  • , Yan Li
  • , Qi Yang
  • , Ling Guo
  • , Baiying Lei
  • , Haojiang Li
  • Guangdong University of Technology
  • University of Macau
  • Shenzhen Polytechnic
  • Sun Yat-Sen University Cancer Center
  • Shenzhen University

研究成果: Article同行評審

摘要

3D medical imaging provides rich volumetric information and plays a critical role in disease diagnosis, however, due to the high-dimensional complexity of 3D radiology images, automatically generating radiology reports remains a challenging task. Specifically, this complexity makes cross-modal feature alignment between images and text challenging, weakening their semantic association mapping. To this end, we propose the Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation (CFCM-Gen), which aims to enhance cross-modal feature co-optimization to improve semantic alignment between images and text, thereby strengthening the reporting consistency and cross-modal semantic alignment of automatically generated reports. In addition, we construct the first 3D MRI-text aligned dataset specifically designed for nasopharyngeal carcinoma diagnosis, named NPC-RG, which provides high-quality data support for the 3D radiology report generation task. In the design of our method, firstly, we design Keywords-Prompt for Local-Global Feature Fusion Module (KPLG), which uses medical keywords as hints to guide the local features and global features of 3D radiology images to be fused. This enhances the semantic representation of key regions and provides more precise feature information for cross-modal alignment. Secondly, the Dynamic Bidirectional Modality Interaction Module (DBMI) is proposed to dynamically optimize bidirectional information flow. It refines the correspondence between image regions and textual words, mitigates semantic conflicts, and enhances the alignment accuracy of cross-modal features. Finally, we propose Block-specific Reinforcement Feedback Mechanism (BSRF) to optimize semantic modeling by text block partitioning and reinforcement learning reward signals. Experimental results show that CFCM-Gen performs well in terms of reporting consistency and semantic correspondence.

指紋

深入研究「CFCM-Gen:Cross-modal Feature Co-optimization Mechanism for 3D Radiology Report Generation」主題。共同形成了獨特的指紋。

引用此