Skip to main navigation Skip to search Skip to main content

3D-IRMM-Net: An Interpretable Radiologist-Mimicking 3D Multimodal Network for Lesion Detection in ABUS with LLM-Based Guidance and Uncertainty Quantification

  • Xin Qian
  • , Wei Ke
  • , Luoxi Zhu
  • , Yue Sun
  • , Jiajia Lu
  • , Zhang Xiong
  • , Lingyun Bao
  • , Tao Tan
  • Macao Polytechnic University
  • Westlake University
  • First People’s Hospital of Xiaoshan District

Research output: Contribution to journalArticlepeer-review

Abstract

Automated breast ultrasound (ABUS) has emerged as a promising tool for breast lesion detection, but most existing deep learning models for ABUS lack transparency and fail to align with clinical reasoning processes. This limits their interpretability and hinders clinical adoption. Therefore, we propose 3D-IRMM-Net, a radiologist-mimicking 3D multimodal network. It integrates visual, textual, and semantic cues to emulate hierarchical diagnostic reasoning. At its core, the network integrates two novel components: the 3D Scale-Expert Convolution (3D-SEC) block, which enables parameter-efficient, topology-aware feature routing via a shared graph convolutional layer; and the Adaptive Spatial-Gaussian (ASG) block, which enhances lesion localization in ABUS by modeling spatial dependencies through multi-scale Gaussian attention. In addition, we propose LLM-Vision activator, a natural-language-driven mechanism that uses language to guide 3D visual attention activation and enables 3D spatial-perceptual feature learning via a pretrained 2D VLM, boosting training efficiency. By aligning AI inference with clinical workflows, 3D-IRMM-Net enhances both detection rates and interpretability. Extensive experiments on ABUS2025, TDSC-ABUS2023, and FPHXD-ABUS datasets confirm its state-of-the-art performance across multiple metrics. At an FPPI of 0.5, our method achieved a patient-level detection rates of 89.7±1.1% (internal) and 85.1±2.0% (external); at an FPPI of 4, they rose to 96.1±1.6% and 91.5 ±1.9%, respectively. Furthermore, the model provides principled uncertainty quantification, improving trustworthiness in real-world deployment. These results confirm the effectiveness of 3D-IRMM-Net in bridging the gap between AI-driven detection and clinical practice.

Original languageEnglish
JournalIEEE Transactions on Circuits and Systems for Video Technology
DOIs
Publication statusAccepted/In press - 2026

Keywords

  • 3D lesion detection
  • ABUS
  • interpretability
  • large language model
  • multimodal

Fingerprint

Dive into the research topics of '3D-IRMM-Net: An Interpretable Radiologist-Mimicking 3D Multimodal Network for Lesion Detection in ABUS with LLM-Based Guidance and Uncertainty Quantification'. Together they form a unique fingerprint.

Cite this