TY - JOUR
T1 - Leveraging modality-guided pre-training for dual-prompt-driven multi-cancer PET-CT segmentation
AU - Liang, Xinglong
AU - Huang, Jiaju
AU - Zhang, Tianyu
AU - Han, Luyi
AU - Wang, Xin
AU - Gao, Yuan
AU - Lu, Chunyao
AU - Sun, Yue
AU - Teuwen, Jonas
AU - Tan, Tao
AU - Mann, Ritse
N1 - Publisher Copyright:
Copyright © 2026. Published by Elsevier B.V.
PY - 2026/9
Y1 - 2026/9
N2 - PET-CT lesion segmentation remains challenging due to heterogeneous lesion appearance, small and dispersed lesions, physiological FDG uptake, and limited annotations. Existing self-supervised methods are mostly designed for unimodal imaging and therefore fail to fully exploit the complementary anatomical and metabolic information in PET-CT. Meanwhile, conventional multi-cancer segmentation strategies often treat different cancer types as a unified task, which weakens cancer-specific features, and existing prompt-based methods still have limited task adaptation and sensitivity to small lesions. To address these limitations, a unified two-stage framework for multi-cancer PET-CT segmentation is presented. First, a modality-guided probabilistic masked autoencoder is introduced to enhance cross-modal PET-CT representation learning through modality-specific masking. Second, a dual-prompt downstream segmentation network is designed to model both cancer-specific characteristics and cross-cancer shared knowledge, with prompt-aware heads further improving task adaptation and small-lesion delineation. Experiments on a multi-cancer PET-CT dataset show consistent improvements over the best-performing non-prompt and prompt-based baselines, with average Dice gains of 2.51% and 2.18%, respectively. The framework is further applied to an unannotated breast cancer cohort for survival analysis, demonstrating promising generalizability and improved risk stratification. The code is available at: https://github.com/XinglongLiang08/DpDNet .
AB - PET-CT lesion segmentation remains challenging due to heterogeneous lesion appearance, small and dispersed lesions, physiological FDG uptake, and limited annotations. Existing self-supervised methods are mostly designed for unimodal imaging and therefore fail to fully exploit the complementary anatomical and metabolic information in PET-CT. Meanwhile, conventional multi-cancer segmentation strategies often treat different cancer types as a unified task, which weakens cancer-specific features, and existing prompt-based methods still have limited task adaptation and sensitivity to small lesions. To address these limitations, a unified two-stage framework for multi-cancer PET-CT segmentation is presented. First, a modality-guided probabilistic masked autoencoder is introduced to enhance cross-modal PET-CT representation learning through modality-specific masking. Second, a dual-prompt downstream segmentation network is designed to model both cancer-specific characteristics and cross-cancer shared knowledge, with prompt-aware heads further improving task adaptation and small-lesion delineation. Experiments on a multi-cancer PET-CT dataset show consistent improvements over the best-performing non-prompt and prompt-based baselines, with average Dice gains of 2.51% and 2.18%, respectively. The framework is further applied to an unannotated breast cancer cohort for survival analysis, demonstrating promising generalizability and improved risk stratification. The code is available at: https://github.com/XinglongLiang08/DpDNet .
KW - Modality-guided probabilistic masking
KW - PET-CT segmentation
KW - Prompt-based segmentation
KW - Survival analysis
UR - https://www.scopus.com/pages/publications/105042683360
U2 - 10.1016/j.media.2026.104182
DO - 10.1016/j.media.2026.104182
M3 - Article
AN - SCOPUS:105042683360
SN - 1361-8415
VL - 113
JO - Medical Image Analysis
JF - Medical Image Analysis
M1 - 104182
ER -