TY - JOUR
T1 - S2DENet
T2 - Shallow suppression and deep enhancement network for general ultrasound image segmentation
AU - Pang, Xintao
AU - Yang, Jinlin
AU - Gao, Zhifan
AU - Lin, Chuan
AU - Sun, Yue
AU - Li, Shuo
AU - de With, Peter H.N.
AU - Tan, Tao
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/9
Y1 - 2026/9
N2 - Ultrasound image segmentation serves as a cornerstone of clinical diagnosis, yet remains a formidable challenge due to inherent image artifacts such as speckle noise and ambiguous boundaries. Existing approaches typically employ uniform feature extraction strategies across all network layers, disregarding the fundamental disparities between noise-dominated shallow stages and semantically-rich deep stages. This monolithic strategy compels networks to concurrently learn noise suppression and feature enhancement, often resulting in over-parameterization and critical compromises in computational efficiency, particularly within resource-constrained environments. To address these limitations, we propose S2DENet, an efficient network architecture that employs noise suppression in shallow layers and semantic feature enhancement in deep layers. Specifically, we design Multi-order Differential Convolution (MDiffConv) to enhance high-frequency feature capture, implementing suppression strategies in shallow layers and enhancement strategies in deep layers. Simultaneously, we introduce a Differential Self-Attention (DiffSA) mechanism to mitigate noise artifacts while preserving structural integrity, with deep layers transitioning to standard self-attention for semantic feature amplification. This depth-differentiated design, combined with differential mechanisms, enables the model to focus on specific tasks at different network stages, thereby reducing learning complexity and improving feature representation efficiency. Experimental results on ten public ultrasound datasets demonstrate that S2DENet achieves an exceptional balance between efficiency and accuracy, attaining state-of-the-art (SOTA) performance on nine public datasets. With merely 0.05M/0.15M parameters (a reduction exceeding 99% compared to SOTA methods), S2DENet maintains real-time inference at 80+ FPS on an NVIDIA RTX A8000 GPU while achieving competitive or superior performance. S2DENet not only establishes a new paradigm for ultrasound image segmentation but also opens new possibilities for deployment in clinical practice. Code is available at https://github.com/PXinTao/S2DENet .
AB - Ultrasound image segmentation serves as a cornerstone of clinical diagnosis, yet remains a formidable challenge due to inherent image artifacts such as speckle noise and ambiguous boundaries. Existing approaches typically employ uniform feature extraction strategies across all network layers, disregarding the fundamental disparities between noise-dominated shallow stages and semantically-rich deep stages. This monolithic strategy compels networks to concurrently learn noise suppression and feature enhancement, often resulting in over-parameterization and critical compromises in computational efficiency, particularly within resource-constrained environments. To address these limitations, we propose S2DENet, an efficient network architecture that employs noise suppression in shallow layers and semantic feature enhancement in deep layers. Specifically, we design Multi-order Differential Convolution (MDiffConv) to enhance high-frequency feature capture, implementing suppression strategies in shallow layers and enhancement strategies in deep layers. Simultaneously, we introduce a Differential Self-Attention (DiffSA) mechanism to mitigate noise artifacts while preserving structural integrity, with deep layers transitioning to standard self-attention for semantic feature amplification. This depth-differentiated design, combined with differential mechanisms, enables the model to focus on specific tasks at different network stages, thereby reducing learning complexity and improving feature representation efficiency. Experimental results on ten public ultrasound datasets demonstrate that S2DENet achieves an exceptional balance between efficiency and accuracy, attaining state-of-the-art (SOTA) performance on nine public datasets. With merely 0.05M/0.15M parameters (a reduction exceeding 99% compared to SOTA methods), S2DENet maintains real-time inference at 80+ FPS on an NVIDIA RTX A8000 GPU while achieving competitive or superior performance. S2DENet not only establishes a new paradigm for ultrasound image segmentation but also opens new possibilities for deployment in clinical practice. Code is available at https://github.com/PXinTao/S2DENet .
KW - Differential mechanism
KW - Hierarchical processing
KW - Lightweight method
KW - Ultrasound image segmentation
UR - https://www.scopus.com/pages/publications/105045300340
U2 - 10.1016/j.media.2026.104224
DO - 10.1016/j.media.2026.104224
M3 - Article
AN - SCOPUS:105045300340
SN - 1361-8415
VL - 113
JO - Medical Image Analysis
JF - Medical Image Analysis
M1 - 104224
ER -