TY - JOUR
T1 - SemanticST
T2 - Semantics-enhanced Spatio-Temporal Modeling for Ejection Fraction Estimation in Echocardiography
AU - Zheng, Dashun
AU - Li, Jiaxuan
AU - Meng, Mingyuan
AU - Ma, Tengfei
AU - Li, Wei
AU - Gao, Zhifan
AU - Pang, Patrick
AU - Tan, Tao
N1 - Publisher Copyright:
© 2013 IEEE.
PY - 2026
Y1 - 2026
N2 - Estimating left ventricular Ejection Fraction (EF) from echocardiography is critical for cardiac systolic function assessment and clinical risk stratification. Unfortunately, existing EF estimation methods are limited by (i) insufficient modeling of temporal clues embedded in video frames and (ii) inadequate exploitation of clinical semantics in easily accessible textual reports. In this study, we propose a unified segmentation and EF estimation framework with Semantics-enhanced spatio-temporal modeling (named SemanticST), integrating spatio-temporal consistency modeling with structured clinical semantic priors. Our SemanticST introduces a spatio-temporal & text-guided neighborhood correlation mining (STT-NCM) encoder, which captures both short- and long-range temporal dependencies via text-modulated 3D neighborhood attention. Further, a text-guided pixel-level semantic projection module (TextSP) is designed to map the key clinical cues extracted by a large language model (LLM) into pixel-level guidance features, enabling the alignment of semantic priors with visual context for optimized EF estimation. Extensive experiments on two public datasets (CAMUS and EchoNet-Dynamic) demonstrate that our SemanticST outperforms state-of-the-art methods in segmentation accuracy, temporal consistency, and EF estimation correlation.
AB - Estimating left ventricular Ejection Fraction (EF) from echocardiography is critical for cardiac systolic function assessment and clinical risk stratification. Unfortunately, existing EF estimation methods are limited by (i) insufficient modeling of temporal clues embedded in video frames and (ii) inadequate exploitation of clinical semantics in easily accessible textual reports. In this study, we propose a unified segmentation and EF estimation framework with Semantics-enhanced spatio-temporal modeling (named SemanticST), integrating spatio-temporal consistency modeling with structured clinical semantic priors. Our SemanticST introduces a spatio-temporal & text-guided neighborhood correlation mining (STT-NCM) encoder, which captures both short- and long-range temporal dependencies via text-modulated 3D neighborhood attention. Further, a text-guided pixel-level semantic projection module (TextSP) is designed to map the key clinical cues extracted by a large language model (LLM) into pixel-level guidance features, enabling the alignment of semantic priors with visual context for optimized EF estimation. Extensive experiments on two public datasets (CAMUS and EchoNet-Dynamic) demonstrate that our SemanticST outperforms state-of-the-art methods in segmentation accuracy, temporal consistency, and EF estimation correlation.
KW - Echocardiography
KW - Ejection fraction Estimation
KW - Large language model
KW - Spatio-temporal modeling
UR - https://www.scopus.com/pages/publications/105043047215
U2 - 10.1109/JBHI.2026.3703643
DO - 10.1109/JBHI.2026.3703643
M3 - Article
C2 - 42295964
AN - SCOPUS:105043047215
SN - 2168-2194
JO - IEEE Journal of Biomedical and Health Informatics
JF - IEEE Journal of Biomedical and Health Informatics
ER -