TY - JOUR
T1 - GPDPose
T2 - Self-supervised transformer with geometry, pose, and depth consistency for multi-view 3D human pose estimation
AU - Song, Jucheng
AU - Zhang, Jie
AU - Yang, Xu
AU - wang, Yapeng
AU - Gao, Hao
AU - Li, Haolun
AU - Im, Sio Kei
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/10/1
Y1 - 2026/10/1
N2 - When using multiple visible-light cameras for 3D human pose estimation(HPE), the scarcity of such datasets arises because acquiring authentic 3D data from visible-light cameras is both complex and costly. This impedes the translation of pose estimation to real-world applications. To address these challenges, we propose GPDPose–a novel self-supervised method with Geometry, Pose, and Depth Consistency for multi-view 3D HPE. The GPDPose introduces three core components: First, we propose the Spatial-Temporal Feature Extraction Module (STEM), which performs coarse extraction of spatial-temporal features for each view in a decoupled manner to avoid noise interference between views. Subsequently, it conducts fine extraction on the fused spatial-temporal features to enhance accuracy. Second, we designed a Feature Fusion Module (FFM) that performs implicit feature matching to facilitate effective interaction and complementarity among multi-camera features. Finally, we propose a self-supervision strategy based on multi-camera geometric consistency, pose consistency, and depth consistency. Specifically, geometric consistency is used to compute pseudo labels, while pose and depth consistency reduce pseudo-label errors caused by occlusions. The integration of these constraints significantly enhances model performance. Results from experiments on two public datasets indicate that our approach attains SOTA performance in the self-supervised domain, reducing accuracy errors to competitive levels and even outperforming some fully supervised methods.
AB - When using multiple visible-light cameras for 3D human pose estimation(HPE), the scarcity of such datasets arises because acquiring authentic 3D data from visible-light cameras is both complex and costly. This impedes the translation of pose estimation to real-world applications. To address these challenges, we propose GPDPose–a novel self-supervised method with Geometry, Pose, and Depth Consistency for multi-view 3D HPE. The GPDPose introduces three core components: First, we propose the Spatial-Temporal Feature Extraction Module (STEM), which performs coarse extraction of spatial-temporal features for each view in a decoupled manner to avoid noise interference between views. Subsequently, it conducts fine extraction on the fused spatial-temporal features to enhance accuracy. Second, we designed a Feature Fusion Module (FFM) that performs implicit feature matching to facilitate effective interaction and complementarity among multi-camera features. Finally, we propose a self-supervision strategy based on multi-camera geometric consistency, pose consistency, and depth consistency. Specifically, geometric consistency is used to compute pseudo labels, while pose and depth consistency reduce pseudo-label errors caused by occlusions. The integration of these constraints significantly enhances model performance. Results from experiments on two public datasets indicate that our approach attains SOTA performance in the self-supervised domain, reducing accuracy errors to competitive levels and even outperforming some fully supervised methods.
KW - Multi-view 3D human pose estimation
KW - Self-supervised learning
KW - Spatial-temporal feature
KW - Transformer
UR - https://www.scopus.com/pages/publications/105039815672
U2 - 10.1016/j.eswa.2026.132909
DO - 10.1016/j.eswa.2026.132909
M3 - Article
AN - SCOPUS:105039815672
SN - 0957-4174
VL - 328
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 132909
ER -