Abstract
When using multiple visible-light cameras for 3D human pose estimation(HPE), the scarcity of such datasets arises because acquiring authentic 3D data from visible-light cameras is both complex and costly. This impedes the translation of pose estimation to real-world applications. To address these challenges, we propose GPDPose–a novel self-supervised method with Geometry, Pose, and Depth Consistency for multi-view 3D HPE. The GPDPose introduces three core components: First, we propose the Spatial-Temporal Feature Extraction Module (STEM), which performs coarse extraction of spatial-temporal features for each view in a decoupled manner to avoid noise interference between views. Subsequently, it conducts fine extraction on the fused spatial-temporal features to enhance accuracy. Second, we designed a Feature Fusion Module (FFM) that performs implicit feature matching to facilitate effective interaction and complementarity among multi-camera features. Finally, we propose a self-supervision strategy based on multi-camera geometric consistency, pose consistency, and depth consistency. Specifically, geometric consistency is used to compute pseudo labels, while pose and depth consistency reduce pseudo-label errors caused by occlusions. The integration of these constraints significantly enhances model performance. Results from experiments on two public datasets indicate that our approach attains SOTA performance in the self-supervised domain, reducing accuracy errors to competitive levels and even outperforming some fully supervised methods.
| Original language | English |
|---|---|
| Article number | 132909 |
| Journal | Expert Systems with Applications |
| Volume | 328 |
| DOIs | |
| Publication status | Published - 1 Oct 2026 |
Keywords
- Multi-view 3D human pose estimation
- Self-supervised learning
- Spatial-temporal feature
- Transformer
Fingerprint
Dive into the research topics of 'GPDPose: Self-supervised transformer with geometry, pose, and depth consistency for multi-view 3D human pose estimation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver