Skip to main navigation Skip to search Skip to main content

GPDPose: Self-supervised transformer with geometry, pose, and depth consistency for multi-view 3D human pose estimation

  • Macao Polytechnic University
  • Nanjing University of Posts and Telecommunications

Research output: Contribution to journalArticlepeer-review

Abstract

When using multiple visible-light cameras for 3D human pose estimation(HPE), the scarcity of such datasets arises because acquiring authentic 3D data from visible-light cameras is both complex and costly. This impedes the translation of pose estimation to real-world applications. To address these challenges, we propose GPDPose–a novel self-supervised method with Geometry, Pose, and Depth Consistency for multi-view 3D HPE. The GPDPose introduces three core components: First, we propose the Spatial-Temporal Feature Extraction Module (STEM), which performs coarse extraction of spatial-temporal features for each view in a decoupled manner to avoid noise interference between views. Subsequently, it conducts fine extraction on the fused spatial-temporal features to enhance accuracy. Second, we designed a Feature Fusion Module (FFM) that performs implicit feature matching to facilitate effective interaction and complementarity among multi-camera features. Finally, we propose a self-supervision strategy based on multi-camera geometric consistency, pose consistency, and depth consistency. Specifically, geometric consistency is used to compute pseudo labels, while pose and depth consistency reduce pseudo-label errors caused by occlusions. The integration of these constraints significantly enhances model performance. Results from experiments on two public datasets indicate that our approach attains SOTA performance in the self-supervised domain, reducing accuracy errors to competitive levels and even outperforming some fully supervised methods.

Original languageEnglish
Article number132909
JournalExpert Systems with Applications
Volume328
DOIs
Publication statusPublished - 1 Oct 2026

Keywords

  • Multi-view 3D human pose estimation
  • Self-supervised learning
  • Spatial-temporal feature
  • Transformer

Fingerprint

Dive into the research topics of 'GPDPose: Self-supervised transformer with geometry, pose, and depth consistency for multi-view 3D human pose estimation'. Together they form a unique fingerprint.

Cite this