Skip to main navigation Skip to search Skip to main content

ECTFormer: Efficient CNN-Transformer Network for Uncalibrated Multiview 3-D Human Pose Estimation

  • Macao Polytechnic University

Research output: Contribution to journalArticlepeer-review

1 Citation (Scopus)

Abstract

Single-view camera sensors suffer from inherent depth blurring issues that hinder progress in 3-D human pose estimation (HPE), sparking widespread research interest in multiview camera sensor systems. However, existing methods typically rely on complex camera calibration processes and are sensitive to dynamic environments. To address these limitations, we propose ECTFormer, an innovative calibration-free multiview 3-D HPE framework that seamlessly combines Transformer-based spatiotemporal modeling with convolutional neural network (CNN)-based local feature extraction. The primary contents of this article are: Firstly, we introduce a hierarchical multiview spatiotemporal feature extraction network. This network avoids interference between noise from different viewpoints through hierarchical learning and employs a transformer to capture spatio-temporal features within views for subsequent fusion. And then, we design a CNN-transformer fusion module (CTFM) that efficiently aggregates multiview features, enabling accurate 3-D pose regression. Extensive experiments on a public 3-D human pose benchmark demonstrate that our approach attains superior performance without relying on calibration. For real-world environments equipped with dynamic camera sensors, ECTFormer can efficiently perform HPE.

Original languageEnglish
Pages (from-to)8487-8498
Number of pages12
JournalIEEE Sensors Journal
Volume26
Issue number6
DOIs
Publication statusPublished - 15 Mar 2026

Keywords

  • Convolutional neural network (CNN)
  • multiview 3-D human pose estimation (HPE)
  • transformer
  • uncalibrated camera

Fingerprint

Dive into the research topics of 'ECTFormer: Efficient CNN-Transformer Network for Uncalibrated Multiview 3-D Human Pose Estimation'. Together they form a unique fingerprint.

Cite this