Abstract
Single-view camera sensors suffer from inherent depth blurring issues that hinder progress in 3-D human pose estimation (HPE), sparking widespread research interest in multiview camera sensor systems. However, existing methods typically rely on complex camera calibration processes and are sensitive to dynamic environments. To address these limitations, we propose ECTFormer, an innovative calibration-free multiview 3-D HPE framework that seamlessly combines Transformer-based spatiotemporal modeling with convolutional neural network (CNN)-based local feature extraction. The primary contents of this article are: Firstly, we introduce a hierarchical multiview spatiotemporal feature extraction network. This network avoids interference between noise from different viewpoints through hierarchical learning and employs a transformer to capture spatio-temporal features within views for subsequent fusion. And then, we design a CNN-transformer fusion module (CTFM) that efficiently aggregates multiview features, enabling accurate 3-D pose regression. Extensive experiments on a public 3-D human pose benchmark demonstrate that our approach attains superior performance without relying on calibration. For real-world environments equipped with dynamic camera sensors, ECTFormer can efficiently perform HPE.
| Original language | English |
|---|---|
| Pages (from-to) | 8487-8498 |
| Number of pages | 12 |
| Journal | IEEE Sensors Journal |
| Volume | 26 |
| Issue number | 6 |
| DOIs | |
| Publication status | Published - 15 Mar 2026 |
Keywords
- Convolutional neural network (CNN)
- multiview 3-D human pose estimation (HPE)
- transformer
- uncalibrated camera
Fingerprint
Dive into the research topics of 'ECTFormer: Efficient CNN-Transformer Network for Uncalibrated Multiview 3-D Human Pose Estimation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver