摘要
The rapid growth of long-form user-generated videos has highlighted the challenges in Video Question Answering (VideoQA), especially for long-form content with complex temporal relationships and event structures. Existing VideoQA methods typically follow a forward reasoning paradigm and struggle to effectively model both short-term and long-term temporal dependencies, making the probe of question-related visual content challenging. To address these issues, we propose event-aware temporal modeling and semantic alignment (ESA), a novel framework designed to learn temporal representations and align question-related semantics in long-form VideoQA scenarios. ESA introduces event-aware temporal modeling to enhance temporal consistency and incorporate event information into frame features. Next, we design question relevance enhancement to progressively capture question-related features while compressing redundant information. Additionally, we introduce answer-driven semantic alignment that treats the forward prediction as a feedback pseudo-answer to improve semantic correlation among “video-question-answer” and facilitates answer reasoning. Extensive experiments on three long-form VideoQA datasets with various durations demonstrate that our proposed ESA method achieves state-of-the-art performance.
| 原文 | English |
|---|---|
| 文章編號 | 114369 |
| 期刊 | Pattern Recognition |
| 卷 | 180 |
| DOIs | |
| 出版狀態 | Published - 12月 2026 |
指紋
深入研究「Event-aware temporal modeling and semantic alignment for long-form video question answering」主題。共同形成了獨特的指紋。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver