Abstract
The rapid growth of long-form user-generated videos has highlighted the challenges in Video Question Answering (VideoQA), especially for long-form content with complex temporal relationships and event structures. Existing VideoQA methods typically follow a forward reasoning paradigm and struggle to effectively model both short-term and long-term temporal dependencies, making the probe of question-related visual content challenging. To address these issues, we propose event-aware temporal modeling and semantic alignment (ESA), a novel framework designed to learn temporal representations and align question-related semantics in long-form VideoQA scenarios. ESA introduces event-aware temporal modeling to enhance temporal consistency and incorporate event information into frame features. Next, we design question relevance enhancement to progressively capture question-related features while compressing redundant information. Additionally, we introduce answer-driven semantic alignment that treats the forward prediction as a feedback pseudo-answer to improve semantic correlation among “video-question-answer” and facilitates answer reasoning. Extensive experiments on three long-form VideoQA datasets with various durations demonstrate that our proposed ESA method achieves state-of-the-art performance.
| Original language | English |
|---|---|
| Article number | 114369 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| Publication status | Published - Dec 2026 |
Keywords
- Long-form video question answering
- Multimodal fusion
- Semantic alignment
- Temporal modeling
Fingerprint
Dive into the research topics of 'Event-aware temporal modeling and semantic alignment for long-form video question answering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver