Skip to main navigation Skip to search Skip to main content

Event-aware temporal modeling and semantic alignment for long-form video question answering

  • Xichun Sheng
  • , Haibo Gong
  • , Lian Zhang
  • , Liang Li
  • , Jiehua Zhang
  • , Chenggang Yan
  • , Tao Tan
  • Macao Polytechnic University
  • Hangzhou Dianzi University
  • Hebei Medical University
  • Hebei Institute of Clinical Artificial Intelligence
  • CAS - Institute of Computing Technology
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

The rapid growth of long-form user-generated videos has highlighted the challenges in Video Question Answering (VideoQA), especially for long-form content with complex temporal relationships and event structures. Existing VideoQA methods typically follow a forward reasoning paradigm and struggle to effectively model both short-term and long-term temporal dependencies, making the probe of question-related visual content challenging. To address these issues, we propose event-aware temporal modeling and semantic alignment (ESA), a novel framework designed to learn temporal representations and align question-related semantics in long-form VideoQA scenarios. ESA introduces event-aware temporal modeling to enhance temporal consistency and incorporate event information into frame features. Next, we design question relevance enhancement to progressively capture question-related features while compressing redundant information. Additionally, we introduce answer-driven semantic alignment that treats the forward prediction as a feedback pseudo-answer to improve semantic correlation among “video-question-answer” and facilitates answer reasoning. Extensive experiments on three long-form VideoQA datasets with various durations demonstrate that our proposed ESA method achieves state-of-the-art performance.

Original languageEnglish
Article number114369
JournalPattern Recognition
Volume180
DOIs
Publication statusPublished - Dec 2026

Keywords

  • Long-form video question answering
  • Multimodal fusion
  • Semantic alignment
  • Temporal modeling

Fingerprint

Dive into the research topics of 'Event-aware temporal modeling and semantic alignment for long-form video question answering'. Together they form a unique fingerprint.

Cite this