跳至主導覽 跳至搜尋 跳過主要內容

Event-aware temporal modeling and semantic alignment for long-form video question answering

  • Xichun Sheng
  • , Haibo Gong
  • , Lian Zhang
  • , Liang Li
  • , Jiehua Zhang
  • , Chenggang Yan
  • , Tao Tan
  • Macao Polytechnic University
  • Hangzhou Dianzi University
  • Hebei Medical University
  • Hebei Institute of Clinical Artificial Intelligence
  • CAS - Institute of Computing Technology
  • Xi'an Jiaotong University

研究成果: Article同行評審

摘要

The rapid growth of long-form user-generated videos has highlighted the challenges in Video Question Answering (VideoQA), especially for long-form content with complex temporal relationships and event structures. Existing VideoQA methods typically follow a forward reasoning paradigm and struggle to effectively model both short-term and long-term temporal dependencies, making the probe of question-related visual content challenging. To address these issues, we propose event-aware temporal modeling and semantic alignment (ESA), a novel framework designed to learn temporal representations and align question-related semantics in long-form VideoQA scenarios. ESA introduces event-aware temporal modeling to enhance temporal consistency and incorporate event information into frame features. Next, we design question relevance enhancement to progressively capture question-related features while compressing redundant information. Additionally, we introduce answer-driven semantic alignment that treats the forward prediction as a feedback pseudo-answer to improve semantic correlation among “video-question-answer” and facilitates answer reasoning. Extensive experiments on three long-form VideoQA datasets with various durations demonstrate that our proposed ESA method achieves state-of-the-art performance.

原文English
文章編號114369
期刊Pattern Recognition
180
DOIs
出版狀態Published - 12月 2026

指紋

深入研究「Event-aware temporal modeling and semantic alignment for long-form video question answering」主題。共同形成了獨特的指紋。

引用此