AI-VQA: Visual Question Answering based on Agent Interaction with Interpretability

Rengang Li, Cong Xu, Zhenhua Guo, Baoyu Fan, Runze Zhang, Wei Liu, Yaqian Zhao, Weifeng Gong, Endong Wang

研究成果: Conference contribution同行評審

7 引文 斯高帕斯(Scopus)

摘要

Visual Question Answering (VQA) serves as a proxy for evaluating the scene understanding of an intelligent agent by answering questions about images. Most VQA benchmarks to date are focused on those questions that can be answered through understanding visual content in the scene, such as simple counting, visual attributes, and even a little challenging questions that require extra encyclopedic knowledge. However, humans have a remarkable capacity to reason dynamic interaction on the scene, which is beyond the literal content of an image and has not been investigated so far. In this paper, we propose Agent Interaction Visual Question Answering (AI-VQA), a task investigating deep scene understanding if the agent takes a certain action. For this task, a model not only needs to answer action-related questions but also to locate the objects in which the interaction occurs for guaranteeing it truly comprehends the action. Accordingly, we make a new dataset based on Visual Genome and ATOMIC knowledge graph, including more than 19,000 manually annotated questions, and will make it publicly available. Besides, we also provide an annotation of the reasoning path while developing the answer for each question. Based on the dataset, we further propose a novel method, called ARE, that can comprehend the interaction and explain the reason based on a given event knowledge base. Experimental results show that our proposed method outperforms the baseline by a clear margin.

原文English
主出版物標題MM 2022 - Proceedings of the 30th ACM International Conference on Multimedia
發行者Association for Computing Machinery, Inc
頁面5274-5282
頁數9
ISBN(電子)9781450392037
DOIs
出版狀態Published - 10 10月 2022
對外發佈
事件30th ACM International Conference on Multimedia, MM 2022 - Lisboa, Portugal
持續時間: 10 10月 202214 10月 2022

出版系列

名字MM 2022 - Proceedings of the 30th ACM International Conference on Multimedia

Conference

Conference30th ACM International Conference on Multimedia, MM 2022
國家/地區Portugal
城市Lisboa
期間10/10/2214/10/22

指紋

深入研究「AI-VQA: Visual Question Answering based on Agent Interaction with Interpretability」主題。共同形成了獨特的指紋。

引用此