All You May Need for VQA are Image Captions

Soravit Changpinyo,Doron Kukliansky,Idan Szpektor,Xi Chen,Nan Ding,Radu Soricut

North American Chapter of the Association for Computational Linguistics (NAACL)（2022）

引用 40|浏览65

暂无评分

摘要

Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. In this paper, we propose a method that automatically derives VQA examples at volume, by leveraging the abundance of existing image-caption annotations combined with neural models for textual question generation. We show that the resulting data is of high-quality. VQA models trained on our data improve state-of-the-art zero-shot accuracy by double digits and achieve a level of robustness that lacks in the same model trained on human-annotated VQA data.

查看译文

关键词

vqa,image,need

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要