测试视觉语言模型实时问答能力,发现仍远低于人类水平。
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
- 构建实时音视频问答数据集IVD,评估模型现场交互能力。
- 现有模型表现远逊于人类,尤其在动态场景理解上差距明显。
- 微调可显著提升模型感知能力,适合开发真实世界助手。
近年来,人工智能模型在描述和回答现实图像问题方面取得了显著进展,并在使用音频输入与用户实时对话方面有所突破。这引出了一个关键问题:当模型连接摄像头和麦克风时,能否实时对镜头前正在发生的场景和事件进行对话?这一目标是实现真实世界智能助手和人形机器人的日常交互前提。本文提出新的数据集与基准测试——高通互动视频数据集(IVD),用于评估现有模型在该任务上的表现,并探索通过微调提升能力的潜力。该数据集采用简单的问答形式,用户提问,系统需基于实时摄像头与音频输入作答。实验表明,现有模型在该任务上的表现远落后于人类,且我们识别出主要性能差距来源。然而,对于许多所需感知技能,通过对此类数据进行微调可显著缩小差距。
原文摘要 · Abstract (English)
AI models have made significant strides in recent years in their ability to describe and answer questions about real-world images. They have also made progress in the ability to converse with users in real-time using audio input. This raises the question: have we reached the point where AI models, connected to a camera and microphone, can converse with users in real-time about scenes and events that are unfolding live in front of the camera? This has been a long-standing goal in AI and is a prerequisite for real-world AI assistants and humanoid robots to interact with humans in everyday situations. In this work, we introduce a new dataset and benchmark, the Qualcomm Interactive Video Dataset (IVD), which allows us to assess the extent to which existing models can support these abilities, and to what degree these capabilities can be instilled through fine-tuning. The dataset is based on a simple question-answering setup, where users ask questions that the system has to answer, in real-time, based on the camera and audio input. We show that existing models fall far behind human performance on this task, and we identify the main sources for the performance gap. However, we also show that for many of the required perceptual skills, fine-tuning on this form of data can significantly reduce this gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。