用实时眼动数据提升大模型对开放问答的歧义消解能力
IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models
- 基于眼动追踪数据,在推理时模拟注视动作来解析问题歧义
- 在模糊问题上准确率从35.2%提升至77.2%,未牺牲清晰问题表现
- 适配各类主流视觉语言模型,适用于需要交互式理解的场景
我们提出IRIS(通过推理时注视行为实现意图解析),一种无需训练的新方法,利用实时眼动数据解决大视觉语言模型中开放问答的歧义问题。通过包含500个独特图像-问题对的用户研究,我们发现参与者开始口头提问时最接近的注视点最具信息量,引入该数据使模糊问题的回答准确率从35.2%提升至77.2%,同时保持对非模糊问题的性能。我们在多个前沿视觉语言模型上验证了该方法,结果显示无论架构差异如何,加入眼动数据均能稳定提升模糊情境下的表现。我们发布了新的基准数据集、实时交互协议及评估套件。
原文摘要 · Abstract (English)
We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolve ambiguity in open-ended VQA. Through a comprehensive user study with 500 unique image-question pairs, we demonstrate that fixations closest to the time participants start verbally asking their questions are the most informative for disambiguation in Large VLMs, more than doubling the accuracy of responses on ambiguous questions (from 35.2% to 77.2%) while maintaining performance on unambiguous queries. We evaluate our approach across state-of-the-art VLMs, showing consistent improvements when gaze data is incorporated in ambiguous image-question pairs, regardless of architectural differences. We release a new benchmark dataset to use eye movement data for disambiguated VQA, a novel real-time interactive protocol, and an evaluation suite.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。