arXiv:2509.21980cs.CV2025-09被引 1

用眼动数据增强视觉助手交互,解决用户提问模糊与眼动噪声问题。

Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm

  • 融合时空眼动信息,设计新方法提升视觉模型对注意力的理解。
  • 在自建数据集上表现优于基线,显著改善模糊提问的识别准确率。
  • 适合智能眼镜、人机交互等需实时理解注意力的应用场景。

随着智能眼镜普及,用户注意力被整合进视觉语言模型(VLMs)以简化日常多模态查询。然而,利用眼动数据建模用户注意力会引入歧义挑战:(1) 用户口语提问常使用代词或省略上下文;(2) 人类眼动模式具有噪声且与话语存在复杂时空关系。现有工作仅使用单张图像作为视觉输入,难以捕捉注意力动态性。本文提出GLARIFY,一种利用时空眼动信息增强现实应用中模型效果的新方法。首先分析数百个带眼动的查询样本,揭示眼动噪声特性;随后通过GPT-4o设计自动数据合成流程,生成包含链式思维(CoT)的GLARIFY-Ambi数据集以处理噪声;最后设计热力图模块,将眼动信息融入前沿VLMs,同时保留其预训练知识。在留出测试集上的实验表明,GLARIFY显著优于基线。通过鲁棒对齐VLM与人类注意力,为视觉助手提供可落地、直观的交互范式。

原文摘要 · Abstract (English)

With the rise in popularity of smart glasses, users' attention has been integrated into Vision-Language Models (VLMs) to streamline multi-modal querying in daily scenarios. However, leveraging gaze data to model users' attention may introduce ambiguity challenges: (1) users' verbal questions become ambiguous by using pronouns or skipping context, (2) humans' gaze patterns can be noisy and exhibit complex spatiotemporal relationships with their spoken questions. Previous works only consider single image as visual modality input, failing to capture the dynamic nature of the user's attention. In this work, we introduce GLARIFY, a novel method to leverage spatiotemporal gaze information to enhance the model's effectiveness in real-world applications. Initially, we analyzed hundreds of querying samples with the gaze modality to demonstrate the noisy nature of users' gaze patterns. We then utilized GPT-4o to design an automatic data synthesis pipeline to generate the GLARIFY-Ambi dataset, which includes a dedicated chain-of-thought (CoT) process to handle noisy gaze patterns. Finally, we designed a heatmap module to incorporate gaze information into cutting-edge VLMs while preserving their pretrained knowledge. We evaluated GLARIFY using a hold-out test set. Experiments demonstrate that GLARIFY significantly outperforms baselines. By robustly aligning VLMs with human attention, GLARIFY paves the way for a usable and intuitive interaction paradigm with a visual assistant.

视觉语言模型眼动追踪人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。