arXiv:2604.08062cs.HCcs.AI2026-04被引 1

用眼动追踪提升AI助手对用户认知困难的识别与响应能力

From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants

  • 结合眼动与视频流,让大模型理解用户阅读时的认知瓶颈
  • 实验显示眼动感知助手显著提升信息回忆率与交互效率
  • 适合教育、无障碍辅助等需要个性化认知支持的场景

当前大语言模型助手虽能回答问题,但缺乏对用户行为上下文的感知,难以发现其在何处或何时遇到困难。本文提出一种基于眼动的多模态大模型助手,利用第一视角视频与眼动叠加信息,识别用户可能的认知难点,并提供针对性的回顾性辅助。在一项受控研究中(n=36),对比了眼动感知助手与纯文本大模型助手的表现。结果显示,相比传统助手,眼动感知助手在评估用户阅读行为时更准确、更具个性化,显著提升了用户的记忆能力;且用户与其交互时说话字数明显减少,表明交互更高效。定性分析也证实了其在理解力提升方面的价值,但当眼动解读错误时仍存在挑战。结果表明,眼动感知的大模型助手能够推理认知需求,从而改善用户认知表现。

原文摘要 · Abstract (English)

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant that uses egocentric video with gaze overlays to identify likely points of difficulty and target follow-up retrospective assistance. We instantiate this vision in a controlled study (n=36) comparing the gaze-aware AI assistant to a text-only LLM assistant. Compared to a conventional LLM assistant, the gaze-aware assistant was rated as significantly more accurate and personalized in its assessments of users' reading behavior and significantly improved people's ability to recall information. Users spoke significantly fewer words with the gaze-aware assistant, indicating more efficient interactions. Qualitative results underscored both perceived benefits in comprehension and challenges when interpretations of gaze behaviors were inaccurate. Our findings suggest that gaze-aware LLM assistants can reason about cognitive needs to improve cognitive outcomes of users.

多模态眼动追踪认知辅助大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。