arXiv:2603.12701cs.HCcs.AI2026-03中稿 · ACM CHI 2026, Barc…被引 2

让AI看用户视角,提升人机协作效率与信任

Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration

  • 用第一人称视角对齐认知,实现自然协作
  • 任务完成时间缩短,交互负担降低,信任度上升
  • 适合人机协同、AR辅助等场景的开发者参考

尽管多模态AI取得进展,当前基于视觉的助手在协作任务中仍效率低下。我们识别出两个关键差距:沟通鸿沟——用户需将丰富的并行意图转化为语言指令,因通道不匹配;理解鸿沟——AI难以解析细微的具身线索。为此,我们提出Eye2Eye框架,利用第一人称视角作为人机认知对齐的通道。该框架集成三个组件:(1)联合注意力协调,实现流畅焦点对齐;(2)可修订记忆,维持动态共享认知基础;(3)反思反馈机制,支持用户澄清和修正AI理解。我们在AR原型中实现该框架,并通过用户研究与后验流程评估验证。结果表明,Eye2Eye显著缩短任务完成时间、降低交互负荷,同时提升信任度,证明各组件协同有效提升协作质量。

原文摘要 · Abstract (English)

Despite advances in multimodal AI, current vision-based assistants often remain inefficient in collaborative tasks. We identify two key gulfs: a communication gulf, where users must translate rich parallel intentions into verbal commands due to the channel mismatch , and an understanding gulf, where AI struggles to interpret subtle embodied cues. To address these, we propose Eye2Eye, a framework that leverages first-person perspective as a channel for human-AI cognitive alignment. It integrates three components: (1) joint attention coordination for fluid focus alignment, (2) revisable memory to maintain evolving common ground, and (3) reflective feedback allowing users to clarify and refine AI's understanding. We implement this framework in an AR prototype and evaluate it through a user study and a post-hoc pipeline evaluation. Results show that Eye2Eye significantly reduces task completion time and interaction load while increasing trust, demonstrating its components work in concert to improve collaboration.

人机协作第一人称视角认知对齐AR交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。