arXiv:2608.25561cs.CL2026-08中稿 · EMNLP

评测视觉语言模型在第一视角任务中的跨模态决策能力。

EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

论文配图:EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports
图 1 · 摘自论文原文
  • 构建五类日常场景的对话视频数据集,评估模型跨模态判断能力。
  • 现有模型在模态冲突时仍难做出可靠决策,准确率不足60%。
  • 揭示当前模态对齐方法局限性,适合关注AI助手部署的实践者。

视觉语言模型(VLMs)正被视作能感知第一人称环境、理解用户对话并决定如何协助的日常助手。现有第一视角基准测试主要孤立评估视觉理解能力,未考察模型在视觉证据与用户语言一致、无关或冲突时,能否有效权衡两者的可靠性。为此,我们提出了EgoArgus,一个由人工标注的、用于评估第一视角助手在五个对话-视频日常场景中理解与决策能力的数据集。实验结果表明,当前VLMs作为可靠的助手仍面临挑战,需准确识别可信模态并判断干预时机。深入分析还显示,现有缓解模态偏见的方法效果有限,为实际部署提供了重要启示。

原文摘要 · Abstract (English)

VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between visual evidence and user-provided language when the two are helpful, irrelevant, or conflicting. We introduce EgoArgus, a human-annotated dataset for evaluating egocentric assistants on understanding and decision tasks in five dialogue-video daily scenarios. Our results demonstrate that it is still challenging for current VLMs as reliable egocentric assistants, which requires identifying which modality is trustworthy and deciding when intervention is warranted. Deeper analysis also shows that existing modality bias mitigation methods are quite restricted to enhance performance, providing insights to aid practioners into the deployment of current VLMs as daily assistants.

视觉语言模型第一视角多模态决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。