arXiv:2607.00547cs.CVcs.AI2026-07

提出新基准EgoGapBench,评估多智能体场景下第一视角动作选择能力。

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

论文配图:EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes
图 1 · 摘自论文原文
  • 构建诊断性基准,分离第一视角理解与动作选择能力。
  • 人类表现可靠,大模型普遍误选他人动作,差距显著。
  • 仅靠第一视角数据微调无效,需专门训练提升视角决策能力。

现有第一人称基准多基于第一视角数据构建,难以独立评估第一视角理解能力。而第一视角输入与视角认知是可分离的能力,尤其在缺乏身体线索或存在其他智能体时更明显。为孤立评估第一视角理解,我们引入EgoGapBench,一个用于衡量多智能体第一视角场景中动作选择的诊断性基准。该基准衡量的能力称为第一视角动作选择(Egocentric Action Selection, EAS):在其他智能体存在时,从自身视角选择合适动作。在EgoGapBench上,人类表现稳定可靠,而开源和专有多模态大模型表现显著更差,系统性地选择可见他者执行的动作。在现有第一视角数据上微调无法缩小差距,甚至可能恶化。相比之下,在EgoGapBench训练数据上微调可提升准确率,但仍未达到人类水平。结果表明,仅从第一视角数据中学习难以掌握EAS能力,大模型评估与训练应不仅关注场景理解,还需专门强化第一视角动作选择能力。

原文摘要 · Abstract (English)

Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation. However, understanding first-person-view input and taking an egocentric perspective are separable abilities, especially when first-person body cues are absent or when other agents are present. To isolate egocentric perspective understanding, we introduce EgoGapBench, a diagnostic benchmark for measuring action selection in multi-agent egocentric scenes. We define the ability measured by this benchmark as Egocentric Action Selection (EAS): selecting an appropriate action from the agent's perspective in the presence of other agents. On EgoGapBench, humans answer reliably, whereas both open-source and proprietary MLLMs perform substantially worse and systematically select actions performed by other visible agents. Fine-tuning on existing egocentric data fails to close this gap and can even be detrimental. In contrast, fine-tuning on EgoGapBench training data improves accuracy but does not reach human performance. These results show that EAS is difficult to acquire from first-person-view data alone, and that MLLMs should be evaluated and trained not only for scene understanding but also for egocentric action selection.

多智能体第一视角动作选择基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。