让多个智能体的视角视频协同理解,提升人机协作效率
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
- 设计多智能体视角视频问答基准MA-EgoQA,支持跨流信息融合
- 提出EgoMAS模型,通过共享记忆与动态检索提升跨视角理解
- 涵盖5类复杂推理任务,适合研究多智能体系统认知建模者参考
随着具身智能体日益强大,未来人类将在工作或家庭中与多个智能体协同。为实现高效沟通,需同时理解多个智能体传来的视觉信息,并为每个问题匹配恰当上下文。现有挑战在于如何压缩并传输大量个体感官输入(视频),以及如何聚合多条第一人称视频构建系统级记忆。本文首次正式定义了同时理解多条长时序第一人称视频的全新任务。为此,我们构建了多智能体第一人称视频问答基准MA-EgoQA,包含1.7k个独特问题,覆盖五类:社交互动、任务协调、心智理论、时间推理和环境交互。我们还提出简单基线模型EgoMAS,利用跨智能体共享记忆与代理专属动态检索机制。在多种基线与EgoMAS上的全面评估表明,当前方法难以有效处理多视角视频流,凸显了系统级理解能力亟待提升。代码与数据集已开源。
原文摘要 · Abstract (English)
As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human users and the multi-agent system, it is crucial to interpret incoming information from agents in parallel and refer to the appropriate context for each query. Existing challenges include effectively compressing and communicating high volumes of individual sensory inputs in the form of video and correctly aggregating multiple egocentric videos to construct system-level memory. In this work, we first formally define a novel problem of understanding multiple long-horizon egocentric videos simultaneously collected from embodied agents. To facilitate research in this direction, we introduce MultiAgent-EgoQA (MA-EgoQA), a benchmark designed to systemically evaluate existing models in our scenario. MA-EgoQA provides 1.7k questions unique to multiple egocentric streams, spanning five categories: social interaction, task coordination, theory-of-mind, temporal reasoning, and environmental interaction. We further propose a simple baseline model for MA-EgoQA named EgoMAS, which leverages shared memory across embodied agents and agent-wise dynamic retrieval. Through comprehensive evaluation across diverse baselines and EgoMAS on MA-EgoQA, we find that current approaches are unable to effectively handle multiple egocentric streams, highlighting the need for future advances in system-level understanding across the agents. The code and benchmark are available at https://ma-egoqa.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。