arXiv:2603.00490cs.AI2026-03被引 2

构建首个从第一视角评估人机协作的多模态基准

LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks

  • 基于第一人称视频流,评估实时交互中的任务理解能力
  • 涵盖4075个问答对,覆盖6大核心能力维度
  • 适合研究人机协同、智能助手机器人方向的学者

多模态大语言模型(MLLMs)的发展为增强人类能力提供了巨大潜力,但其在动态真实环境中的有效辅助能力仍鲜有研究。现有视频基准多聚焦于事后分析或孤立感知任务,难以反映实时人机协作的互动性与适应性。为此,我们提出LifeEval,一个面向第一人称日常任务场景的多模态基准,旨在评估实时、任务导向的人机协同。该基准强调三方面:任务导向的整体评估、来自连续第一人称视频流的实时感知,以及通过自然对话实现的人机协作。通过严格标注流程,构建了包含4,075个高质量问答对的数据集,覆盖6个核心能力维度。对26个前沿MLLMs在LifeEval上的评测显示,实现及时、有效且自适应的交互仍面临巨大挑战,指明了以人为中心的交互智能发展的关键方向。

原文摘要 · Abstract (English)

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective assistance in dynamic, real-world environments remains largely underexplored. Existing video benchmarks predominantly assess passive understanding through retrospective analysis or isolated perception tasks, failing to capture the interactive and adaptive nature of real-time user assistance. To bridge this gap, we introduce LifeEval, a multimodal benchmark designed to evaluate real-time, task-oriented human-AI collaboration in daily life from an egocentric perspective. LifeEval emphasizes three key aspects: task-oriented holistic evaluation, egocentric real-time perception from continuous first-person streams, and human-assistant collaborative interaction through natural dialogues. Constructed via a rigorous annotation pipeline, the benchmark comprises 4,075 high-quality question-answer pairs across 6 core capability dimensions. Extensive evaluations of 26 state-of-the-art MLLMs on LifeEval reveal substantial challenges in achieving timely, effective and adaptive interaction, highlighting essential directions for advancing human-centered interactive intelligence.

人机协作多模态第一人称基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。