arXiv:2605.07299cs.CVcs.AI2026-05

构建首个个性化主动交互基准,提升智能设备预判用户意图能力

EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams

论文配图:EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
图 1 · 摘自论文原文
  • 用模拟用户画像生成12类场景的主动交互数据
  • 在2400段视频上实现精准时机识别,准确率显著提升
  • 适合研发下一代主动式人机交互系统的研究者

现有多模态大模型仍以被动响应为主,无法持续感知环境或主动辅助用户。尽管已有基准尝试评估主动性,但多局限于警报场景,忽视个性化上下文,且未衡量人机交互(HMI)的精确时机。本文提出EgoPro-Bench,一个基于流式第一视角视频的主动交互评测基准,包含2,400段评估视频和超过12,000段训练视频。不同于以往工作,该基准通过模拟用户画像生成多样化用户意图,在12个不同领域构建高保真HMI数据。我们设计专用评估协议与指标,训练适用于流式视频的高效推理、低延迟交互模型,并引入“短思考,好交互”原则——在意图识别前分配有限令牌预算,提升交互性能。实验表明,EgoPro-Bench显著增强MLLM对意图的理解能力,可准确识别人机交互的合适时机,为下一代以用户为中心的主动交互代理奠定基础。

原文摘要 · Abstract (English)

Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users. While emerging benchmarks address proactivity, they are largely confined to alert scenarios, neglect personalized context, and fail to evaluate the precise timing of human-machine interactions (HMI).In this paper, we introduce EgoPro-Bench, a novel benchmark for training and evaluating proactive interaction capabilities based on streaming egocentric videos; it comprises 2,400 videos in the evaluation set and over 12,000 videos in the training set.Unlike previous works, EgoPro-Bench leverages simulated user profiles to generate diverse user intentions and to construct high-fidelity HMI data across 12 distinct domains.Subsequently, we propose a specialized evaluation protocol and metrics, train proactive interaction models designed for efficient reasoning and low-latency interaction on streaming video data, and conduct comprehensive evaluations.Furthermore, we introduce an interaction principle termed "short thinking, better interaction", which allocates a limited token budget prior to intent recognition, thereby enhancing interaction performance.The experiments demonstrate that EgoPro-Bench substantially enhances the intention understanding capabilities of MLLMs and enables accurate identification of appropriate timings for HMI, thereby laying a solid foundation for next-generation user-centric proactive interactive agents.

主动交互第一视角视频多模态模型人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。