构建首个穿戴式设备目标推理基准,用多模态数据评估模型理解用户意图能力。
Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents
- 基于3477段多模态数据构建新基准WAGIBench,含视觉、音频、数字及时间序列信息。
- 人类准确率达93%,最优视觉语言模型仅84%,生成目标相关性不足55%。
- 多模态融合有效提升性能,无关模态干扰小,适合智能助老/辅助设备研究者。
近年来,助人可穿戴代理(如智能眼镜)受到关注,其能根据用户需求执行辅助动作。本文聚焦于从多模态上下文观察中推断用户目标的“目标推理”问题,以减少交互负担。为此,我们构建了WAGIBench,一个基于视觉语言模型(VLMs)的强基准。由于该领域研究有限,我们收集了来自348名参与者、总计29小时的多模态数据,涵盖3,477段记录,包含真实目标及对应的视觉、音频、数字与纵向上下文信息。结果显示,人类在多选任务中达到93%准确率,而最佳VLM仅为84%。生成式评估表明,更大模型表现更优,但仍远未实用——仅55%时间内生成相关目标。模态消融分析显示,相关模态提供显著增益,而无关模态影响极小。
原文摘要 · Abstract (English)
There has been a surge of interest in assistive wearable agents: agents embodied in wearable form factors (e.g., smart glasses) who take assistive actions toward a user's goal/query (e.g. "Where did I leave my keys?"). In this work, we consider the important complementary problem of inferring that goal from multi-modal contextual observations. Solving this "goal inference" problem holds the promise of eliminating the effort needed to interact with such an agent. This work focuses on creating WAGIBench, a strong benchmark to measure progress in solving this problem using vision-language models (VLMs). Given the limited prior work in this area, we collected a novel dataset comprising 29 hours of multimodal data from 348 participants across 3,477 recordings, featuring ground-truth goals alongside accompanying visual, audio, digital, and longitudinal contextual observations. We validate that human performance exceeds model performance, achieving 93% multiple-choice accuracy compared with 84% for the best-performing VLM. Generative benchmark results that evaluate several families of modern vision-language models show that larger models perform significantly better on the task, yet remain far from practical usefulness, as they produce relevant goals only 55% of the time. Through a modality ablation, we show that models benefit from extra information in relevant modalities with minimal performance degradation from irrelevant modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。