让AI理解第一视角下的隐含意图,提升助手机能。
Visual Intention Grounding for Egocentric Assistants
- 构建首个第一人称视角意图定位数据集EgoIntention
- 模型对上下文物体误判率高,缺乏功能推理能力
- 提出链式推理训练法,兼顾显式查询与隐含意图
视觉定位将文本描述与图像中的物体关联。传统方法针对第三人称图像和明确物体查询。在如AI助手等应用中,视角转为第一人称,物体常通过需求和意图间接指代。为此,我们提出EgoIntention,首个第一人称视觉意图定位数据集。该数据集挑战多模态大模型:1)识别并忽略无关上下文物体;2)推理不常见物体的功能。基准测试显示,当前模型易误识上下文物体,且缺乏第一人称视角下的功能理解能力。我们还提出一种名为Reason-to-Ground(RoG)的指令微调方法,结合正常描述与第一人称意图进行混合训练,采用链式意图推理与物体定位机制。RoG显著优于简单微调和普通混合训练,在第一人称意图定位任务上表现更优,同时保持或轻微提升常规描述定位性能。该方法实现第一人称与第三人称输入的统一视觉定位,可处理显式物体查询与隐含人类意图。
原文摘要 · Abstract (English)
Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are egocentric, and objects may be referred to implicitly through needs and intentions. To bridge this gap, we introduce EgoIntention, the first dataset for egocentric visual intention grounding. EgoIntention challenges multimodal LLMs to 1) understand and ignore unintended contextual objects and 2) reason about uncommon object functionalities. Benchmark results show that current models misidentify context objects and lack affordance understanding in egocentric views. We also propose Reason-to-Ground (RoG) instruction tuning; it enables hybrid training with normal descriptions and egocentric intentions with a chained intention reasoning and object grounding mechanism. RoG significantly outperforms naive finetuning and hybrid training on EgoIntention, while maintaining or slightly improving naive description grounding. This advancement enables unified visual grounding for egocentric and exocentric visual inputs while handling explicit object queries and implicit human intentions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。