arXiv:2605.14742cs.CVcs.RO2026-05中稿 · ICML被引 1

提出统一框架,让模型更准理解第一视角互动并精确定位像素。

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding

论文配图:EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
图 1 · 摘自论文原文
  • 分两阶段解析:先整体理解互动,再根据问题生成答案与像素掩码。
  • 在Ego-IRGBench上像素定位准确率达65.48%,比之前方法高8.37%。
  • 适合需要精准互动理解与视觉定位的智能机器人应用。

从第一视角视觉理解人与环境交互对辅助机器人和具身智能体至关重要,但现有多模态大语言模型在交互推理与细粒度像素定位方面仍存不足。本文提出EARL框架,一种面向第一视角交互推理与像素定位的分析引导强化学习方法,通过显式传递粗粒度交互语义至查询导向的回答与定位。该框架采用两阶段解析:第一阶段全局解析第一视角交互,生成结构化文本描述;第二阶段针对用户问题生成文本回答与像素级掩码。为衔接两阶段,提取全局交互描述符作为语义先验,并通过新型分析引导特征合成器(AFS)实现查询导向推理。为优化文本、边界框、定位掩码等异构输出,设计多维度奖励函数,并使用GRPO训练响应阶段。在Ego-IRGBench上的实验表明,EARL达到65.48%的cIoU像素定位准确率,较此前基于强化学习的方法提升8.37%;在EgoHOS上的域外定位结果也显示其具备强泛化能力,适用于未见的第一视角定位场景。

原文摘要 · Abstract (English)

Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing multimodal large language models (MLLMs) still struggle with accurate interaction reasoning and fine-grained pixel grounding. To this end, this paper introduces EARL, an Egocentric Analysis-guided Reinforcement Learning framework that explicitly transfers coarse interaction semantics to query-oriented answering and grounding. Specifically, EARL adopts a two-stage parsing framework including coarse-grained interpretation and fine-grained response. The first stage holistically interprets egocentric interactions and generates a structured textual description. The second stage produces the textual answer and pixel-level mask in response to the user query. To bridge the two stages, we extract a global interaction descriptor as a semantic prior, which is integrated via a novel Analysis-guided Feature Synthesizer (AFS) for query-oriented reasoning. To optimize heterogeneous outputs, including textual answers, bounding boxes, and grounding masks, we design a multi-faceted reward function and train the response stage with GRPO. Experiments on Ego-IRGBench show that EARL achieves 65.48% cIoU for pixel grounding, outperforming previous RL-based methods by 8.37%, while OOD grounding results on EgoHOS indicate strong transferability to unseen egocentric grounding scenarios.

第一视角强化学习像素定位交互推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。