arXiv:2504.01472cs.CV2025-04CVPR被引 9

提出新任务与数据集,实现第一人称交互的文本与像素级联合理解。

ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction

  • 构建分析-回答-像素定位三步统一框架,支持多粒度交互理解。
  • 创建包含20,000张图像、160万查询的Ego-IRGBench数据集。
  • 基于大模型的ANNEXE模型可生成流畅文本与精准像素响应,适合智能系统研发。

第一人称交互感知是研究人-环境交互的重要方向,为下一代智能系统奠定基础。然而,现有方法无法根据用户查询同时生成连贯的文本和像素级响应,难以适应多样下游需求。为此,本文提出首个综合任务Ego-IRG(Egocentric Interaction Reasoning and pixel Grounding),以第一人称图像与查询为输入,通过分析、回答与像素定位三步,实现流畅文本与细粒度像素响应的联合输出。针对现有数据集不满足该任务的问题,本文基于大量人工标注构建了Ego-IRGBench数据集,包含超过20,000张第一人称图像、160万条查询及对应的多模态响应。此外,设计统一的ANNEXE模型,利用多模态大语言模型生成文本与像素级输出,实现对第一人称交互的全面解析。在Ego-IRGBench上的实验验证了该模型的有效性。

原文摘要 · Abstract (English)

Egocentric interaction perception is one of the essential branches in investigating human-environment interaction, which lays the basis for developing next-generation intelligent systems. However, existing egocentric interaction understanding methods cannot yield coherent textual and pixel-level responses simultaneously according to user queries, which lacks flexibility for varying downstream application requirements. To comprehend egocentric interactions exhaustively, this paper presents a novel task named Egocentric Interaction Reasoning and pixel Grounding (Ego-IRG). Taking an egocentric image with the query as input, Ego-IRG is the first task that aims to resolve the interactions through three crucial steps: analyzing, answering, and pixel grounding, which results in fluent textual and fine-grained pixel-level responses. Another challenge is that existing datasets cannot meet the conditions for the Ego-IRG task. To address this limitation, this paper creates the Ego-IRGBench dataset based on extensive manual efforts, which includes over 20k egocentric images with 1.6 million queries and corresponding multimodal responses about interactions. Moreover, we design a unified ANNEXE model to generate text- and pixel-level outputs utilizing multimodal large language models, which enables a comprehensive interpretation of egocentric interactions. The experiments on the Ego-IRGBench exhibit the effectiveness of our ANNEXE model compared with other works.

第一人称视觉多模态理解像素定位大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。