让AI通过逐步获取3D几何和2D语义信息,更准地理解复杂场景中的动作可能性。
A3R: Agentic Affordance Reasoning via Cross-Dimensional Evidence in 3D Gaussian Scenes
- 将动作推理变为多步获取证据的过程,结合3D与2D信息。
- 在多个场景基准上优于静态方法,准确率显著提升。
- 适合做智能机器人场景理解、具身AI任务规划的开发者参考。
3D Gaussian场景中的动作可能性推理旨在从复杂环境中识别出支持给定文本指令所指定动作的区域。现有方法通常将其视为基于静态场景观察的一次性预测,假设任务相关证据已充分存在。然而,在复杂3D场景中,许多失败并非源于预测能力不足,而是固定观测下任务相关证据不完整所致。为此,我们将细粒度动作可能性推理重新建模为一个序列化证据获取过程,通过互补的3D几何与2D语义证据逐步消除模糊性。在此框架下,我们提出A3R,一种基于多模态大模型(MLLM)策略的智能体式动作可能性推理框架,可迭代选择证据获取动作,并通过跨维度证据融合更新动作信念。为优化此类序列决策,我们进一步引入基于GRPO的策略学习方法,提升证据获取效率与推理准确性。在场景级基准上的大量实验表明,A3R持续超越静态一次性基线,验证了智能体式跨维度证据获取在复杂3D Gaussian场景中细粒度动作可能性推理中的优势。
原文摘要 · Abstract (English)
Affordance reasoning in 3D Gaussian scenes aims to identify the region that supports the action specified by a given text instruction in complex environments. Existing methods typically cast this problem as one-shot prediction from static scene observations, assuming sufficient evidence is already available for reasoning. However, in complex 3D scenes, many failure cases arise not from weak prediction capacity, but from incomplete task-relevant evidence under fixed observations. To address this limitation, we reformulate fine-grained affordance reasoning as a sequential evidence acquisition process, where ambiguity is progressively reduced through complementary 3D geometric and 2D semantic evidence. Building on this formulation, we propose A3R, an agentic affordance reasoning framework that enables an MLLM-based policy to iteratively select evidence acquisition actions and update the affordance belief through cross-dimensional evidence acquisition. To optimize such sequential decision making, we further introduce a GRPO-based policy learning strategy that improves evidence acquisition efficiency and reasoning accuracy. Extensive experiments on scene-level benchmarks show that A3R consistently surpasses static one-shot baselines, demonstrating the advantage of agentic cross-dimensional evidence acquisition for fine-grained affordance reasoning in complex 3D Gaussian scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。