通过分步推理提升复杂场景中细粒度3D功能区域定位精度
ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

- 先生成高召回率的交互区域候选,再结合语言指令精炼定位
- 在SceneFun3D上达25.46% AP25,优于现有2D到3D基线方法
- 适合需要精准操作理解的机器人视觉任务,如智能家居交互
面向自然语言指令的任务驱动3D功能区域定位,旨在从杂乱场景中找到可执行动作的功能区域。现有方法或直接预测3D掩码,或通过选择和融合中间2D/3D区域构建掩码,但易受两类耦合失败影响:预测区域可能遗漏目标交互区或粒度过粗;语言定位在关系性指令下易混淆视觉相似项。为此,提出ThinkAfford,将高召回率的功能区域生成与指令感知推理解耦。首先,通过可学习的功能提示与多层级视觉特征生成交互条件热图,提取可变数量的细粒度候选,无需物体或部件名称作为分割提示。其次,利用完整指令对标注候选叠加进行视觉-提示推理,以“思考-作答”结构输出标识符。此外,采用提案级奖励的组相对策略优化(GRPO),基于提升后的3D重叠度对推理选择进行对齐。在SceneFun3D验证集上,ThinkAfford取得10.69% AP50与25.46% AP25,超越同类3D开放词汇及基于视觉语言模型的2D-to-3D基线。模块诊断显示,APG在25%交并比下达到77.5%召回率,而GRPO训练的VPAR在覆盖查询上实现72.1%选择准确率,高于监督微调下的63.4%。
原文摘要 · Abstract (English)
Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing methods either predict 3D masks directly or construct them by selecting and fusing intermediate 2D/3D regions. However, they remain vulnerable to two intertwined failure modes: the predicted or selected regions may miss the target interaction area or have unsuitable granularity, while language grounding may confuse visually similar alternatives under relational instructions. To this end, we introduce ThinkAfford, which decouples high-recall affordance proposal generation from instruction-grounded reasoning. Specifically, the Affordance Proposal Generation module first uses learnable affordance prompts and multi-level visual features to predict interaction-conditioned heatmaps, extracting a variable number of fine-grained proposals without parsed object or part names as segmentation prompts. Visual-Prompted Affordance Reasoning then reasons over labeled proposal overlays using the full instruction, returning identifiers in a structured "think-then-answer" response. Moreover, Group Relative Policy Optimization uses proposal-level rewards from lifted 3D overlap to align VPAR selection with final 3D grounding. On the SceneFun3D validation split, ThinkAfford achieves 10.69% AP50 and 25.46% AP25 under the official evaluator, outperforming comparable 3D open-vocabulary and vision-language-model-based 2D-to-3D baselines. Module-level diagnostics further show that APG attains 77.5% recall at 25% intersection-over-union, while GRPO-trained VPAR achieves 72.1% selection accuracy on APG-covered queries, compared with 63.4% under supervised fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。