arXiv:2608.20720cs.CV2026-08

仅用一张照片就能识别物体功能部位,支持开放世界语言查询。

AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning

论文配图:AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning
图 1 · 摘自论文原文
  • 用单目图像构建文本引导的3D部件监督信号
  • 未见过物体和类别时达0.428和0.315的定位精度
  • 无需人工标注即可通过伪标签自训练拓展新物体

开放世界3D功能定位需在给定自由语言查询时定位3D物体的功能部件。现有方法通常依赖预构建的以物体为中心的3D几何结构和封闭的功能本体,限制了从原始RGB图像直接部署的能力。我们提出AffordAny,一个端到端框架,仅使用单张RGB图像即可构建大规模文本条件的3D部件监督数据,利用冻结的视觉-语言模型(VLM)引导解码器进行功能定位,并通过伪标签自训练提升开放世界泛化能力。自动化流水线生成包含5,334个物体和10,633个部件级样本的数据集,覆盖473个类别,类别多样性较之前提升一个数量级。解码器通过空间投影、指令条件语义压缩及双向几何-语义交互,逐步融合冻结的Cosmos-2B特征与3D几何信息。最小扰动伪标签自训练进一步实现无人工标注的新物体扩展。在系统性泛化评估协议下(涵盖未见物体、未见类别、未见指令改写),自训练后在未见物体上达到0.428 IoU,未见类别上达0.315 IoU,未见类别mIoU相对提升6.3%(p<0.01),指令敏感度差距仅为0.105,验证了方法的有效性与鲁棒性。

原文摘要 · Abstract (English)

Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assume pre-built object-centric 3D geometry and closed affordance ontologies, limiting deployment from raw RGB observations. We present AffordAny, an end-to-end framework that uses one monocular RGB image to construct large-scale text-conditioned 3D part supervision, ground affordances with a frozen vision-language model (VLM) guided decoder, and improve open-world generalization through pseudo-label self-training. Our automated pipeline produces a benchmark of 5,334 objects and 10,633 part-level samples spanning 473 categories, an order-of-magnitude increase in categorical diversity over prior work. The decoder progressively fuses frozen Cosmos-2B features with 3D geometry through spatial projection, instruction-conditioned semantic compression, and bidirectional geometry-semantics interaction. Minimal-perturbation pseudo-label self-training further adds new objects without human annotation. Under a systematic generalization protocol evaluating unseen objects, unseen categories, and unseen instruction paraphrases, our approach achieves 0.428 IoU on unseen objects and 0.315 IoU on unseen categories after self-training, with unseen-category mIoU improving by 6.3% relative (p<0.01) and an instruction sensitivity gap of only 0.105, demonstrating effectiveness and robustness of our method.

3D定位视觉语言开放世界自训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。