仅用一张照片就能识别物体功能部位,支持开放世界语言查询。
AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning

- 用单目图像构建文本引导的3D部件监督信号
- 未见过物体和类别时达0.428和0.315的定位精度
- 无需人工标注即可通过伪标签自训练拓展新物体
开放世界3D功能定位需在给定自由语言查询时定位3D物体的功能部件。现有方法通常依赖预构建的以物体为中心的3D几何结构和封闭的功能本体,限制了从原始RGB图像直接部署的能力。我们提出AffordAny,一个端到端框架,仅使用单张RGB图像即可构建大规模文本条件的3D部件监督数据,利用冻结的视觉-语言模型(VLM)引导解码器进行功能定位,并通过伪标签自训练提升开放世界泛化能力。自动化流水线生成包含5,334个物体和10,633个部件级样本的数据集,覆盖473个类别,类别多样性较之前提升一个数量级。解码器通过空间投影、指令条件语义压缩及双向几何-语义交互,逐步融合冻结的Cosmos-2B特征与3D几何信息。最小扰动伪标签自训练进一步实现无人工标注的新物体扩展。在系统性泛化评估协议下(涵盖未见物体、未见类别、未见指令改写),自训练后在未见物体上达到0.428 IoU,未见类别上达0.315 IoU,未见类别mIoU相对提升6.3%(p<0.01),指令敏感度差距仅为0.105,验证了方法的有效性与鲁棒性。
原文摘要 · Abstract (English)
Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assume pre-built object-centric 3D geometry and closed affordance ontologies, limiting deployment from raw RGB observations. We present AffordAny, an end-to-end framework that uses one monocular RGB image to construct large-scale text-conditioned 3D part supervision, ground affordances with a frozen vision-language model (VLM) guided decoder, and improve open-world generalization through pseudo-label self-training. Our automated pipeline produces a benchmark of 5,334 objects and 10,633 part-level samples spanning 473 categories, an order-of-magnitude increase in categorical diversity over prior work. The decoder progressively fuses frozen Cosmos-2B features with 3D geometry through spatial projection, instruction-conditioned semantic compression, and bidirectional geometry-semantics interaction. Minimal-perturbation pseudo-label self-training further adds new objects without human annotation. Under a systematic generalization protocol evaluating unseen objects, unseen categories, and unseen instruction paraphrases, our approach achieves 0.428 IoU on unseen objects and 0.315 IoU on unseen categories after self-training, with unseen-category mIoU improving by 6.3% relative (p<0.01) and an instruction sensitivity gap of only 0.105, demonstrating effectiveness and robustness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。