让模型零样本识别新交互类型,提升真实场景泛化能力
Zero-shot 2D Grounding with Novel Affordance Types

- 利用物体子部件与交互区域强相关性,通过分割线索实现零样本定位
- 在新交互类型基准上比当前最佳方法提升12.3%([email protected])
- 无需训练的AffordAnything适合快速部署,可扩展的AffordAnything+适合优化
2D交互属性定位旨在识别人类可交互的物体区域。现有研究仅关注训练中见过的交互类型,未考察模型对新交互类型的泛化能力,而这在真实应用中至关重要。我们提出零样本2D交互属性定位新任务(NAT),并构建了NAT基准数据集。提出AffordAnything方法,无需训练,利用分割线索,基于交互区域与物体子部件的强相关性。为进一步提升性能,开发可训练的变体AffordAnything+,学习融合这些线索。在提出的AGD20K-NAT基准上,最优模型AffordAnything+在[email protected]指标上比当前最优方法OOAL提升12.3个百分点。
原文摘要 · Abstract (English)
2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types seen during training and does not study models' ability to generalize to novel affordances, which is crucial for real-world applications. We propose the task of zero-shot 2D grounding with novel affordance types (NAT) and introduce the NAT benchmarks. We then propose AffordAnything, a training-free method that leverages segmentation cues, motivated by the strong correlation between affordance regions and object subparts. To further improve performance, we develop AffordAnything+, a trainable variant that learns to combine these cues. On the proposed AGD20K-NAT benchmark, our best model AffordAnything+ achieves a substantial improvement of 12.3% (absolute) in [email protected] over the SOTA affordance grounding method, OOAL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。