arXiv:2608.08929cs.CV2026-08

让模型零样本识别新交互类型,提升真实场景泛化能力

Zero-shot 2D Grounding with Novel Affordance Types

论文配图:Zero-shot 2D Grounding with Novel Affordance Types
图 1 · 摘自论文原文
  • 利用物体子部件与交互区域强相关性,通过分割线索实现零样本定位
  • 在新交互类型基准上比当前最佳方法提升12.3%([email protected]
  • 无需训练的AffordAnything适合快速部署,可扩展的AffordAnything+适合优化

2D交互属性定位旨在识别人类可交互的物体区域。现有研究仅关注训练中见过的交互类型,未考察模型对新交互类型的泛化能力,而这在真实应用中至关重要。我们提出零样本2D交互属性定位新任务(NAT),并构建了NAT基准数据集。提出AffordAnything方法,无需训练,利用分割线索,基于交互区域与物体子部件的强相关性。为进一步提升性能,开发可训练的变体AffordAnything+,学习融合这些线索。在提出的AGD20K-NAT基准上,最优模型AffordAnything+在[email protected]指标上比当前最优方法OOAL提升12.3个百分点。

原文摘要 · Abstract (English)

2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types seen during training and does not study models' ability to generalize to novel affordances, which is crucial for real-world applications. We propose the task of zero-shot 2D grounding with novel affordance types (NAT) and introduce the NAT benchmarks. We then propose AffordAnything, a training-free method that leverages segmentation cues, motivated by the strong correlation between affordance regions and object subparts. To further improve performance, we develop AffordAnything+, a trainable variant that learns to combine these cues. On the proposed AGD20K-NAT benchmark, our best model AffordAnything+ achieves a substantial improvement of 12.3% (absolute) in [email protected] over the SOTA affordance grounding method, OOAL.

零样本交互属性目标定位分割线索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。