通过学习提示分布提升交互检测的细粒度识别能力
Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection
- 用多组软提示学习类别内动态与跨类别关系
- 在HICO-DET和vcoco上达到领先性能,参数增量极小
- 可插拔适配主流检测器,适合关注复杂交互识别的研究者
基于查询-变换器架构的人机交互(HOI)检测器虽表现良好,但对罕见视觉模式的识别和模糊交互的区分仍存困难。我们发现这可能源于传统检测查询在表征类别内多样性与类别间依赖关系上的能力有限。为此,提出交互提示分布学习(InterProDa)方法:学习多组软提示,并从不同提示中估计类别分布;将这些分布融入HOI查询,使其能表征近乎无限的类别内动态与通用的跨类别关系。所提模型在HICO-DET与vcoco基准上表现优异,且可无缝集成至多数基于Transformer的HOI检测器中,在几乎不增加参数的前提下显著提升性能。
原文摘要 · Abstract (English)
Human-object interaction (HOI) detectors with popular query-transformer architecture have achieved promising performance. However, accurately identifying uncommon visual patterns and distinguishing between ambiguous HOIs continue to be difficult for them. We observe that these difficulties may arise from the limited capacity of traditional detector queries in representing diverse intra-category patterns and inter-category dependencies. To address this, we introduce the Interaction Prompt Distribution Learning (InterProDa) approach. InterProDa learns multiple sets of soft prompts and estimates category distributions from various prompts. It then incorporates HOI queries with category distributions, making them capable of representing near-infinite intra-category dynamics and universal cross-category relationships. Our InterProDa detector demonstrates competitive performance on HICO-DET and vcoco benchmarks. Additionally, our method can be integrated into most transformer-based HOI detectors, significantly enhancing their performance with minimal additional parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。