用分割掩码增强人物交互预测,支持零样本新提示
Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
- 将分割基础模型与交互任务结合,引入包含掩码的四元组
- 在两个公开数据集上达到顶尖性能,零样本下仍有效
- 可响应未训练过的文本/视觉提示,适合多场景应用
本文提出一种名为 Seg2HOI 的新框架,通过整合基于分割的视觉基础模型与人-物交互(HOI)任务,区别于传统检测式方法。该方法不仅预测标准三元组,还引入四元组,即在人-物对基础上增加分割掩码信息。Seg2HOI继承基础模型的可提示性与交互能力,通过专用解码器将其应用于HOI任务。尽管仅针对HOI任务训练,无需额外机制,其特性仍能高效运行。在两个公开基准数据集上的大量实验表明,即使在零样本场景下,性能也与当前最先进方法相当。最后,我们证明该框架可生成未训练过的文本和视觉提示下的四元组及交互分割结果,具备广泛的应用灵活性。
原文摘要 · Abstract (English)
In this work, we introduce Segmentation to Human-Object Interaction (\textit{\textbf{Seg2HOI}}) approach, a novel framework that integrates segmentation-based vision foundation models with the human-object interaction task, distinguished from traditional detection-based Human-Object Interaction (HOI) methods. Our approach enhances HOI detection by not only predicting the standard triplets but also introducing quadruplets, which extend HOI triplets by including segmentation masks for human-object pairs. More specifically, Seg2HOI inherits the properties of the vision foundation model (e.g., promptable and interactive mechanisms) and incorporates a decoder that applies these attributes to HOI task. Despite training only for HOI, without additional training mechanisms for these properties, the framework demonstrates that such features still operate efficiently. Extensive experiments on two public benchmark datasets demonstrate that Seg2HOI achieves performance comparable to state-of-the-art methods, even in zero-shot scenarios. Lastly, we propose that Seg2HOI can generate HOI quadruplets and interactive HOI segmentation from novel text and visual prompts that were not used during training, making it versatile for a wide range of applications by leveraging this flexibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。