用图像扩散模型生成交互检测结果,提升复杂场景下的识别准确率。
An Image-like Diffusion Method for Human-Object Interaction Detection
- 将人-物交互输出视为图像,用扩散模型生成检测结果。
- 在HICO-DET数据集上达到新高,显著优于传统方法。
- 适合研究视觉理解与生成模型融合的学者使用。
人-物交互(HOI)检测常因语义模糊、外观差异大而面临挑战,尤其在遮挡和背景杂乱时更为明显。本文提出新思路:将每个交互对的检测输出重新建模为一种图像形式。受图像扩散模型强大生成能力启发,我们构建了HOI-IDiff框架,通过类图像扩散过程生成交互检测结果。针对此类“重建成像”与真实图像的差异,设计了定制化的扩散流程与切片式分块架构,专门优化交互图像生成。大量实验表明该方法在标准数据集上表现优异,验证了其有效性。
原文摘要 · Abstract (English)
Human-object interaction (HOI) detection often faces high levels of ambiguity and indeterminacy, as the same interaction can appear vastly different across different human-object pairs. Additionally, the indeterminacy can be further exacerbated by issues such as occlusions and cluttered backgrounds. To handle such a challenging task, in this work, we begin with a key observation: the output of HOI detection for each human-object pair can be recast as an image. Thus, inspired by the strong image generation capabilities of image diffusion models, we propose a new framework, HOI-IDiff. In HOI-IDiff, we tackle HOI detection from a novel perspective, using an Image-like Diffusion process to generate HOI detection outputs as images. Furthermore, recognizing that our recast images differ in certain properties from natural images, we enhance our framework with a customized HOI diffusion process and a slice patchification model architecture, which are specifically tailored to generate our recast ``HOI images''. Extensive experiments demonstrate the efficacy of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。